Skip to content

feat(mssql): fail over an availability group (SE-247, part 1) - #200

Merged
Lionear merged 1 commit into
developfrom
feature/SE-247-ag-failover
Sep 30, 2026
Merged

Lionear merged 1 commit into
developfrom
feature/SE-247-ag-failover

Conversation

@Lionear

@Lionear Lionear commented Sep 30, 2026

Copy link
Copy Markdown
Owner

What

Fail Over… on an availability group node (SE-247, part 1 of 2 — the new-AG wizard follows in its own PR). Pick the target replica, map the replicas the plan runs on to saved connections, read the exact plan, confirm.

Which plan you get follows cluster_type_desc only (Rick's decision), never the OS:

Cluster type Plan DataTray
WSFC one statement on the target: FAILOVER when both ends are synchronous-commit and every database on the target is failover-ready, otherwise FORCE_FAILOVER_ALLOW_DATA_LOSS with a data-loss warning and per-database "behind by" runs it
NONE (read-scale) the documented sequence: both replicas synchronous → wait for SYNCHRONIZED → REQUIRED_SYNCHRONIZED_SECONDARIES_TO_COMMIT = 1 → OFFLINE on the primary → FORCE_FAILOVER_ALLOW_DATA_LOSS on the target → SET (ROLE = SECONDARY) on the old primary → HADR RESUME per database → listener re-created runs it, as a visible plan
EXTERNAL (Pacemaker) refused; shows the pcs/crm commands with the target node filled in never runs anything

Safety

  • Nothing runs unseen. The view shows every statement and the instance it runs on. A forced or read-scale failover also needs the group name typed; the host's destructive-action confirm comes on top (IsDestructive = true).
  • Re-checked at confirm. ExecuteAsync re-reads the DMVs, rebuilds the plan and refuses unless it is byte-identical to the script the user reviewed — a target that fell behind turns a planned failover into a forced one, and that must not run on the old confirmation.
  • Right instance. Every picked connection has to prove it is that replica (SERVERPROPERTY('ServerName')) before it is used. All connections a plan needs are resolved before the first statement, so a read-scale plan can never stop after OFFLINE for want of the target's connection.
  • Waits before the irreversible part. The read-scale plan polls until the target is SYNCHRONIZED (2 min cap) and aborts before OFFLINE if it never gets there.
  • Not reachable over MCP. It is a tool; the MCP surface exposes no tools. Its statements are all DDL to McpSqlClassifier, which no AI access mode but a transient Sandbox connection allows — now pinned in McpSqlClassifierTests. I did not change the Sandbox rule itself.

Design notes

  • Plan as data. AvailabilityGroupFailover.Plan(topology, target) is pure: every branch (three cluster types, planned vs forced, refusals, escaping, the listener, the review mismatch) is unit-tested without a server. AvailabilityGroupQueries parses DMV rows with the SE-284 lessons (lowercase _desc columns, small ints through Convert, names from the cluster-states DMV).
  • Reuse of SE-284. The plugin reaches SQL Server only through the host's IDbProvider and loads in its own ALC, so it cannot reference the provider. AvailabilityGroupStatus (the "behind by" derivation) is source-linked, like plugins/Shared.Schema. The queries themselves are modelled on the dashboard's but could not be literally shared — the dashboard uses SqlConnection with parameters.
  • One additive SDK member: IToolUiContext.OpenConnection (default null), mirroring IToolHost.OpenConnection, so the view can read the group's state from the primary when it was opened on a secondary. Folded into tool API 8 by the existing "does it add types?" rule; mssql-admin now declares 8.
  • Checked against Microsoft Learn (the sql-docs sources): the planned/forced preconditions, the read-scale sequence, and the Pacemaker commands (pcs resource move … --master for both the -clone and -master resource forms, crm resource migrate ag_cluster … for SLES, and removing the constraint each move leaves behind). After a forced failover the plan advises a database snapshot before resuming, as the docs do, not a backup.
  • One deliberate deviation: the read-scale page says to run SET HADR RESUME "on the primary"; Resume an availability database says a locally suspended secondary database is resumed on its own replica. The plan does the latter (the old primary, now a secondary), and that is what worked against the lab.
  • Deliberate limits: read-scale failover only for exactly two replicas (the documented and lab-verified shape); EXTERNAL's forced emergency path stays out of scope; neither the Pacemaker resource name nor its node name is visible from SQL Server, so the commands use the documented ag_cluster and the replica name, and say to check both.

Found by running it

Against a real two-node CLUSTER_TYPE = NONE group (tools/mssql-ag-lab, own containers on free ports): SET (ROLE = SECONDARY) issued the moment the promotion returns fails on the old primary — "the availability group resource did not come online due to a previous error" — and succeeds a few seconds later. That step now retries for up to 30 s; nothing else is retried.

Verification

  • DataTray.Tools.MsSqlAdmin.Tests: 111 passed (26 new). DataTray.Core.Tests 629 passed, DataTray.App.Tests 56 passed, solution builds with 0 warnings. DataTray.Core.Tests MCP classifier: +8 AG statements pinned as refused under ReadOnly/ReadWrite.
  • Live, against two throwaway SQL Server 2025 containers forming a read-scale AG: forward failover from the primary (target async → switched to sync first), failback launched from the secondary with the primary via a picked connection, a stale review refused, a wrong group name refused, a connection pointing at the wrong replica detected. Roles swapped each time, both replicas back to HEALTHY/SYNCHRONIZED, data intact.
  • Dialog rendered headlessly (DataTray.Screenshots, temporary scene, not committed) in light and dark.
  • Not live-tested: the WSFC and Pacemaker paths — covered by the generated SQL/commands only. Live testing is planned on Rick's Proxmox lab.

Fail Over… on an availability group builds the plan from cluster_type_desc:
one FAILOVER (or FORCE_FAILOVER_ALLOW_DATA_LOSS) on the target for WSFC, the
documented multi-instance sequence for read-scale (NONE), and a refusal with
the pcs/crm commands for Pacemaker (EXTERNAL). The plan is data, shown in full
before anything runs, re-read and compared at confirm time, and every picked
connection must prove it is the replica it was picked for.

The old primary's SET (ROLE = SECONDARY) is retried for up to 30 s: issued the
moment the promotion returns it fails while that replica is still resolving,
seen against a real read-scale group.

IToolUiContext gains OpenConnection so the view can read the primary's state
when opened on a secondary (additive, folded into tool API 8).
@Lionear
Lionear merged commit ad58804 into develop Sep 30, 2026
2 checks passed
@Lionear
Lionear deleted the feature/SE-247-ag-failover branch September 30, 2026 08:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant