feat(mssql): fail over an availability group (SE-247, part 1) - #200
Merged
Merged
Conversation
Fail Over… on an availability group builds the plan from cluster_type_desc: one FAILOVER (or FORCE_FAILOVER_ALLOW_DATA_LOSS) on the target for WSFC, the documented multi-instance sequence for read-scale (NONE), and a refusal with the pcs/crm commands for Pacemaker (EXTERNAL). The plan is data, shown in full before anything runs, re-read and compared at confirm time, and every picked connection must prove it is the replica it was picked for. The old primary's SET (ROLE = SECONDARY) is retried for up to 30 s: issued the moment the promotion returns it fails while that replica is still resolving, seen against a real read-scale group. IToolUiContext gains OpenConnection so the view can read the primary's state when opened on a secondary (additive, folded into tool API 8).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Fail Over… on an availability group node (SE-247, part 1 of 2 — the new-AG wizard follows in its own PR). Pick the target replica, map the replicas the plan runs on to saved connections, read the exact plan, confirm.
Which plan you get follows
cluster_type_desconly (Rick's decision), never the OS:WSFCFAILOVERwhen both ends are synchronous-commit and every database on the target is failover-ready, otherwiseFORCE_FAILOVER_ALLOW_DATA_LOSSwith a data-loss warning and per-database "behind by"NONE(read-scale)SYNCHRONIZED→REQUIRED_SYNCHRONIZED_SECONDARIES_TO_COMMIT = 1→OFFLINEon the primary →FORCE_FAILOVER_ALLOW_DATA_LOSSon the target →SET (ROLE = SECONDARY)on the old primary →HADR RESUMEper database → listener re-createdEXTERNAL(Pacemaker)pcs/crmcommands with the target node filled inSafety
IsDestructive = true).ExecuteAsyncre-reads the DMVs, rebuilds the plan and refuses unless it is byte-identical to the script the user reviewed — a target that fell behind turns a planned failover into a forced one, and that must not run on the old confirmation.SERVERPROPERTY('ServerName')) before it is used. All connections a plan needs are resolved before the first statement, so a read-scale plan can never stop afterOFFLINEfor want of the target's connection.SYNCHRONIZED(2 min cap) and aborts beforeOFFLINEif it never gets there.McpSqlClassifier, which no AI access mode but a transient Sandbox connection allows — now pinned inMcpSqlClassifierTests. I did not change the Sandbox rule itself.Design notes
AvailabilityGroupFailover.Plan(topology, target)is pure: every branch (three cluster types, planned vs forced, refusals, escaping, the listener, the review mismatch) is unit-tested without a server.AvailabilityGroupQueriesparses DMV rows with the SE-284 lessons (lowercase_desccolumns, small ints throughConvert, names from the cluster-states DMV).IDbProviderand loads in its own ALC, so it cannot reference the provider.AvailabilityGroupStatus(the "behind by" derivation) is source-linked, likeplugins/Shared.Schema. The queries themselves are modelled on the dashboard's but could not be literally shared — the dashboard usesSqlConnectionwith parameters.IToolUiContext.OpenConnection(defaultnull), mirroringIToolHost.OpenConnection, so the view can read the group's state from the primary when it was opened on a secondary. Folded into tool API 8 by the existing "does it add types?" rule;mssql-adminnow declares 8.sql-docssources): the planned/forced preconditions, the read-scale sequence, and the Pacemaker commands (pcs resource move … --masterfor both the-cloneand-masterresource forms,crm resource migrate ag_cluster …for SLES, and removing the constraint each move leaves behind). After a forced failover the plan advises a database snapshot before resuming, as the docs do, not a backup.SET HADR RESUME"on the primary"; Resume an availability database says a locally suspended secondary database is resumed on its own replica. The plan does the latter (the old primary, now a secondary), and that is what worked against the lab.EXTERNAL's forced emergency path stays out of scope; neither the Pacemaker resource name nor its node name is visible from SQL Server, so the commands use the documentedag_clusterand the replica name, and say to check both.Found by running it
Against a real two-node
CLUSTER_TYPE = NONEgroup (tools/mssql-ag-lab, own containers on free ports):SET (ROLE = SECONDARY)issued the moment the promotion returns fails on the old primary — "the availability group resource did not come online due to a previous error" — and succeeds a few seconds later. That step now retries for up to 30 s; nothing else is retried.Verification
DataTray.Tools.MsSqlAdmin.Tests: 111 passed (26 new).DataTray.Core.Tests629 passed,DataTray.App.Tests56 passed, solution builds with 0 warnings.DataTray.Core.TestsMCP classifier: +8 AG statements pinned as refused under ReadOnly/ReadWrite.HEALTHY/SYNCHRONIZED, data intact.DataTray.Screenshots, temporary scene, not committed) in light and dark.