Skip to content

feat(mssql): new availability group wizard (SE-247, part 2) - #201

Merged
Lionear merged 2 commits into
developfrom
feature/SE-247-new-ag-wizard
Sep 30, 2026
Merged

Lionear merged 2 commits into
developfrom
feature/SE-247-new-ag-wizard

Conversation

@Lionear

@Lionear Lionear commented Sep 30, 2026

Copy link
Copy Markdown
Owner

Stacked on #200. Merge #200 first. Until then this PR's diff also shows #200's commit; the new work is the second commit (feat(mssql): create an availability group…).

What

New Availability Group… on the Availability Groups folder (SE-247, part 2 of 2). The instance the tool was opened on becomes the primary; the user ticks the saved connections that hold the secondaries.

The pre-flight is a gate, not a page

While anything blocks, the dialog shows only the checks and why DataTray cannot fix them from a connection — no form at all. Blocking:

  • fewer than two instances, or two connections reaching the same instance
  • Always On not enabled (with where to switch it on: Configuration Manager / mssql-conf)
  • an edition without availability groups; SQL Server older than 2016
  • different major versions or server collations across instances (Learn: "same version", "same collation")
  • Standard mixed with Enterprise/Developer; more than two replicas on Standard (basic group)
  • a login without CONTROL SERVER
  • mirroring endpoints on only some instances, or a stopped one
  • no database that can join

Once it passes, the form offers only what can work: WSFC only when every instance is a node of the same Windows cluster; NONE (read-scale) only on 2017+, and never preselected because it is not HA; EXTERNAL never (Pacemaker is not something DataTray configures); automatic failover only under WSFC and only on synchronous replicas; backup preference restricted to PRIMARY for a basic group. Ineligible databases are listed with the reason instead of a disabled box (not FULL recovery, no full backup, already in a group, mirrored, read-only, AUTO_CLOSE, not multi-user, or a same-named database already on a secondary, which automatic seeding needs absent).

The plan

Per instance, in order, shown in full and rebuilt on every change:

  1. When no instance has an endpoint: master key (if missing), a certificate and a certificate-authenticated DATABASE_MIRRORING endpoint on each; then each instance's public certificate read with CERTENCODED() and created on every other one FROM BINARY with a login/user and GRANT CONNECT ON ENDPOINT. When every instance already has an endpoint, they are reused untouched.
  2. CREATE AVAILABILITY GROUP … SEEDING_MODE = AUTOMATIC on the primary.
  3. On each secondary: JOIN (JOIN WITH (CLUSTER_TYPE = NONE) for read-scale), then GRANT CREATE ANY DATABASE.

ExecuteAsync re-reads every instance, re-runs the pre-flight and refuses unless the rebuilt script equals the reviewed one. Generated passwords travel with the choices so the script is byte-identical.

Decisions on the ticket's open points

  • Certificate exchange, cross-domain: no file moves at all — CERTENCODED() → CREATE CERTIFICATE … FROM BINARY (documented since 2012). No BACKUP CERTIFICATE, no UNC share. Endpoints DataTray creates always use certificates, the one mode that works in a domain, across domains, without one, and on Linux.
  • Seeding progress: not in the wizard. The run ends once the secondaries have joined; the SE-284 dashboard already shows per-database state and polls.
  • Read-only routing, listeners, distributed and contained groups: out of scope. A WSFC listener needs AD rights DataTray cannot check; read-only routing is its own feature (and needs ApplicationIntent in the connection fields first — see Codebase notes).
  • Existing endpoints: reuse all or none; a mix blocks, because joining an endpoint DataTray did not make to ones it did fails in ways nobody can diagnose from here.

Shared with #200

The step runner moves to AgStepRunner (wait steps, retried steps, and now captured values: $(cert:NAME) is filled at run time from an earlier step's query; only that prefix is substituted, and an uncaptured token throws instead of sending empty SQL). The failover tool uses the same runner and script rendering.

Verification

  • DataTray.Tools.MsSqlAdmin.Tests: 140 passed (29 new, on top of feat(mssql): fail over an availability group (SE-247, part 1) #200). DataTray.Core.Tests 629, DataTray.App.Tests 56. Solution builds with 0 warnings.
  • Live, against throwaway SQL Server 2025 containers on free ports (my own, removed afterwards):
    • fresh pair, no endpoints → 9 steps: endpoints + certificate exchange + create + join + grant; group HEALTHY, database SYNCHRONIZED on both sides, endpoints CERTIFICATE/STARTED;
    • pair that already had endpoints → reuse path, 3 steps, second group next to the first, both HEALTHY;
    • the group the wizard made, failed over with feat(mssql): fail over an availability group (SE-247, part 1) #200's tool → roles swapped, data intact on the new primary.
  • Both dialog states (gate, filled form) rendered headlessly with a temporary DataTray.Screenshots scene (not committed).
  • Not live-tested: WSFC creation (needs a Windows cluster), a basic group on Standard edition, and whether a basic group accepts CLUSTER_TYPE = NONE — the docs do not say, so the wizard currently offers it. Live testing is planned on Rick's Proxmox lab.

Fail Over… on an availability group builds the plan from cluster_type_desc:
one FAILOVER (or FORCE_FAILOVER_ALLOW_DATA_LOSS) on the target for WSFC, the
documented multi-instance sequence for read-scale (NONE), and a refusal with
the pcs/crm commands for Pacemaker (EXTERNAL). The plan is data, shown in full
before anything runs, re-read and compared at confirm time, and every picked
connection must prove it is the replica it was picked for.

The old primary's SET (ROLE = SECONDARY) is retried for up to 30 s: issued the
moment the promotion returns it fails while that replica is still resolving,
seen against a real read-scale group.

IToolUiContext gains OpenConnection so the view can read the primary's state
when opened on a secondary (additive, folded into tool API 8).
…s folder

New Availability Group… makes the launch instance the primary and picked
saved connections the secondaries. A pre-flight over every instance gates the
form: nothing is offered that the topology cannot do (one instance, Always On
off, mixed versions/collations/editions, missing CONTROL SERVER, endpoints on
only some instances), WSFC only when all instances are nodes of one cluster,
read-scale never preselected, a basic group on Standard edition, and only
databases that can join.

Endpoints are created with certificate authentication and the public
certificates move as CERTENCODED() -> CREATE CERTIFICATE FROM BINARY, so no
file is written on a server and the cross-domain case needs no share.
Existing endpoints are reused when every instance has one. Seeding is
automatic.

The plan runner is shared with the failover tool and gains captured values:
a step can read a certificate on one instance for a later step to create on
another.
@Lionear
Lionear merged commit 553a198 into develop Sep 30, 2026
2 checks passed
@Lionear
Lionear deleted the feature/SE-247-new-ag-wizard branch September 30, 2026 08:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant