Skip to content

feat: the service VIP on the WireGuard mesh (0049, third round) - #81

Open
marcos-mendez wants to merge 5 commits into
mainfrom
feat/mesh-vip
Open

marcos-mendez wants to merge 5 commits into
mainfrom
feat/mesh-vip

Conversation

@marcos-mendez

@marcos-mendez marcos-mendez commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

The service VIP of a replicated appliance pair on the WireGuard mesh, as decision 0049 and its third round (handbook#48, and the precision in handbook#49) decide. Reference: docs/vip.md.

Design

  • The spec field. appliance.vip is one IPv6 address, the same on both nodes of the pair. It is a /128 of the overlay prefix. spec validate refuses it if it is outside the prefix, is the node's own address, or falls inside a peer's allowed_ips. The primary's installer picks it, and the replica receives it with the other values the primary generates (0028). It is a /128 inside the holder region's /112, so it stays compatible with 0051's regions.

  • The role. There is no role field (third round, point 2). The primary is the node that holds the VIP: the node of the pair with the newest claim, neither fenced nor released. The other node is the replica. A node that declares no VIP replicates nothing. This works the same with or without a database. A Keel Web pair is the case it covers today.

  • Claims and the counter. A claim is signed with the node's Ed25519 key (label keel vip 1\n). It carries the VIP, the holder's overlay address and an epoch that only grows; equal epochs go to the lower public key. A node takes a claim only if:

    • it is newer than the one it holds;
    • it is signed by the key its trust store holds for the holder (0048's amendment);
    • it is for this mesh;
    • it comes from a peer at the address the claim names.

    keel signs a claim only for its own appliance.vip, so the signed VIP is the pair's declaration (point 3).

  • Who may move it.

    • Before etcd, keel vip promote does: release, then take, then announce on the members' channel (POST /v1/vip). Nodes that are not cloud advanced always learn of moves this way (point 4).

    • With etcd, a claim is one transaction that writes two keys:

      • /keel/<mesh>/vip/<vip>/epoch, the counter, kept for good;
      • /keel/<mesh>/vip/<vip>/holder, attached to the holder's lease.

      The transaction requires the counter key to still be at the revision the claimant read, and the holder key to be absent. A stale claim fails that comparison and is never written.

  • etcd: the lease is the fence. keel vip tend (keel-vip.service) works as follows:

    • The holder renews its lease every 2 s.
    • It drops the VIP 10 s after the last renewal etcd answered, counted from when that renewal was sent, and is then fenced.
    • Every member follows the keys.
    • The other node of the pair claims once the holder's lease has expired, but only if it has taken the counter's claim and is not fenced.
    • On a node with etcd, only the controller adds the address. Its ExecStopPost drops every VIP the node carries, so a dead controller never leaves an unrenewed VIP behind.
  • TTL 20 s, release at 10 s. etcd cannot expire the lease earlier than 20 s after the last renewal it answered. A holder cut off from the majority has therefore dropped the VIP at least 10 s before anyone can win it.

    • The 5 s election timeout: a re-election in the majority takes 5 to 10 s. Renewals fail during it, but etcd gives every lease its full TTL again on a leader change, so one re-election does not move the VIP.
    • A holder that is etcd's leader renews locally until it steps down, one election timeout. It still drops the VIP by about 15 s after the cut, while the new leader's lease expires no sooner than 30 s after it.
    • etcd's minimum TTL with this configuration is 8 s, well below 20 s.
  • The 250 ms partition. The cut-off primary drops the VIP within the release time (measured: 9.4 s). The replica claims after the lease expires (measured: 20.4 s). On heal, the old primary takes the newer claim, stays fenced and never claims again by itself (third round; 0049 "must never re-claim"). Only a promote run on that node clears the fence.

  • Manual promote.

    • Without etcd, every peer is asked for its epoch. The holder is asked to release for the next epoch; it drops the address first and answers after. If it does not answer, promote refuses unless --old-primary-gone is given.
    • With etcd, release revokes the holder's own lease. With --old-primary-gone, promote waits for the lease to expire, up to the TTL plus one election timeout, and never revokes another node's lease (point 5). If the claimant's own controller claims first after a release, promote reports that this node holds the VIP.
    • keel database promote moves the VIP first. Writes then never reach two primaries; a failed database step leaves a read-only primary that refuses them.
  • What status, inspect and diff show.

    • keel vip status, and the end of keel mesh status, show the role, epoch and holder, fenced or released, the lease, whether wg0 carries the VIP, and which peer it is routed to.
    • keel inspect writes appliance.vip from the emitted spec. Beside the spec it reports one line per VIP: the role, epoch, holder, fenced, carried and routed-to. tracker#57's Webmin follow-up reads this.
    • apply renders wg0.conf with the VIP routed to its holder, and inspect reads the peers back without it. A move is therefore never an overlay change and never drift.
  • The 0018 window and privileges. The move is exempt from the 0018 window (0029, 0049): only the VIP /128 in allowed_ips and the holder's own address change. The address is added with preferred_lft 0, so it answers connections but is never chosen as a source. The members' channel keeps the node's own address.

    • The network-facing side keeps keel#75's split. The unprivileged listener of keel-mesh-members forwards POST /v1/vip, and the root helper verifies the claim and runs ip and wg set.
    • The controller and the check are not listeners. They are sandboxed in units with only CAP_NET_ADMIN, writing only /var/lib/keel/vip and /etc/wireguard, shipped by keel-overlay-vip (Keel-Linux/common PR).

Seams (tests written against them)

  • keel.mesh.vip: claim order, role, the state file, routing in and out of wg0.conf. Pure (test_mesh_vip.py).
  • keel.mesh.vipmsg: signed messages (test_mesh_vip.py).
  • keel.mesh.vipnet: ip and wg through run and output, against FakeNet (test_mesh_vipmove.py).
  • keel.mesh.vipnode and vipserve: take, release, hold, announce, epochs, and every refusal of the channel. A pair and a third node in one process (vip_helpers.py, test_mesh_vipmove.py, test_mesh_vipedges.py).
  • keel.mesh.vipetcd: the compare-and-swap, the controller's renewal, fence, following and failover, all on FakeKv. The client's lease and transaction calls are tested against the fake gateway over real TLS (test_mesh_vipetcd.py, test_mesh_etcdclient.py).
  • The wiring: the members' channel, keel vip, keel database promote, spec validation, inspect, and apply's rendering (test_mesh_vipwiring.py).
  • End to end: tests/test_vip_netns.py and tests/vip_netns.py, in their own CI job vip / trixie.

Coverage: keel/mesh/vip* is 99 to 100% per module, and the whole package is 99%.

Integration test, measured

The test uses three network namespaces with real wg-quick and trixie's etcd 3.5.16. Every leg runs at 250 ms ±25 ms with 2% loss, measured first and judged by the minimum RTT, as in test_etcd_netns.py. In the last local run the link measured min 226 ms and mean 252 ms, with 3% round-trip loss. Every node runs the members' channel, with its listener under setpriv without capabilities, and the controller. The holders are sampled every 200 ms (max step 0.201 s, 439 samples).

Case Result
(a) A promotes C reaches the VIP at A
(b) planned promote to B A released first; downtime seen from C (longest gap in 50 ms pings) 2.8 s (4.4 s in another run)
(c) B cut off from the majority B dropped the VIP at 9.4 s; A carried it at 20.4 s; C reached it again at 21.2 s
at every sample at most one holder, never two
(d) healed for 30 s B never carried it again (fenced); C routes to A
(e) stale claim (B, epoch 2 < 3) C refused it with 409 on the channel; etcd's compare refused it; the counter is unchanged

Also in this PR

  • test: date the etcd tests' clock by the real one. Three etcd tests fail on main since 2026-10-06: openssl dates the certificates with the real clock while the tests judge renewal against a fixed NOW of 2026-10-04.

After the security review (c2e80b8, 76be548)

  • CRITICAL, the persisted monotonic time. Nothing about the last renewal is kept on disk now. The controller times it in memory on CLOCK_BOOTTIME and carries nothing after a start until it has renewed the lease itself. Any etcd error counts as no renewal, and carry_held checks the renewal's age again before adding the address. Tests cover a restarted holder, a rebooted holder whose state names its old lease while the replica holds the VIP, and an etcd error.
  • HIGH, pair only. keel vip pair ADDRESS creates a pair record signed by both members, or by a trust root in place of a member. Every claim carries the record, and every claim and release is checked against it. The following are refused:
    • a VIP that is a member's own address;
    • a VIP outside its region's /112;
    • a record that names other members than the one a node already keeps;
    • an epoch jump of more than 64 (finding 2 asked for a small window).
  • HIGH, release. A release is taken only from the other member named in the pair record.
  • HIGH, etcd RBAC: not enabled; this is a question for the maintainer. etcd takes a client certificate's CN as its user. Every member issues its own client certificates with its own intermediate CA (0048, third round, point 1), so any member could name itself any etcd user, root included. RBAC would stop no member until certificate issuance changes. Instead the defence in depth asked for is in place: every node verifies the signed claim and pair record of every etcd value. A holder key whose value is not a valid claim neither holds the VIP nor blocks a failover, because the claim's transaction compares that key's revision.
  • MEDIUM, promote. Before etcd, a promote carries the VIP only when a majority of the peers that answered took the claim. If no peer answers, or a majority refuses, it exits non-zero and does not carry. With etcd, the compare-and-swap is the acceptance.
  • MEDIUM, the allocator. It reserves every VIP the node knows or declares.
  • MEDIUM, the split. keel vip tend is now the root helper. It alone changes wg0, after checking the pair record. It starts keel vip control as a transient unit with a dynamic user and no capability, in the helper's network namespace. The socket is 0600 in a 0700 runtime directory, the helper checks the peer by SO_PEERCRED against the unit's MainPID, and the client key is handed over as a memfd. The helper holds CAP_NET_ADMIN, plus CAP_DAC_OVERRIDE so it can connect to the controller's 0600 socket, which the dynamic user owns. The netns CI job runs both as real systemd units and reads their hardening from the kernel: controller uid ≠ 0, CapEff and CapBnd 0, NoNewPrivs 1; helper CapBnd = NET_ADMIN|DAC_OVERRIDE.
  • LOW. carry_held re-checks the lease's age. StartLimitIntervalSec=0 is set in common#40.
  • Race found by the new test. A node that released the VIP is no longer fenced when its revoked lease is later found gone.
  • Test gap, the cut-off etcd leader. The partition now runs twice: once after moving etcd leadership to the primary (what etcdctl move-leader does), and once with the primary as a follower.

Local run (250 ms ±25 ms, 2% loss; link min 226 ms; 794 samples, max step 0.202 s, never two holders):

the primary drops the VIP the replica carries it the third node reaches it again
cut-off primary is etcd's leader 9.9 s 36.5 s 37.9 s
cut-off primary is a follower 8.1 s 24.3 s 24.4 s

Planned promote: 2.6 s of downtime from the third node.

For the security review

  • A trusted member of the mesh that is not in the pair can sign a claim naming the pair's VIP. Receivers cannot tell, because the pair restriction is the claimant's own signed declaration, made only for its own spec. Its power is the same as 0048's "any trusted member can vouch for a new node". Pinning the pair's keys at the first claim would close this; that is not decided.
  • A controller that is frozen rather than dead, such as a stopped container, keeps the address. 0020's watchdog is not built, and a container has none.
  • The rejoin of an old primary's data, and Webmin per role (tracker#57), are follow-ups. Here the old primary only drops the VIP and never claims it again.

Companion: Keel-Linux/common PR for keel-overlay-vip (units only).

navigator added 2 commits October 7, 2026 14:42
openssl dates the certificates these tests issue by the real clock,
while the renewal is judged against a fixed NOW of 2026-10-04: from
2026-10-06 a leaf issued "now" had more than a third of its life left
at NOW plus two thirds of it, and three tests failed on main. NOW is
the real time, to the second, so the two clocks agree.
appliance.vip names a replicated pair's VIP, a /128 of the overlay
prefix. The primary is its holder: it carries the address on wg0
(deprecated, so never a source address), and every other node routes it
to the primary through that peer's allowed_ips. A move is one wg set per
node and the address on the holder, outside 0018's window (0029, 0049).

Claims are signed with the node's Ed25519 key and ordered by an epoch
that only grows (ties to the lower key); a node takes only a newer claim
signed by a key its trust store holds, from a peer at the address the
claim names. keel vip promote asks the old primary to release first,
refuses when it does not answer unless --old-primary-gone, then claims
at the next epoch and announces it on the members' channel
(POST /v1/vip). keel database promote moves the VIP first.

With etcd, keel vip tend renews the holder's lease (TTL 20 s) every 2 s
and drops the VIP 10 s after the last renewal etcd answered; a claim is
one transaction on the counter's key and the lease-bound holder's key,
so a stale claim is never written; the other node of the pair claims
once the lease expired, and a node that lost the VIP never claims it
again by itself. keel vip check, at boot and every minute, learns newer
claims from the peers.

keel vip status, keel mesh status and keel inspect show the role, epoch,
holder, lease, whether wg0 carries the VIP and where it is routed; apply
renders wg0.conf with the VIP routed to its holder and inspect reads the
peers back without it, so a move is never drift.

tests/test_vip_netns.py runs it end to end on three namespaces at 250 ms
±25 ms with 2% loss, measured first: a planned promote, the primary cut
off from the majority, the heal and a stale claim, the holders sampled
every 200 ms. CI job "vip / trixie".
navigator added 3 commits October 7, 2026 15:07
etcd's leader renews a lease by itself, so a leader cut off from the
majority answers renewals until it steps down, up to two election
timeouts: in a local run the cut-off primary, etcd's leader, dropped the
VIP 19 s after the cut, inside the lease but not within the release time.
A renewal now counts only once a linearizable read, which needs the
majority, finds the holder's key on this lease. The netns test reports
whether the cut primary was etcd's leader, asserts the drop within the
TTL, and its agent reports a failed command instead of dying; the stale
claim's read retries while the healed member catches up.
The review of keel#81 found, and this fixes:

- a persisted CLOCK_MONOTONIC time let a holder that rebooted, or met
  an etcd error, never reach its release deadline. Nothing about the
  last renewal is kept on disk now; the controller times it in memory on
  CLOCK_BOOTTIME (suspend counts), carries nothing after a start until
  it renewed its lease itself, and any error of etcd is no renewal.
  carry_held checks the renewal's age again before it adds the address;
- "pair only" was not enforced. keel vip pair records the pair, signed
  by both members (or a trust root); every claim carries the record,
  every claim and release is checked against it, a VIP that is a
  member's address or outside its region's /112 is refused, and an
  epoch more than 64 above the one a node knows is refused;
- a promote with no answering peer reported success. Before etcd it
  carries the VIP only when a majority of the peers that answered took
  the claim, and exits non-zero otherwise;
- keel mesh invite's allocator now reserves every VIP the node knows
  or declares;
- the controller was one root process. keel vip tend is now the root
  helper, the only one to change wg0, after checking the pair record;
  it starts keel vip control as a transient unit with a dynamic user and
  no capability, over a 0600 socket in a 0700 runtime directory checked
  by SO_PEERCRED against the unit's MainPID, the client key as a memfd;
- a holder key that is no valid claim neither holds the VIP nor blocks
  a failover: the claim's transaction compares its revision. A node
  that released the VIP is not fenced when its old lease is found gone.

etcd's RBAC is not enabled: every member issues its own client
certificates, so any member could name itself any etcd user. That needs
a decision on who issues them; docs/vip.md says so.
keel vip tend runs as a transient systemd unit joined to each node's
namespace with keel-vip.service's sandbox, and starts the controller as
its own unit; the controller's uid, capabilities and no-new-privs, and
the helper's bounding set, are read from the kernel. A and B are paired
with keel vip pair. The partition runs twice: with the cut-off primary
made etcd's leader first (what etcdctl move-leader does), the case the
majority-confirmed renewal fixed, and with it a follower. The CI job
boots systemd.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant