Repository navigation
feat: etcd over the mesh, formed at the third cloud advanced member (0025, 0048) - #78
Merged
Merged
Conversation
added 3 commits
October 4, 2026 04:00
…0025, 0048) etcd, the mesh's registry, on the members whose installation mode is cloud_advanced, over the WireGuard mesh, as handbook decision 0048 and its third round decide: - its own TLS, peer and client, on the overlay addresses and ::1 only: the first node makes the mesh's root CA, every member holds an intermediate CA signed by its inviter and carried in the join's authenticated answer, and issues its own member and client certificates (keel.mesh.etcdpki, keel.mesh.etcdstate); - the member list and the certificates are state under /var/lib/keel/etcd; the spec says only overlays.etcd: enabled, which spec validate refuses outside cloud advanced; - the cluster forms at the third member's join; from the fourth the inviter adds the new node as a learner before it answers and promotes it once in sync; keel mesh etcd tend, from keel-mesh-etcd.timer, promotes, removes learners that never started, renews the leaves; - keel mesh etcd form forms it on a mesh that never saw a third join, such as one adopted under 0.18 with no CA: it asks every member first, refuses two roots or a cluster it is not in, enrolls everyone before it starts anything, and has --dry-run; - the members' channel gains POST /v1/etcd, signed with the sender's Ed25519 key and checked against the receiver's trust store; - keel mesh remove removes the etcd member only when this node admitted the node or it is a trust root; - keel mesh status shows the members, the leader and their health; - apply renders /etc/default/etcd, the certificates for the etcd user and a sandboxing drop-in from the state, and starts etcd --no-block; before the cluster exists etcd waits without failing the run. Heartbeat 300 ms and election timeout 5000 ms for 250 ms ±25 ms with 2% loss (decision 0050), measured end to end in tests/test_etcd_netns.py with three namespaces, the real WireGuard and etcd 3.5, run in CI as "etcd / trixie".
…ember From a review of the etcd branch: - a cluster message is refused when this node holds another cluster (its token is always the mesh's identity, so comparing tokens let any member replace it), and while a network change waits; - an etcd CA is taken only from this node's inviter or a trust root, and its root is written first; - the members another member says are ready are kept only when they are this node's peers at those addresses, at a join too; - a join never fails on etcd fields it cannot read, and a join forms no cluster of more than seven; - the fingerprint of what is no certificate is an error, and the receiver answers 503 rather than raising when openssl or a write fails; - the netns test waits out etcd's strict reconfiguration check before it adds its learner, right after a member rejoined.
added 5 commits
October 4, 2026 05:21
…a CRL From the security review of keel#78: - only the node that holds the mesh's root CA forms the cluster, once (keel.mesh.etcdca): it signs a formation record (cluster token, members, root fingerprint) with the root and reserves the formation under its lock; a member takes a cluster, from a join's answer or a cluster message, only with that record, and a member in a cluster takes no other. Two concurrent invites can no longer start two clusters; another inviter of the third member admits it and names the holder; keel mesh etcd form runs only on the holder, and on a mesh with no CA makes the root where the operator runs it (a trust root); - leaves last 30 days and are renewed by their member; intermediates a year, renewed by the holder, which signs them with the root (an inviter relays the request; only when the holder is unreachable does it sign one level deeper, re-anchored at the first renewal); - every intermediate is name constrained to the mesh's /64 and ::1; - the holder keeps a CRL signed by the root, written to etcd's --peer-crl-file and --client-crl-file and carried by grants and rosters; keel mesh remove has the holder revoke the removed node's intermediates, and so its leaves and what it issued; - etcd serves /health on a plain metrics listener on [::1]:2381 for Monit (Keel-Linux/common#39); - apply refuses a member in /var/lib/etcd/default that keel never started, and the join or form that starts etcd removes it first.
From the third review of keel#78: - a revocation names no serial: the holder revokes the intermediates it recorded for that address and WireGuard key, and only for the node's admitter, a trust root or the node itself; it signs an intermediate for an address only with that address's own key; - records carry etcd's cluster ID, an epoch that only grows, a nonce and an expiry, name exactly the cluster's members for `new` and `existing` alike, and a member takes none no newer than the last; only the holder adds a learner, so every record is the root's; - only the root signs intermediates, relayed by the inviter (0048's third round, point 1, narrowed): one intermediate deep, path length 0, name constrained to its member's own /128 and ::1, each recorded. When the holder cannot be reached the join completes without etcd, the request is queued, keel-mesh-etcd.timer asks again, and keel mesh status says so; - a member directory keel never started is removed only when keel-overlay-etcd marked it as etcd-server's own (Keel-Linux/common#39); any other is refused, named; - a renewal that fails is alerted through the monitor's channels and shown by keel mesh status and keel diff (etcd.certificates), as is a certificate within seven days of its expiry.
…scenario, and revokes by record
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
etcd, the mesh registry (0025), over the WireGuard mesh: formed at the third cloud advanced member, its own TLS, learners promoted, members removed under the amendment's rule, and a status section. Decisions: 0020, 0025, 0029, 0048 with its third round (handbook#46) and its admission evidence amendment, 0049 (VIP out of scope), 0050 (the design case).
Design
Who counts. Only a member whose spec says
installation.mode: cloud_advancedand whose appliance carries theetcdoverlay (0048, third round, point 3). The join request says whether the joining node can run etcd: it carries a certificate request only when it can.CA and certificates (third round, point 1). All P-256, made with openssl on the machine, never in an image.
keel mesh createon a cloud advanced node, orkeel mesh etcd formon a mesh that has none (a mesh adopted under 0.18). Its key stays on that node,/var/lib/keel/etcd/root.key, 0600./var/lib/keel/etcd/root's, 0700, files 0600.keel spec applycopies the member's leaf with its chain, its key and the root to/etc/etcd/keel/, owned by etcd, 0600.keel-mesh-etcd.timerrenews a leaf with a third of its life left; etcd reads its certificate files at each handshake, so no restart. Rotating an intermediate or the root is follow-up work, stated in docs/mesh.md.Peer and client TLS, overlay only. Peers on
https://[overlay]:2380, clients onhttps://[overlay]:2379andhttps://[::1]:2379, never a wildcard address.client-cert-authandpeer-client-cert-authboth on, trusted CA the mesh's root, TLS 1.3 minimum. WireGuard is not what the registry relies on: a process on a member that holds no certificate signed in the mesh cannot write it (0048, second round, point 3).Birth at the third member. The token's etcd state says
formswhen the inviter and exactly one other member are ready (credentials, cloud advanced). At admission the inviter issues the new member's intermediate and writes a three member initial cluster (overlay addresses, port 2380, nameskeel-<address>, the mesh identity as cluster token) into its state; the answer carries it. Once the join is confirmed, the new node and the inviter each setoverlays.etcd: enabledin their own spec and apply. The inviter sends the cluster to the third member over the members' channel, signed with its Ed25519 key, and the third member does the same. Each starts withinitial-cluster-state new. The start is--no-block: etcd reports ready only once there is quorum.Fourth and fifth nodes. The token says
running. Before it answers, the inviter runsmember add --learnerfor the new node's peer URL. The answer saysexistingwith the member list. The new node starts as a learner, which does not count toward quorum, so a join that never finishes does not lower fault tolerance. The inviter's helper promotes the learner once etcd accepts the promotion (in sync).keel-mesh-etcd.timeron every member does the same every 5 minutes, and removes a learner that never started after an hour.An existing mesh:
keel mesh etcd form. Safe on a mesh adopted under 0.18 with no CA:probe: mode, credentials and their root, cluster), changes nothing, and refuses on a different identity, a different root, fewer than three ready members or a network change waiting;--dry-runstops there and prints the plan;A run that stopped part way runs again: what is held is kept and the cluster is sent again. On a formed cluster it adds the ready members that are not in it as learners: the recovery path, and the fallback join's.
The members' channel gains
POST /v1/etcd(probe,enroll,cluster), with its existing limits. WireGuard says who sends. The message also carries an Ed25519 signature by the sender's signing key as the receiver's trust store holds it (the amendment's rule for messages that act). It is refused when stale (5 minutes), for another mesh, or on a node that cannot run etcd.Leaving (third round, point 4). After
keel mesh removehas dropped the peer under its window, it runs etcd'smember removefor the node's member only when this node may remove it mesh-wide: this node admitted it (the evidence it keeps), or the node is a trust root of this node (an adopted mesh, where the operator made the members each other's roots). Otherwise it says the etcd member stays and the removal is local only. With two voters left it warns that the cluster has no fault tolerance.Timeouts. Heartbeat 300 ms, election timeout 5000 ms. etcd's tuning guide sets the heartbeat interval "around the round-trip time between members", and the election timeout "at least 10 times the round-trip time" to absorb its variance. The design case is 250 ms ±25 ms with 2% loss.
Two members. Never formed by a join (0048, Q3). Reached only by removal or failure: quorum is 2 of 2 and
keel mesh statussays the cluster has no fault tolerance.Partition. A member cut off from a three member cluster cannot win a pre-vote, so it does not raise the term. It serves no linearizable request, and catches up from the leader's log on heal. A leader cut off steps down after an election timeout (check-quorum), and the majority elects.
After the third review (fifth and sixth commits)
This narrows the approved third round, Q1. The decision was "an intermediate CA per member, signed by its inviter". It is now "signed by the root CA, the request relayed by the inviter". Why:
The root's key still stays on the first node. If the holder is unreachable, the join completes without etcd, and the request is queued: keel-mesh-etcd.timer retries it, and
keel mesh statussays the node waits. Grants carry no chain (one intermediate deep, path length 0).Fixes, item by item:
newandexistingalike. Members refuse a record no newer than the last they took, and refuse expired ones. Only the holder adds learners, so every record is the root's.keel mesh statusshows them, andkeel diffreportsetcd.certificatesas drift, as it does for a certificate within 7 days of expiry.After the security review (third and fourth commits)
keel mesh etcd formruns only on the holder. On a mesh with no CA,formmakes the root where it runs, so run it on a trust root. Tested with two concurrent invites: one formation, never a second cluster.--peer-crl-fileand--client-crl-file. It travels in grants and rosters, and a removal has the holder revoke the node's intermediates./healthon a plain metrics listener onhttp://[::1]:2381(fix: etcd's health on its metrics listener, and no lone member kept common#39, merge it first).Seams (each written before its tests)
keel.mesh.etcdpki: the root, intermediates, leaves, CSR checks, chain verification, expiry, with the real openssl (tests/test_mesh_etcdpki.py).keel.mesh.etcdstate: the state directory, credentials, cluster and ready members, member names (tests/test_mesh_etcdstate.py).keel.mesh.etcdconf:/etc/default/etcdand the systemd drop-in rendered from state, pure (tests/test_mesh_etcdconf.py).keel.mesh.etcdclient: etcd's v3 JSON gateway (the API etcdctl calls) over TLS with keel's client certificate; tested against a fake gateway over real TLS on the loopback, and against real etcd in the netns test (tests/test_mesh_etcdclient.py).keel.mesh.etcdmsg: the join request's CSR, the answer's grant and cluster, the members' channel messages and their signatures, pure (tests/test_mesh_etcdmsg.py).keel.mesh.etcd: the flows (token state, issue at admission, take the grant, enable, form, bring a member in, tend, remove, status), each against scratch roots with a recording client and channel (tests/test_mesh_etcd.py,tests/test_mesh_etcd_form.py).keel.system.etcdwith the appliance step: files, owner, drop-in, daemon-reload, start--no-block, nothing started before the cluster exists (tests/test_system_etcd.py); the spec rule (tests/test_spec_appliance.py), the diff (tests/test_diff_appliance.py).tests/test_etcd_netns.pyrunstests/etcd_netns.pyas root in a network namespace that routes for three child namespaces: real WireGuard, real etcd 3.5, keel's CA, intermediates (one issued by a non-root member), leaves and rendered configuration, under netem.Tests and evidence
commands.pythat only the live-system netns job reaches (as before this PR). keel/system/etcd.py 100%.etcd / trixie(new CI job): the netns test with 60 s of stability and a 45 s partition.ETCD_SOAK=full: three namespaces, real WireGuard (wg-quick) and trixie's etcd 3.5.16, keel's CA with C's intermediate signed by B's, and keel's rendered configuration. Each leg wasnetem delay 125ms 12.5ms loss 2%, and the measured overlay RTT A→B was min/avg/max 229/254/310 ms with 5% ping loss.4 passed in 500.06s. The raw result line:nobody) reached "ready to serve client requests". Its only complaint was the info-level "failed to detect default host" (netlink is not among the address families allowed), which has no effect with explicit URLs.Review before opening, fixed in the second commit
A cluster message no longer replaces a cluster already held (tokens are always the mesh identity). An etcd CA is taken only from the inviter or a trust root. Ready lists from other members are kept only for this node's own peers. A join can't fail on etcd fields.
fingerprintraises on a non-certificate. The receiver refuses while a network change waits.Before the first real run on web-1..3
keel mesh etcd form --dry-runon one node first. It needs every node to beinstallation.mode: cloud_advancedwith the etcd overlay in its spec, and each node's trust store to hold the others' signing keys (an adopted mesh gets them atkeel mesh sync). A refusal says which is missing.