Skip to content

feat: etcd over the mesh, formed at the third cloud advanced member (0025, 0048) - #78

Merged
marcos-mendez merged 8 commits into
mainfrom
feat/mesh-etcd
Oct 4, 2026
Merged

marcos-mendez merged 8 commits into
mainfrom
feat/mesh-etcd

Conversation

@marcos-mendez

@marcos-mendez marcos-mendez commented Oct 4, 2026 •

Copy link
Copy Markdown
Collaborator

etcd, the mesh registry (0025), over the WireGuard mesh: formed at the third cloud advanced member, its own TLS, learners promoted, members removed under the amendment's rule, and a status section. Decisions: 0020, 0025, 0029, 0048 with its third round (handbook#46) and its admission evidence amendment, 0049 (VIP out of scope), 0050 (the design case).

Design

Who counts. Only a member whose spec says installation.mode: cloud_advanced and whose appliance carries the etcd overlay (0048, third round, point 3). The join request says whether the joining node can run etcd: it carries a certificate request only when it can.

CA and certificates (third round, point 1). All P-256, made with openssl on the machine, never in an image.

  • The root CA is made by the first node: keel mesh create on a cloud advanced node, or keel mesh etcd form on a mesh that has none (a mesh adopted under 0.18). Its key stays on that node, /var/lib/keel/etcd/root.key, 0600.
  • Every etcd member holds its own intermediate CA. Its key is made on the member; the CSR goes in the join request, and the inviter signs it with its own intermediate (or with the root on the first node). The join answer, under the invite's HMAC, carries the intermediate, the issuer chain and the root. A node that joined by the fallback, or a member of an adopted mesh, gets the same over the members' channel (below).
  • Each member issues its own leaves with its intermediate: the member certificate (serverAuth and clientAuth, SAN its overlay address and ::1, 1 year), and keel's client certificate (clientAuth, root's alone). Every certificate can be traced to the member whose intermediate signed it.
  • State, not spec (third round, point 2): /var/lib/keel/etcd/ root's, 0700, files 0600. keel spec apply copies the member's leaf with its chain, its key and the root to /etc/etcd/keel/, owned by etcd, 0600.
  • Rotation: keel-mesh-etcd.timer renews a leaf with a third of its life left; etcd reads its certificate files at each handshake, so no restart. Rotating an intermediate or the root is follow-up work, stated in docs/mesh.md.

Peer and client TLS, overlay only. Peers on https://[overlay]:2380, clients on https://[overlay]:2379 and https://[::1]:2379, never a wildcard address. client-cert-auth and peer-client-cert-auth both on, trusted CA the mesh's root, TLS 1.3 minimum. WireGuard is not what the registry relies on: a process on a member that holds no certificate signed in the mesh cannot write it (0048, second round, point 3).

Birth at the third member. The token's etcd state says forms when the inviter and exactly one other member are ready (credentials, cloud advanced). At admission the inviter issues the new member's intermediate and writes a three member initial cluster (overlay addresses, port 2380, names keel-<address>, the mesh identity as cluster token) into its state; the answer carries it. Once the join is confirmed, the new node and the inviter each set overlays.etcd: enabled in their own spec and apply. The inviter sends the cluster to the third member over the members' channel, signed with its Ed25519 key, and the third member does the same. Each starts with initial-cluster-state new. The start is --no-block: etcd reports ready only once there is quorum.

Fourth and fifth nodes. The token says running. Before it answers, the inviter runs member add --learner for the new node's peer URL. The answer says existing with the member list. The new node starts as a learner, which does not count toward quorum, so a join that never finishes does not lower fault tolerance. The inviter's helper promotes the learner once etcd accepts the promotion (in sync). keel-mesh-etcd.timer on every member does the same every 5 minutes, and removes a learner that never started after an hour.

An existing mesh: keel mesh etcd form. Safe on a mesh adopted under 0.18 with no CA:

  • it asks every peer first (probe: mode, credentials and their root, cluster), changes nothing, and refuses on a different identity, a different root, fewer than three ready members or a network change waiting;
  • --dry-run stops there and prints the plan;
  • then it makes the root if the mesh has none, enrolls every cloud advanced peer (its CSR over the members' channel, signed with its intermediate), and only once every enrollment succeeded sends each the cluster.

A run that stopped part way runs again: what is held is kept and the cluster is sent again. On a formed cluster it adds the ready members that are not in it as learners: the recovery path, and the fallback join's.

The members' channel gains POST /v1/etcd (probe, enroll, cluster), with its existing limits. WireGuard says who sends. The message also carries an Ed25519 signature by the sender's signing key as the receiver's trust store holds it (the amendment's rule for messages that act). It is refused when stale (5 minutes), for another mesh, or on a node that cannot run etcd.

Leaving (third round, point 4). After keel mesh remove has dropped the peer under its window, it runs etcd's member remove for the node's member only when this node may remove it mesh-wide: this node admitted it (the evidence it keeps), or the node is a trust root of this node (an adopted mesh, where the operator made the members each other's roots). Otherwise it says the etcd member stays and the removal is local only. With two voters left it warns that the cluster has no fault tolerance.

Timeouts. Heartbeat 300 ms, election timeout 5000 ms. etcd's tuning guide sets the heartbeat interval "around the round-trip time between members", and the election timeout "at least 10 times the round-trip time" to absorb its variance. The design case is 250 ms ±25 ms with 2% loss.

  • 300 ms covers the RTT plus its jitter.
  • 3000 ms (10× the heartbeat) is what 0050's bench used, and it re-elected under loss. Peer traffic is a TCP stream, so a lost segment stalls it for one retransmission timeout (200 ms minimum plus the RTT, about 450 to 500 ms), and back-to-back losses double it.
  • 5000 ms is about 17 heartbeats and 20× the mean RTT, and absorbs three successive retransmission timeouts. Detecting a dead leader takes 5 to 10 s (the timeout is randomized in [T, 2T)), small beside DNS's 1 to 5 minutes (0020).
  • Pre-vote stays on (etcd 3.5's default), so a member coming back from a partition does not depose the leader.

Two members. Never formed by a join (0048, Q3). Reached only by removal or failure: quorum is 2 of 2 and keel mesh status says the cluster has no fault tolerance.

Partition. A member cut off from a three member cluster cannot win a pre-vote, so it does not raise the term. It serves no linearizable request, and catches up from the leader's log on heal. A leader cut off steps down after an election timeout (check-quorum), and the majority elects.

After the third review (fifth and sixth commits)

This narrows the approved third round, Q1. The decision was "an intermediate CA per member, signed by its inviter". It is now "signed by the root CA, the request relayed by the inviter". Why:

  • an intermediate that may sign intermediates cannot be limited to its own member;
  • nor can it be recorded where a removal finds it;
  • chains through members could not be revoked as a unit.
    The root's key still stays on the first node. If the holder is unreachable, the join completes without etcd, and the request is queued: keel-mesh-etcd.timer retries it, and keel mesh status says the node waits. Grants carry no chain (one intermediate deep, path length 0).

Fixes, item by item:

  1. Revocations. A revocation names no serial. The holder revokes only what it recorded for that address and WireGuard key, and only when asked by the node's admitter, a trust root, or the node itself. It signs an intermediate for an address only with that address's own key. Tested: a member cannot revoke another's certificates.
  2. Records. Records carry etcd's cluster ID, an epoch that only grows, a nonce and a 1-hour expiry. They name exactly the cluster's members for new and existing alike. Members refuse a record no newer than the last they took, and refuse expired ones. Only the holder adds learners, so every record is the root's.
  3. Stray member. A stray member directory is removed only if keel-overlay-etcd marked it as etcd-server's own (common#39 now marks it rather than deleting it). Any other is refused, by name.
  4. Root-only signing. Every intermediate is signed by the root, as above, and recorded.
  5. Name constraints. Each intermediate is constrained to its member's own /128 and ::1. Tested: a member cannot certify another member's address.
  6. Renewal failures. They are alerted through the monitor's channels (decision 0021) and kept. keel mesh status shows them, and keel diff reports etcd.certificates as drift, as it does for a certificate within 7 days of expiry.
  7. Test validity. Before the scenario the netns test pings the overlay 400 times. It asserts a round trip between 225 and 290 ms and round-trip loss between 0.5% and 8%, and fails otherwise. Locally: avg 250.4 ms, 2.25% loss.

After the security review (third and fourth commits)

  • Only the root CA's holder forms, once. It picks the members, signs a formation record (token, members, root fingerprint) with the root, and reserves the formation under its lock. A member accepts a cluster only with a valid record signed by the root, and a member already in a cluster refuses any other. If another node invites the third member, it admits that node to the mesh and names the holder; keel mesh etcd form runs only on the holder. On a mesh with no CA, form makes the root where it runs, so run it on a trust root. Tested with two concurrent invites: one formation, never a second cluster.
  • Short lives and a CRL. Leaves last 30 days and are renewed by their member. Intermediates last 1 year and are renewed by the holder, which signs them with the root.
    • Chain depth (point 6): an inviter relays a joiner's request to the holder. Only if the holder is unreachable does it sign one level deeper, and the first renewal re-anchors the chain under the root. MAX_CHAIN stays 8.
    • Name constraints: every intermediate permits only the mesh's /64 and ::1. pathLen can't be 0, because of that fallback.
    • The CRL is signed by the root and written to --peer-crl-file and --client-crl-file. It travels in grants and rosters, and a removal has the holder revoke the node's intermediates.
    • The netns test shows a revoked member's client certificate refused by etcd.
  • Monit: /health on a plain metrics listener on http://[::1]:2381 (fix: etcd's health on its metrics listener, and no lone member kept common#39, merge it first).
  • Stray member: common#39's postinst removes the member that etcd-server's own start made. keel refuses to start over a member it never started, and the join or form that starts etcd removes it, logging it, only while the node was never a member.
  • Leader partition in the CI etcd test: the leader is cut off for 120 s at 250 ms ±25 ms with 2% loss. The test checks a new leader is elected, writes resume, and the old leader rejoins as a follower.

Seams (each written before its tests)

  • keel.mesh.etcdpki: the root, intermediates, leaves, CSR checks, chain verification, expiry, with the real openssl (tests/test_mesh_etcdpki.py).
  • keel.mesh.etcdstate: the state directory, credentials, cluster and ready members, member names (tests/test_mesh_etcdstate.py).
  • keel.mesh.etcdconf: /etc/default/etcd and the systemd drop-in rendered from state, pure (tests/test_mesh_etcdconf.py).
  • keel.mesh.etcdclient: etcd's v3 JSON gateway (the API etcdctl calls) over TLS with keel's client certificate; tested against a fake gateway over real TLS on the loopback, and against real etcd in the netns test (tests/test_mesh_etcdclient.py).
  • keel.mesh.etcdmsg: the join request's CSR, the answer's grant and cluster, the members' channel messages and their signatures, pure (tests/test_mesh_etcdmsg.py).
  • keel.mesh.etcd: the flows (token state, issue at admission, take the grant, enable, form, bring a member in, tend, remove, status), each against scratch roots with a recording client and channel (tests/test_mesh_etcd.py, tests/test_mesh_etcd_form.py).
  • keel.system.etcd with the appliance step: files, owner, drop-in, daemon-reload, start --no-block, nothing started before the cluster exists (tests/test_system_etcd.py); the spec rule (tests/test_spec_appliance.py), the diff (tests/test_diff_appliance.py).
  • The changed join, accept, serve, remove, status and members' channel paths in their existing test files.
  • End to end, tests/test_etcd_netns.py runs tests/etcd_netns.py as root in a network namespace that routes for three child namespaces: real WireGuard, real etcd 3.5, keel's CA, intermediates (one issued by a non-root member), leaves and rendered configuration, under netem.

Tests and evidence

  • Unit and flow tests: 2951 passed locally; keel/mesh at 100% line and branch coverage, except two lines of commands.py that only the live-system netns job reaches (as before this PR). keel/system/etcd.py 100%.
  • etcd / trixie (new CI job): the netns test with 60 s of stability and a 45 s partition.
  • The brief's run, local, ETCD_SOAK=full: three namespaces, real WireGuard (wg-quick) and trixie's etcd 3.5.16, keel's CA with C's intermediate signed by B's, and keel's rendered configuration. Each leg was netem delay 125ms 12.5ms loss 2%, and the measured overlay RTT A→B was min/avg/max 229/254/310 ms with 5% ping loss.
Phase Result
Forms quorum and one agreed leader 10.1 s after the three etcds started
Stable, 300 s every member saw one leader and term 2 in every sample (158, 160, 215 samples), no flap; 267 writes, 0 failed
Partition of B, a follower, 120 s, 100% loss both ways A and C kept the same leader and term 2: 64 writes, 0 failed. B: no leader, 21 writes refused
Heal B reported the leader again after 47.2 s, still term 2 (pre-vote: no disruption), same leader
Learner added with keel's client, listed unstarted, removed

4 passed in 500.06s. The raw result line:

RESULT {"netem_leg": ["125ms", "12.5ms", "2%"], "stable_s": 300, "partition_s": 120, "heartbeat_ms": 300, "election_ms": 5000, "formed_s": 10.1, "leader_at_formation": "13236511977862390820", "ping_a_b_overlay": ["20 packets transmitted, 19 received, 5% packet loss, time 9532ms", "rtt min/avg/max/mdev = 228.775/254.356/309.877/17.097 ms"], "stable": {"A": {"samples": 158, "leaders": ["13236511977862390820"], "terms": [2], "puts": 79, "put_failures": 0, "unhealthy": 0}, "B": {"samples": 160, "leaders": ["13236511977862390820"], "terms": [2], "puts": 80, "put_failures": 0, "unhealthy": 0}, "C": {"samples": 215, "leaders": ["13236511977862390820"], "terms": [2], "puts": 108, "put_failures": 0, "unhealthy": 0}}, "partitioned": "B", "partitioned_was_leader": false, "during_partition": {"A": {"samples": 55, "leaders": ["13236511977862390820"], "terms": [2], "puts": 28, "put_failures": 0, "unhealthy": 0}, "B": {"samples": 41, "leaders": [], "terms": [2], "puts": 0, "put_failures": 21, "unhealthy": 41}, "C": {"samples": 72, "leaders": ["13236511977862390820"], "terms": [2], "puts": 36, "put_failures": 0, "unhealthy": 0}}, "rejoined_s": 47.2, "after_heal": {"A": {"samples": 28, "leaders": ["13236511977862390820"], "terms": [2], "puts": 14, "put_failures": 0, "unhealthy": 0}, "B": {"samples": 17, "leaders": ["13236511977862390820"], "terms": [2], "puts": 3, "put_failures": 5, "unhealthy": 11}, "C": {"samples": 38, "leaders": ["13236511977862390820"], "terms": [2], "puts": 19, "put_failures": 0, "unhealthy": 0}}, "leader_after_heal": "13236511977862390820", "learner": {"added": true, "listed_unstarted": true, "removed": true}}
  • Sandbox: etcd 3.5.16 under every property of the drop-in (as a transient unit, as nobody) reached "ready to serve client requests". Its only complaint was the info-level "failed to detect default host" (netlink is not among the address families allowed), which has no effect with explicit URLs.

Review before opening, fixed in the second commit

A cluster message no longer replaces a cluster already held (tokens are always the mesh identity). An etcd CA is taken only from the inviter or a trust root. Ready lists from other members are kept only for this node's own peers. A join can't fail on etcd fields. fingerprint raises on a non-certificate. The receiver refuses while a network change waits.

Before the first real run on web-1..3

keel mesh etcd form --dry-run on one node first. It needs every node to be installation.mode: cloud_advanced with the etcd overlay in its spec, and each node's trust store to hold the others' signing keys (an adopted mesh gets them at keel mesh sync). A refusal says which is missing.

navigator added 3 commits October 4, 2026 04:00
…0025, 0048)

etcd, the mesh's registry, on the members whose installation mode is
cloud_advanced, over the WireGuard mesh, as handbook decision 0048 and
its third round decide:

- its own TLS, peer and client, on the overlay addresses and ::1 only:
  the first node makes the mesh's root CA, every member holds an
  intermediate CA signed by its inviter and carried in the join's
  authenticated answer, and issues its own member and client
  certificates (keel.mesh.etcdpki, keel.mesh.etcdstate);
- the member list and the certificates are state under
  /var/lib/keel/etcd; the spec says only overlays.etcd: enabled, which
  spec validate refuses outside cloud advanced;
- the cluster forms at the third member's join; from the fourth the
  inviter adds the new node as a learner before it answers and promotes
  it once in sync; keel mesh etcd tend, from keel-mesh-etcd.timer,
  promotes, removes learners that never started, renews the leaves;
- keel mesh etcd form forms it on a mesh that never saw a third join,
  such as one adopted under 0.18 with no CA: it asks every member first,
  refuses two roots or a cluster it is not in, enrolls everyone before
  it starts anything, and has --dry-run;
- the members' channel gains POST /v1/etcd, signed with the sender's
  Ed25519 key and checked against the receiver's trust store;
- keel mesh remove removes the etcd member only when this node admitted
  the node or it is a trust root;
- keel mesh status shows the members, the leader and their health;
- apply renders /etc/default/etcd, the certificates for the etcd user
  and a sandboxing drop-in from the state, and starts etcd --no-block;
  before the cluster exists etcd waits without failing the run.

Heartbeat 300 ms and election timeout 5000 ms for 250 ms ±25 ms with
2% loss (decision 0050), measured end to end in tests/test_etcd_netns.py
with three namespaces, the real WireGuard and etcd 3.5, run in CI as
"etcd / trixie".
…ember

From a review of the etcd branch:

- a cluster message is refused when this node holds another cluster
  (its token is always the mesh's identity, so comparing tokens let any
  member replace it), and while a network change waits;
- an etcd CA is taken only from this node's inviter or a trust root,
  and its root is written first;
- the members another member says are ready are kept only when they are
  this node's peers at those addresses, at a join too;
- a join never fails on etcd fields it cannot read, and a join forms no
  cluster of more than seven;
- the fingerprint of what is no certificate is an error, and the
  receiver answers 503 rather than raising when openssl or a write fails;
- the netns test waits out etcd's strict reconfiguration check before
  it adds its learner, right after a member rejoined.
navigator added 5 commits October 4, 2026 05:21
…a CRL

From the security review of keel#78:

- only the node that holds the mesh's root CA forms the cluster, once
  (keel.mesh.etcdca): it signs a formation record (cluster token,
  members, root fingerprint) with the root and reserves the formation
  under its lock; a member takes a cluster, from a join's answer or a
  cluster message, only with that record, and a member in a cluster
  takes no other. Two concurrent invites can no longer start two
  clusters; another inviter of the third member admits it and names the
  holder; keel mesh etcd form runs only on the holder, and on a mesh
  with no CA makes the root where the operator runs it (a trust root);
- leaves last 30 days and are renewed by their member; intermediates a
  year, renewed by the holder, which signs them with the root (an
  inviter relays the request; only when the holder is unreachable does
  it sign one level deeper, re-anchored at the first renewal);
- every intermediate is name constrained to the mesh's /64 and ::1;
- the holder keeps a CRL signed by the root, written to etcd's
  --peer-crl-file and --client-crl-file and carried by grants and
  rosters; keel mesh remove has the holder revoke the removed node's
  intermediates, and so its leaves and what it issued;
- etcd serves /health on a plain metrics listener on [::1]:2381 for
  Monit (Keel-Linux/common#39);
- apply refuses a member in /var/lib/etcd/default that keel never
  started, and the join or form that starts etcd removes it first.
From the third review of keel#78:

- a revocation names no serial: the holder revokes the intermediates
  it recorded for that address and WireGuard key, and only for the
  node's admitter, a trust root or the node itself; it signs an
  intermediate for an address only with that address's own key;
- records carry etcd's cluster ID, an epoch that only grows, a nonce
  and an expiry, name exactly the cluster's members for `new` and
  `existing` alike, and a member takes none no newer than the last;
  only the holder adds a learner, so every record is the root's;
- only the root signs intermediates, relayed by the inviter (0048's
  third round, point 1, narrowed): one intermediate deep, path length
  0, name constrained to its member's own /128 and ::1, each recorded.
  When the holder cannot be reached the join completes without etcd,
  the request is queued, keel-mesh-etcd.timer asks again, and keel
  mesh status says so;
- a member directory keel never started is removed only when
  keel-overlay-etcd marked it as etcd-server's own (Keel-Linux/common#39);
  any other is refused, named;
- a renewal that fails is alerted through the monitor's channels and
  shown by keel mesh status and keel diff (etcd.certificates), as is a
  certificate within seven days of its expiry.
@marcos-mendez
marcos-mendez merged commit 4136e02 into main Oct 4, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant