diff --git a/ARCHITECTURE.md b/ARCHITECTURE.md index 43911afd..e1076342 100644 --- a/ARCHITECTURE.md +++ b/ARCHITECTURE.md @@ -539,13 +539,13 @@ A group id must outlive the replica object that minted it — a receiver buckets **A filter that withholds a member destrands the survivors.** Redaction is per op: a doc-ACL read verdict, a zone scope, or a migration rewrite drops individual members out of a batch. The rest then carry a count their bucket can never reach, so a recipient holds them against a member that will never arrive — invisible to it forever, and still counted among the ids it holds. Every seam that withholds a member therefore delivers the group's survivors untagged, so they merge standalone. Delivering them is the convergence requirement: every op a recipient may receive has to reach it, or it diverges from the correct projection of the sender's state. The atomic view is lost at such a recipient, unavoidably — it cannot see the member that was withheld — and the ops still merge. A group a filter carries whole keeps its tags and stays atomic. One seam cannot follow the rule: a read projection of a snapshot has no per-op verdict to apply to buffered ops, whose paths may not resolve at all, so it drops the buffer entire rather than destranding it — the survivors go with the withheld member. -**A rewritten `count` is a memory-retention instruction, and the answer to it is eviction, not a repair rule.** `count` tells the receiver to hold the group's members until that many arrive, so a size no arrival meets tells it to hold them for the life of the replica — and the buffer rides the state encoding, so the next replica holds them too. Two judgements over a member's own envelope are safe to make locally. A size outside the cap is **refused** where it arrives — at the decode boundary, and at the apply seam an in-process caller reaches without crossing one — because no honest sender mints one, the judgement is on that member alone, and refusing holds nothing. And a group's size is what its members *agree* it is: read off whichever member the buffer happens to hold first, a rewritten count decides when the group commits, so a bucket whose members disagree names no group and is never complete. Unanimity is judged over the members that have *arrived*, so it bounds a rewrite rather than closing it — a unanimous **subset** can reach its own declared count before the dissenting member lands, which is a further shape of the same defect the record below closes for its own three. +**A rewritten `count` is a memory-retention instruction, and the answer to it is eviction, not a repair rule.** `count` tells the receiver to hold the group's members until that many arrive, so a size no arrival meets tells it to hold them for the life of the replica — and the buffer rides the state encoding, so the next replica holds them too. Two judgements over a member's own envelope are safe to make locally. A size outside the cap is **refused** where it arrives — at the decode boundary, and at the apply seam an in-process caller reaches without crossing one — because no honest sender mints one, the judgement is on that member alone, and refusing holds nothing. And a group's size is what its members *agree* it is: read off whichever member the buffer happens to hold first, a rewritten count decides when the group commits, so a bucket whose members disagree names no group and can never complete. Holding such a bucket is not the answer to that, because unanimity is judged over the members that have *arrived*: a unanimous **subset** reaches its own declared count and commits before the dissenting member lands, so whether a bucket ever holds the disagreement at all belongs to the arrival order, and one op set folds to two states. A disagreement therefore **spends the bucket key** exactly where a commit spends it (the record below) — the members it holds are released untagged, and every member still to come is a stray of a resolved group. What that costs is the atomic view of a group a rewrite has already made unservable, which is the price the record pays for the same reason. Neither of those reaches a rewrite consistent across every member: a group of three retagged to declare two is, to a receiver, an honest group of two followed by a stray. It commits at the size it was told, and *which* members that is belongs to the arrival order, so the third is left holding a size its bucket has already met and no arrival can meet again. Two more shapes land in the same place: an unrelated op of the same author carrying a live group's id, and two copies of one op id carrying envelopes that disagree — in group id or in declared size — where the bucket reads whichever the dedup kept. What is common to them is that a bucket key is *consumed* when it resolves, so a late member of a resolved group is indistinguishable from the first member of a fresh one — and the replica folds one op set to two states. -So the judgement that closes them is over the group rather than the member: **record the key**. A `(author, group id)` set marks a bucket resolved, and a member arriving under a resolved key is untagged and merges standalone. A key is spent at each of the four points a bucket resolves: when it **commits**; when the author **mints** the group, since the author applies its own edits as it makes them and buckets nothing, so without this it would hold a stray every receiver merged; when **eviction** gives up on it, or a member arriving after one would wait on a group the replica has already released, and two replicas on one policy would disagree over nothing but which had ticked first; and when a member arrives naming a group other than the one the buffer is holding that same id under, which spends **both** — only one of the two can ever hold the id, and which one is the arrival order's. Each of the three shapes then lands the same op set from every arrival order, which is what the law asks; what it costs is the atomic *view* of a group a rewrite has already made unservable, and only for the members that follow the commit. The record is persisted, carried in the state encoding beside the buffer it rules — a group resolved before a restart is one whose stray still has to land after it — and every entry is charged to a bucket the replica held, committed, evicted or minted, so it is bounded by the ops it holds rather than by what arrives. +So the judgement that closes them is over the group rather than the member: **record the key**. A `(author, group id)` set marks a bucket resolved, and a member arriving under a resolved key is untagged and merges standalone. A key is spent at each of the five points a bucket resolves: when it **commits**; when its members **disagree**, since a bucket without unanimity completes on no delivery (above); when the author **mints** the group, since the author applies its own edits as it makes them and buckets nothing, so without this it would hold a stray every receiver merged; when **eviction** gives up on it, or a member arriving after one would wait on a group the replica has already released, and two replicas on one policy would disagree over nothing but which had ticked first; and when a member arrives naming a group other than the one the buffer is holding that same id under, which spends **both** — only one of the two can ever hold the id, and which one is the arrival order's. Each of the three shapes then lands the same op set from every arrival order, which is what the law asks; what it costs is the atomic *view* of a group a rewrite has already made unservable — for the members that follow a commit, and for every member of a bucket that disagrees. The record is persisted, carried in the state encoding beside the buffer it rules — a group resolved before a restart is one whose stray still has to land after it — and every entry is charged to a bucket the replica held, committed, evicted or minted. Its bound is the **dedup set**, not the buffer: a key outlives the ops that earned it exactly as a `seen` entry does, so an eviction empties the buffer and keeps the key, and what caps the record is that each entry costs the sender fresh op ids it can never re-spend. -Two rules the record deliberately does not take. It does not read a member whose id is merely **applied**: a resend is ordinary traffic on every transport that retries, so a delivery that spent a key would make state a function of how often an op arrived rather than of which ops did — the same law, broken in the other dimension. And it does not release a bucket the moment it *looks* unreachable, because whether it looks that way is a function of which members have landed, so replicas served the same ops in different orders would release different sets. What those two leave is three further shapes. Two are order-dependent before the record existed too; the third the record itself opens, because the record is per-replica *evidence* and a destranding seam destroys the evidence at exactly the recipients it serves. First, a second envelope of one id that the buffer holds **nothing tagged** to contradict — because the other copy already committed out of the buffer under a different group, or because it carries **no** tag at all, which is what a filtering seam's destranding produces. The honest group's own member is then left holding on an id that will never join it. Its members converge on eviction; which keys each replica has spent does not, and on some arrival orders the record adds a *third* reading where the un-recorded replica had two — the conflict rule fires on the orders that buffer the disagreeing copy and not on the rest. Measured against the un-recorded replica over 392 forged pools it is better on 156 and worse on 4, so the record improves this shape without closing it. And a **minority** count rewrite, where the members left unanimous are exactly as many as they now declare: that subset commits and the dissenter lands as a stray, or the dissenter arrives first and the bucket names no group at all. Closing the second means a disagreeing bucket spends its key rather than merely never completing, which is a change to what unanimity *is* and wants its own decision. And third: a recipient served a group **destranded** never buckets it, so it never spends the key, while the author spends it at the mint and a whole-delivery recipient spends it on commit — a later stray under that id then merges at those two and is held at the destranded one, where before the record all three held it alike. Eviction closes it; the destranding seams knowing the keys they cut is the other way, and it has to answer what a projection may reveal about a group that straddles its cut. A projection drops the record whole: a key names an author and a group, never a partition, so a kept one would count the groups a withheld partition resolved — the same inference the causal-frontier scrub closes. +Two rules the record deliberately does not take. It does not read a member whose id is merely **applied**: a resend is ordinary traffic on every transport that retries, so a delivery that spent a key would make state a function of how often an op arrived rather than of which ops did — the same law, broken in the other dimension. And it does not release a bucket that merely *looks* unreachable on the count it is short of, because whether the members it lacks will ever come is a function of what has not arrived, so replicas served the same ops in different orders would release different sets. A **disagreement** is not that: two members the buffer already holds contradicting each other is a property of what *has* landed and no arrival repairs it, so spending that key is the rule above rather than an exception to this one. What those two leave is two further shapes. One is order-dependent before the record existed too; the other the record itself opens, because the record is per-replica *evidence* and a destranding seam destroys the evidence at exactly the recipients it serves. First, a second envelope of one id that the buffer holds **nothing tagged** to contradict — because the other copy already committed out of the buffer under a different group; because it carries **no** tag at all, which is what a filtering seam's destranding produces; or because the other copy was **released by a disagreement**, the door the rule above opens, where whether the buffer still holds it when the second envelope lands is the arrival order's. The honest group's own member is then left holding on an id that will never join it. Its members converge on eviction; which keys each replica has spent does not, and on some arrival orders the record adds a *third* reading where the un-recorded replica had two — the conflict rule fires on the orders that buffer the disagreeing copy and not on the rest. Measured against the un-recorded replica over 392 forged pools it is better on 156 and worse on 4, so the record improves this shape without closing it. And second: a recipient served a group **destranded** never buckets it, so it never spends the key, while the author spends it at the mint and a whole-delivery recipient spends it on commit — a later stray under that id then merges at those two and is held at the destranded one, where before the record all three held it alike. Eviction closes it; the destranding seams knowing the keys they cut is the other way, and it has to answer what a projection may reveal about a group that straddles its cut. A projection drops the record whole: a key names an author and a group, never a partition, so a kept one would count the groups a withheld partition resolved — the same inference the causal-frontier scrub closes. A key spent on a **disagreement** is dropped there too, and reaches that recipient without any cut at all, so the same seam owes the same answer for a group that never straddled it. So the residue is what no local judgement separates: a member that never arrives looks exactly like one still in flight. The replica exposes a way to give up rather than a rule — **eviction** untags every group still waiting, and how long to wait first is the caller's policy, the core reading no clock. Eviction untags rather than discards for the reason a filter destrands its survivors: the members are ops the replica holds and no peer will send again, so dropping them diverges, while untagging costs only the atomic view a group that never completes was never going to deliver. A replica that never evicts holds those members and does not converge with one that does — which is what makes eviction a policy every deployment runs, not an optional cleanup. Eviction **spends the bucket key** it gives up on, which is what makes a bare period over a seam with no notion of a bucket's age safe for a caller that keeps ticking: a tick landing between two members of an *honest* group untags it, and the member that follows is then a stray of a key the replica has already spent rather than the first member of a fresh group, so replicas ticking out of phase converge on the state a reader sees rather than diverging on it. What remains is narrower and still real — a replica that never evicts at all does not converge with one that does; a single tick placement can still leave a group's member unread until the next tick, so it is a policy that repeats rather than one tick that carries the guarantee; and the buffer residue of members untagged but not yet ready still differs by tick placement, so snapshot bytes are not preserved even where the reading is. diff --git a/DECISIONS.md b/DECISIONS.md index 53382571..5d2dc7f2 100644 --- a/DECISIONS.md +++ b/DECISIONS.md @@ -7,6 +7,28 @@ Log of design changes to [ARCHITECTURE.md](ARCHITECTURE.md) that implementation The entries below (2026-07-02) are a backfill: design changes made during the v0.1→v0.2 build that predate this log, recovered from the sessions and commit history. +## 2026-08-09 · C47 minority `count` rewrite · a bucket whose members disagree **spends its key** rather than holding — reversing what C3 decided a disagreement means + +**Changed:** ARCHITECTURE §Opt-In: Atomic, and the change is an **inversion rather than an extension** — a reader six months out should see that C3's "a disagreement means hold" was overturned deliberately. The sentence "a bucket whose members disagree names no group and is never complete" stood alone as the bound on a rewritten `count`; it now continues into the rule that such a bucket resolves its key at the point it disagrees. Three tests that pinned the hold inverted with it, named below; the C3 entry in this file and the C3 and C21 entries on the board gained forward pointers so the superseded rule is not read as current, and C21's entry here is corrected where this unit overtakes it. The paragraph listing what the C21 record deliberately does not take loses the minority rewrite from its residue and narrows "does not release a bucket that merely *looks* unreachable" to the count a bucket is short of, which is the case that argument was ever about. No wire or state format moves: the key lands in `resolved_tx`, which `STATE_VERSION` 13 already carries. + +**The defect, reproduced before it was fixed.** Unanimity is judged over the members that have *arrived*, so it bounds a rewrite rather than closing it. Take an honest three-member group and rewrite **two** members to declare 2, leaving the third at 3 — every member legal on its own terms, `is_admissible` passes on each. Delivered `x,y,z` or `y,x,z` the unanimous pair reaches its own declared size, commits, spends the key, and the dissenter lands as a stray of a resolved group: all three present. Delivered any of the other four ways the dissenter is in the bucket from the start, `tx_declared_count` answers `None` forever, and nothing lands at all. Measured on `main` before the change: **two distinct `encode_state` readings over the six orders**, split 2/4. It is cheaper for an attacker than rewriting every member, though not the cheapest of the four shapes on rewrite count — re-tagging one unrelated op into a live group needs a single envelope. + +**The rule, and why it is the reversal rather than a repair.** C3 read a disagreement as a reason to *hold*, on the argument that whether a bucket looks unreachable is a function of which members have landed. That argument is correct about a bucket short of its count and wrong about a disagreement, and the difference is monotonicity: two members the buffer **already holds** contradicting each other is a property of what has arrived, and no later arrival repairs it — a group whose members disagree cannot honestly complete on any delivery. Holding it is therefore the order-dependent choice, not the safe one, because the arrival order decides whether the bucket ever holds the disagreement at all. So `drain_buffer` spends the key of every disagreeing bucket at the same point a commit spends one: the members it holds are released untagged and every member still to come is a stray of a resolved key. Every order then lands all three and spends the same key, and the byte-comparison across all six orders collapses to one reading. + +**This reversal was not available to C3, and that is why it reads as a reversal rather than as C3 having been wrong.** C3's cold review rejected releasing on disagreement with a concrete counterexample — `m1(3), m2(3), m3(2)` delivered `m1,m3,m2` sees the disagreement at `m3`, releases `m1` and `m3`, and then `m2` arrives into an **empty** bucket where it looks perfectly sound at 1 of 3 and is held, while `m1,m2,m3` releases all three. That counterexample is exact, and what disarms it is C21's `resolved_tx`: the release spends the key, so `m2` arrives under a resolved key, is untagged, and merges. Releasing on disagreement is safe **only** with the record underneath it — which is why C3 was right to refuse it, C21 had to land first, and this unit is its own decision rather than a line C21 could have taken. + +**Three tests inverted, and each inversion is a claim.** `a_rewritten_first_member_count_does_not_commit_the_group_at_the_wrong_size` pinned the pair held until eviction; the release is now `a_disagreeing_bucket_releases_what_it_holds_the_moment_it_disagrees`, which pins it at the contradicting arrival rather than at eviction and pins eviction finding nothing left. `a_bucket_whose_members_disagree_on_the_size_never_completes` pinned the same hold with the disagreement arriving last; it never completing is still true and is no longer the observable, so it is `..._spends_its_key`. `a_rewritten_count_holds_the_same_set_whatever_order_it_arrives_in` pinned order-independence *of the hold*, which held only for the rewrite it fixtured — 1 of 3 members rewritten, leaving a unanimous remainder of 2 that still declares 3, so no subset can complete. The split needs the remainder's size to equal what it now declares, which is the distinction and not majority-versus-minority. It is now `..._lands_the_same_set_...` and compares canonical bytes across every order rather than reading three keys. + +**The half of that first test which was *not* the hold needed a new fixture, and a falsification pass is what showed it.** The original's live claim was that a group's size is not read off whichever member the buffer happens to hold first, and the replacement above does **not** pin it: on a rewrite *downward*, "commit the pair at the rewritten size" and "spend the key and release the pair untagged" land the same members, spend the same key and leave eviction the same nothing, so both rules pass every one of those fixtures. Deleting the unanimity test from `tx_declared_count` reddened three tests before this unit and none of the three replacements after it. It is still killed by the convergence fuzz, so it would not have vanished outright — what would have vanished is any *fixture* naming the invariant, which is the difference between a suite that localises a regression and one that only reports it. It is pinned again by `a_bucket_reads_no_size_off_the_member_it_happens_to_hold_first`, which separates the two rules where a downward rewrite cannot: a two-member group with one member rewritten *larger*, where reading off the rewritten member leaves the bucket short and waiting and reading off the honest one completes it, so the two delivery orders part company. It kills that mutant. + +**The oracle is the convergence fuzz, not the fixture.** `rewritten_group_counts_converge_under_every_order` gains a third rewrite shape: every member of each group but the last rewritten to declare one fewer, so the rewritten members are unanimous among themselves and reach their own size before the dissenter lands. It reddens against `main`'s rule and passes under this one. Neither existing shape leaves a unanimous subset that can complete, so neither moves a test — but "unchanged by the rule" would be wrong about the first, and a falsification pass measured it: its relay picks `count + 1` one time in three, which *is* a disagreement, and on the fixture (member 0 at 4, members 1-2 at 3) the base commit lands nothing in any order where this one lands all three in every order. It converges either way; the rule fires. + +**What it costs, stated rather than discovered later.** The atomic view of a group a rewrite has already made unservable, for every member of it rather than only the ones following a commit — which is the price the C21 record already pays for the three shapes it closes, taken here for the fourth. An honest group is unanimous and a one-member bucket cannot disagree, so no honest traffic reaches the rule — with one unmeasured edge worth naming: a `TxId` collision between two groups of one author would now break both groups' atomicity silently rather than hold them for eviction. `TxId::derive` is a UUIDv5 digest XOR-folded from 128 bits to 64, so two groups over *distinct* member sets collide at about 2^-64 with no op-id collision at all, and the key carries the `ClientId` so the edge is confined to one author's own groups. Negligible, and owned by this rule rather than by the id-space record — a first draft said the opposite. Each spent key is charged to a bucket the replica was holding at least two members of. The bound is C21's and is the **dedup set, not the buffer** — a key outlives the ops that earned it exactly as a `seen` entry does, so eviction empties the buffer and keeps the key; a first draft of this entry said "bounded by the ops it holds", which a falsification pass measured false (200 forged ops in 100 disagreeing pairs leave 2400 B of record with the buffer empty), and ARCHITECTURE's copy of that sentence is corrected with it. What caps it is that each key costs the sender two fresh op ids it can never re-spend, which makes it *cheaper* than the commit path it joins: 24 B/op for a per-op commit against 12 B/op for a per-pair disagreement, measured, against a replica that under the old rule retained all 200 ops in the persisted buffer instead. + +**And it widens C48, which this entry should not read as though it left alone.** The record is per-replica *evidence*, and both projections (`project_zones`, `project_read_paths`) clear it whole — a key names an author and a group and never a partition, so a kept one would count the groups a withheld partition resolved. Under the old rule a disagreement produced no record, so a projection had nothing to lose there; it produces one now, so a projected recipient re-buckets the member still to come while a verbatim recipient merges it as a stray. Measured on an all-one-zone group, so nothing straddles and no destranding is involved: a verbatim and a projected recipient of the same bytes read the same on the base commit (both holding) and differently after it. It is C48's shape reached by a second door rather than a new one, eviction still collapses it, and the fix is C48's — spending the keys a projection cuts, which has to answer what re-adding them reveals about a withheld partition. Filed onto **C48** rather than settled here. + +**And it widens C46, through a door the enumeration of that shape did not have.** A second envelope of one op id is answerable only where the buffer is still *holding* the first copy — `apply` reads `buffered_tx`, which answers `None` for an id the buffer does not hold under a group. C46 enumerates two ways that happens: the other copy committed out of the buffer under a different group, or it carries no tag at all. A disagreement now takes a member out of the buffer too, so there is a third, and whether it fires is the delivery order's. Measured on three admissible envelopes — `x` under `(T1,2)`, its group-mate `y` under `(T1,3)`, and `x` again under a second group `(T2,2)`: the base commit reads **one** state over all six orders, this one reads **two**, split 2/4, because the two orders that contradict before the second envelope arrives release `x` and never spend `T2`. Add a stray under `T2` and it is content-visible — present in all 24 orders on the base commit, absent in 8 of 24 here. Eviction collapses the document every order reads; the spent-key sets stay split, which is exactly the state C46 already describes for its other two doors ("the members still converge on eviction, and which keys each replica has spent does not"). Closing it needs the same per-op-id evidence C46 is filed for — which group a member was released *under* — so it goes onto **C46**, pinned as `a_copy_released_by_a_disagreement_leaves_the_second_envelope_nothing_to_contradict`. C21 accepted the same trade on this shape explicitly ("better on 156 and worse on 4 over 392 forged pools"); this is another 4. One smaller cost with it, stated precisely because a first draft overstated it: `apply` gains **no** new branch — the contradicting arrival takes the existing tagged path and normally answers `true` for itself. What is new is that the ops it *releases* are reported nowhere, because the FFI and wasm folds count `apply`'s `true`s one op at a time: a two-op disagreeing batch folded that way reports 1 where 2 applied. Filed as **C147**. + ## 2026-08-09 · C14 redacted-delta frontier (#398) · the carrier C9 refused for the snapshot seam is the right answer one seam over — and a redaction owes the recipient **both** records its mint reads, without burying the ops it names **Changed:** ARCHITECTURE §Wire-Level Redaction gains a paragraph ruling what a redaction owes its recipient's own authorship and how each of the three seams pays it. Wire: one new frame, `Message::Frontier { channel, seqs, reach }` (tag 54), server→client only — 53 went to C55's `ReplicateMeta` while both branches were open. State: `STATE_VERSION` 15, carrying the published-but-unheld run. @@ -528,7 +550,7 @@ And **both snapshot projections stop redacting writes made into a displaced cont **What that leaves, and why each is a separate unit rather than a compromise here.** Three shapes. Two are order-dependent on `main` before the record existed; the third the record opens, and it is the one cost this unit adds rather than removes. First, a second envelope of one id that the buffer holds **nothing to contradict**. That covers two ways in: the other envelope names a *different* group whose bucket already completed and applied the id, so it is out of the buffer before the honest copy arrives; or it carries **no tag at all**, which is what a filtering seam's destranding produces, so there is no held group to disagree with. Either way the honest group's own member waits on an id that will never join it. The record only reaches the case where the buffer is *holding* the id under a group, so what it closes is a second envelope that **disagrees with the copy the buffer holds** — in group id or in declared size, either one — and what it misses is a second envelope with nothing tagged to disagree with. On this shape the record does not merely narrow the split: on the orders that buffer the disagreeing copy the conflict rule fires and everything lands, so a pool that read two ways without the record can read three with it. Measured over 392 forged pools it is better on 156, unchanged on 232 and worse on 4, and eviction still collapses every one to a single reading. It is order-dependent on `main` before this change too, and the record narrows it from a split in what a replica reads to a split in which keys it has spent — the members still converge on eviction. Closing it needs per-op-id evidence (which group a member was *committed under*, bounded by the dedup set) rather than per-key, which is a second persisted structure and wants its own decision. Filed as **C46** and pinned as a passing test. -Second, a **minority** count rewrite, which a cold review found and which shows C3's unanimity rule bounds a rewrite rather than closing it: unanimity is judged over the members that have *arrived*, so rewriting exactly k of a group's members to declare k lets that subset reach its count and commit before the dissenter lands, while the orders that deliver the dissenter first name no group at all. Measured: two distinct states over the six orders of a three-member group with two members rewritten — cheaper for an attacker than any of the three shapes above, since it rewrites a minority. Closing it means a disagreeing bucket **spends** its key rather than merely never completing, which converges every order but reverses what C3 decided a disagreement means and the three tests that pin it. That is its own decision, filed as **C47**; the ARCHITECTURE sentence that implied unanimity protected the group is corrected here rather than left to overclaim. +Second, a **minority** count rewrite, which a cold review found and which shows C3's unanimity rule bounds a rewrite rather than closing it: unanimity is judged over the members that have *arrived*, so rewriting exactly k of a group's members to declare k lets that subset reach its count and commit before the dissenter lands, while the orders that deliver the dissenter first name no group at all. Measured: two distinct states over the six orders of a three-member group with two members rewritten — cheaper for an attacker than any of the three shapes above, since it rewrites a minority. Closing it means a disagreeing bucket **spends** its key rather than merely never completing, which converges every order but reverses what C3 decided a disagreement means and the three tests that pin it. That is its own decision, **taken by C47 (2026-08-09, #402)** — which also widens the first shape above and the destranding shape below, both recorded there; the ARCHITECTURE sentence that implied unanimity protected the group is corrected here rather than left to overclaim. Third, and the one the record *adds*: it is per-replica evidence, and a **destranded** recipient has none. The author spends the key at the mint and a whole-delivery recipient spends it on commit, but a recipient the zone or read filter served the same group untagged never buckets it and never spends it — so a later stray under that group's id merges at the first two and is held at the third, where before the record all three held it alike. Measured, and it needs the forged stray this unit exists for, so it is attacker-triggered rather than ordinary traffic. Eviction closes it. The alternative is for the destranding seams to spend the keys they cut, which they do know — but a snapshot projection would then have to say what it may reveal about a group straddling its cut, and that is a privacy decision of its own. Filed as **C48**. @@ -753,9 +775,9 @@ One seam interaction worth naming, since C26 (#370) scopes its own caveat more n **The line that mattered most: only a judgement on a member's own declared size is order-free.** The first cut untagged a whole *bucket* the moment it looked unreachable — a size out of range, members disagreeing, more members than the size admits. Cold review broke it, and the break is the interesting part of this unit. Whether a bucket looks unreachable is a property of *which of its members have landed*, so it is a function of the delivery prefix, not of the op set. Members `m1(count 3)`, `m2(count 3)`, `m3(count 2)`: delivered `m1,m2,m3` the disagreement is seen with all three present and all three are released; delivered `m1,m3,m2` it is seen at `m3`, `m1` and `m3` are released, and `m2` then arrives into an *empty* bucket where it looks perfectly sound at 1 of 3 — and is held. Two replicas, one op set, different states. Worse than the defect it replaced: the old first-member rule stranded on 2 of the 6 orders, the bucket rule on 4. -So the rule is now split by what it reads. A member declaring a size outside the cap is **refused** where it arrives — a per-member predicate every replica evaluates identically whatever else has landed, and the answer the codec already gives at the wire boundary. Refusing rather than untagging is the point of a bound whose purpose is retention: the op is not held at all. (The two seams differ in blast radius, deliberately: the codec fails the whole framed batch it is decoding, while `apply` drops the one op and lets its honest group-mates through. The wire path has a frame to reject and no way to say which member spoiled it; the in-process path has neither.) And a bucket's size is what its members *agree* it is, which is a completeness test rather than a release: a bucket without unanimity is never complete, so it is held. +So the rule is now split by what it reads. A member declaring a size outside the cap is **refused** where it arrives — a per-member predicate every replica evaluates identically whatever else has landed, and the answer the codec already gives at the wire boundary. Refusing rather than untagging is the point of a bound whose purpose is retention: the op is not held at all. (The two seams differ in blast radius, deliberately: the codec fails the whole framed batch it is decoding, while `apply` drops the one op and lets its honest group-mates through. The wire path has a frame to reject and no way to say which member spoiled it; the in-process path has neither.) And a bucket's size is what its members *agree* it is, which is a completeness test rather than a release: a bucket without unanimity is never complete, so it is held. (**Superseded by C47, 2026-08-09** — the hold, not the unanimity: once C21's resolved-key record exists, the counterexample above is answered and a bucket that disagrees spends its key instead.) -**What that cannot reach, and why it is filed rather than fixed.** A rewrite consistent across *every* member is, to a receiver, an honest group of the smaller size followed by a stray: it commits at the size it was told, and which members that is belongs to the arrival order — so a three-group retagged to two commits `{x,y}` delivered one way and `{y,z}` delivered the other. Two more shapes reach the same place. An unrelated op of the same author carrying a live group's id completes the bucket with the wrong membership. And two *copies of one op id* carrying different envelopes — a hostile duplicate, or a destranded copy racing the tagged one across two delivery seams — leave the bucket reading whichever envelope won the dedup, which is the arrival order again; that one is not what unanimity introduced, since reading the count off the first-arriving member had it too, on a different set of orders. What is common to all three is that a bucket key is *consumed* when it resolves and nothing records that it did, so a receiver cannot tell a late member of a resolved group from the first member of a fresh one. Closing them needs that record, persisted alongside the buffer: new state, and its own decision (C18). The alternative is the one already rejected above — releasing a bucket that merely looks unreachable is arrival-order dependent and strictly worse. +**What that cannot reach, and why it is filed rather than fixed.** A rewrite consistent across *every* member is, to a receiver, an honest group of the smaller size followed by a stray: it commits at the size it was told, and which members that is belongs to the arrival order — so a three-group retagged to two commits `{x,y}` delivered one way and `{y,z}` delivered the other. Two more shapes reach the same place. An unrelated op of the same author carrying a live group's id completes the bucket with the wrong membership. And two *copies of one op id* carrying different envelopes — a hostile duplicate, or a destranded copy racing the tagged one across two delivery seams — leave the bucket reading whichever envelope won the dedup, which is the arrival order again; that one is not what unanimity introduced, since reading the count off the first-arriving member had it too, on a different set of orders. What is common to all three is that a bucket key is *consumed* when it resolves and nothing records that it did, so a receiver cannot tell a late member of a resolved group from the first member of a fresh one. Closing them needs that record, persisted alongside the buffer: new state, and its own decision (C18). The alternative is the one already rejected above — releasing a bucket that merely looks unreachable is arrival-order dependent and strictly worse. (C47 later takes exactly that alternative for the *disagreement* case, which the record this paragraph asks for is what makes safe.) **Everything left over waits for eviction, and eviction is a way to give up rather than a rule.** A member that never arrives is indistinguishable from one still in flight; no local predicate separates them, and the shapes above collapse into that same case once nothing may be released early. `Document::evict_partial_transactions` untags every group still waiting and reports how many it gave up on; how long to wait first is the caller's policy, since the core reads no clock (`Host` carries entropy and a wall clock for UUIDv7, not a scheduler). A member arriving after its group was evicted forms a group of its own again — the tag it carries is the only record either side keeps — so this is a periodic policy, not a one-shot repair. **The consequence is stated rather than hidden: a replica that never evicts holds those members and does not converge with one that does.** Eviction untags rather than discards, for the reason C11 gave the filter seams: the members are ops the replica holds and no peer will send again, so dropping them diverges, while untagging costs only the atomic view a group that never completes was never going to deliver. It leaves the id accounting C6/C9 depend on untouched — an untagged member is applied (moving `buffered` → `seen`) or still buffered, never free for the sequence counter — and the ordinary readiness gate still runs, so a member whose container has not arrived waits on alone (C1's rule, unchanged). diff --git a/KANBAN.md b/KANBAN.md index 4375e6f4..c39e41ac 100644 --- a/KANBAN.md +++ b/KANBAN.md @@ -77,7 +77,9 @@ _Derived from code + git; a convenience view, not the source of truth._ **C5 — fan-out skipped the whole writing *connection*, so two channels of one session on one room never exchanged ops (crates/server) — DONE (#377).** `Registry`'s broadcast, its redacted sibling and the awareness fan-out all `continue`d on `*peer == id` — the writing **connection**, not the writing channel. That rested on one connection = one replica ("nothing echoes back to the sender"), a premise C4 broke by giving each channel its own `ClientId::for_channel` author: channel B's replica never received channel A's writes while the connection lived, and B's `last_seen_seq` never advanced over them, self-healing only on reconnect via `resume(B)`'s log delta. **Fix:** a `WriteOrigin { conn, channel }` — the channel the write arrived on — threaded into both op fan-outs; each connection's recipient set is its whole `(room, branch)` subscription minus that one channel, on the writing connection alone (handles are numbered per connection, so a peer holding the same handle is untouched). Resolving the recipient set *before* the read verdict and the migration translation also drops the redact-and-translate work for a connection with no channel to send on. **Awareness is ruled the other way, deliberately:** presence is one entry per room keyed by `(Hello client id, authenticated actor)`, both connection-scoped, so a connection's channels share one presence rather than holding two, and `AwarenessUpdate` carries only the actor — an echo would hand a client its own presence back as a peer's, indistinguishably. The whole connection stays excluded there; ARCHITECTURE §Connection / Multiplexing + DECISIONS record the asymmetry. Spec `crates/server/tests/channel_fanout.rs`: a sibling channel receiving a write with no reconnect, its `last_seen_seq` advancing, the redacting path taking the same exclusion, and the unchanged no-echo-to-self / stream-scoping / peer-delivery behaviour. -**C21 — a bucket resolved at a rewritten size committed the wrong members and held the right ones, order-dependently (crates/core) — DONE (#372).** A group's bucket key is `(author, TxId)`, and committing it *spent* the key with no record kept, so a later member of that key started a fresh bucket at an arrival count the group had already met and no arrival could meet again — and **which** members committed was the delivery order's, so one op set folded to two states. Three rewrites reached it, none malformed on any member's own terms: every member of a three-group rewritten to declare two; an unrelated op of the same author re-tagged with a live group's id; and one op id delivered under two envelopes that disagree, since `apply` dedups on op id before reading the tag. **The fix is the record C3 named and did not build:** `Document::resolved_tx: HashSet<(ClientId, TxId)>`, and a member arriving under a resolved key is untagged and merges standalone. A key is spent at each of the four points a bucket resolves — when it **commits** (`drain_buffer`); when the author **mints** the group (`tag_atomic`), because the author applies its own edits as it makes them and buckets nothing, so without this it holds a stray every receiver merges; when **eviction** gives up on it (`evict_partial_transactions`), or a later member waits on a bucket already released and two replicas on one policy disagree over which of them had ticked; and when a member arrives naming a group other than the one the buffer holds that same id under, which spends **both**, since only one of the two can ever hold the id. That fourth point fires only where `Document::apply` sees the second envelope: `Hub::ingest_records` drops an already-seen id from the batch before the fold, so a room replica never reaches it — it is the client, FFI/wasm/SDK and offline seams it protects. Rides `encode_state` (`STATE_VERSION` 12 → 13) so a restart does not re-open it; both projections drop it whole, a key naming an author and a group but never a partition. **Two rules deliberately not taken, and the first is what the review pass turned on.** The record does not fire on an id merely **applied**: a resend is ordinary traffic on any transport that retries, so spending a key there makes state a function of how often an op arrived rather than of which ops did — the same law broken in the other dimension — and a fabricated group id on a duplicate would grow the persisted record without bound (the first pass had exactly this, and a cold review caught it). And it does not release a bucket that merely *looks* unreachable, C3's standing rejection. `is_admissible` is now judged **before** the dedup, so an envelope no replica may hold decides nothing about the groups this one holds. Spec: the convergence law itself — every arrival order of one op set compared on `encode_state` bytes, for each of the three rewrites, plus the author/receiver pair and a snapshot restore. Five cold-review passes drove it: the record firing on a merely-*applied* id, an eviction regression against `main`, a 44x per-duplicate buffer scan, four new lines no test killed, an unpinned encode order, and a claim that the destranded race was closed when it is not. The residue is three further shapes, filed as **C46**, **C47** and **C48** — the first two order-dependent on `main` already, the third the one cost the record adds. Miri-clean. See DECISIONS (2026-07-29). → *Transactions*. +**C47 — a minority `count` rewrite folded one op set to two states, because unanimity is judged over the members that have *arrived* (crates/core) — DONE (#402).** C3 bounded a rewrite with "a bucket whose members disagree names no group and is never complete", and that bounds rather than closes: a unanimous **subset** reaches its own declared count before the dissenting member lands. **Reproduced before fixing**, on the entry's own fixture — an honest three-member group with two members rewritten to declare 2 and the third left at 3, every member legal on its own terms. Delivered `x,y,z` or `y,x,z` the pair commits at 2, spends the key, and `z` lands as a stray of a resolved group; delivered any of the other four ways the dissenter is in the bucket from the start, `tx_declared_count` answers `None` forever and nothing lands. Measured: **two distinct `encode_state` readings over the six orders, split 2/4**. **The fix is one rule and it is a reversal:** a bucket whose members disagree can never honestly complete on any delivery, so `drain_buffer` **spends its key** where a commit spends one (`resolve_disagreeing_tx`) — the members it holds are released untagged and every member still to come is a stray of a resolved key. What separates a disagreement from the bucket C3 refused to release is **monotonicity**: two members the buffer already holds contradicting each other is a property of what *has* landed and no arrival repairs it, where a bucket merely short of its count is a claim about what has not arrived. No format moves — the key lands in `resolved_tx`, which `STATE_VERSION` 13 already carries. **Three tests inverted**, each restated rather than deleted: the two `..._never_completes` / `..._does_not_commit_...` holds become `a_disagreeing_bucket_releases_what_it_holds_the_moment_it_disagrees` and `a_bucket_whose_members_disagree_on_the_size_spends_its_key`, pinning the release at the contradicting arrival and eviction finding nothing left; `a_rewritten_count_holds_the_same_set_...` becomes `..._lands_the_same_set_...` and compares canonical bytes across every order instead of reading three keys. A **fourth** test was added rather than inverted, and a falsification pass is why: the first test's live half — that a group's size is not read off whichever member the buffer holds first — is *not* pinned by any of the three replacements, because on a rewrite downward both rules land the same members and spend the same key. Deleting the unanimity test from `tx_declared_count` reddened three tests before this unit and none after it. `a_bucket_reads_no_size_off_the_member_it_happens_to_hold_first` pins it again, separating the rules with a rewrite *larger* on one member of a pair, and kills that mutant. The oracle is the convergence fuzz, not the fixture: `rewritten_group_counts_converge_under_every_order` gains a third rewrite shape (every member but the last rewritten to declare one fewer, so the rewritten members are unanimous among themselves), which reddens against `main`'s rule and passes under this one. ARCHITECTURE §Opt-In: Atomic's unanimity sentence changed with it, and the reversal is only safe because C21's record is under it — C3's own counterexample against releasing on disagreement is answered by the spent key, not by this rule. **Widens C48** (a projection clears the record, so a disagreement's key is evidence the projected recipient does not get) **and C46** (a disagreement releases a member out of the buffer, so a second envelope of that id finds nothing tagged to contradict — one state over six orders on the base commit, two with the rule); both named on their own entries and in DECISIONS, pinned by tests, and closed by neither this unit nor eviction, which does collapse what a reader sees. Residue filed as **C147** and **C148**. Miri-clean. See DECISIONS (2026-08-09). → *Transactions*. + +**C21 — a bucket resolved at a rewritten size committed the wrong members and held the right ones, order-dependently (crates/core) — DONE (#372).** A group's bucket key is `(author, TxId)`, and committing it *spent* the key with no record kept, so a later member of that key started a fresh bucket at an arrival count the group had already met and no arrival could meet again — and **which** members committed was the delivery order's, so one op set folded to two states. Three rewrites reached it, none malformed on any member's own terms: every member of a three-group rewritten to declare two; an unrelated op of the same author re-tagged with a live group's id; and one op id delivered under two envelopes that disagree, since `apply` dedups on op id before reading the tag. **The fix is the record C3 named and did not build:** `Document::resolved_tx: HashSet<(ClientId, TxId)>`, and a member arriving under a resolved key is untagged and merges standalone. A key is spent at each of the four points a bucket resolves — when it **commits** (`drain_buffer`); when the author **mints** the group (`tag_atomic`), because the author applies its own edits as it makes them and buckets nothing, so without this it holds a stray every receiver merges; when **eviction** gives up on it (`evict_partial_transactions`), or a later member waits on a bucket already released and two replicas on one policy disagree over which of them had ticked; and when a member arrives naming a group other than the one the buffer holds that same id under, which spends **both**, since only one of the two can ever hold the id. That fourth point fires only where `Document::apply` sees the second envelope: `Hub::ingest_records` drops an already-seen id from the batch before the fold, so a room replica never reaches it — it is the client, FFI/wasm/SDK and offline seams it protects. Rides `encode_state` (`STATE_VERSION` 12 → 13) so a restart does not re-open it; both projections drop it whole, a key naming an author and a group but never a partition. **Two rules deliberately not taken, and the first is what the review pass turned on.** The record does not fire on an id merely **applied**: a resend is ordinary traffic on any transport that retries, so spending a key there makes state a function of how often an op arrived rather than of which ops did — the same law broken in the other dimension — and a fabricated group id on a duplicate would grow the persisted record without bound (the first pass had exactly this, and a cold review caught it). And it does not release a bucket that merely *looks* unreachable, C3's standing rejection — narrowed by **C47** to a bucket short of its count, a disagreement being a property of what has already landed. `is_admissible` is now judged **before** the dedup, so an envelope no replica may hold decides nothing about the groups this one holds. Spec: the convergence law itself — every arrival order of one op set compared on `encode_state` bytes, for each of the three rewrites, plus the author/receiver pair and a snapshot restore. Five cold-review passes drove it: the record firing on a merely-*applied* id, an eviction regression against `main`, a 44x per-duplicate buffer scan, four new lines no test killed, an unpinned encode order, and a claim that the destranded race was closed when it is not. The residue is three further shapes, filed as **C46**, **C47** and **C48** — the first two order-dependent on `main` already, the third the one cost the record adds. Miri-clean. See DECISIONS (2026-07-29). → *Transactions*. **C27 — `DiffQuery` served an unredacted diff of the same bytes `VersionFetch` redacts (crates/server) — DONE (#375).** `Message::DiffQuery` resolved two states — `DiffKind::Versions` over two named versions, `DiffKind::Branches` over two `materialize_branch` results — and diffed them **raw**, replying with `Change`s carrying full `core::path`s and the scalar values at them, behind the abstaining room-read tier and with no channel binding at all. A zone-limited reader read any withheld partition as a diff (create `va`, wait for a write into the hidden zone, create `vb`, diff); `DiffKind::Branches` needed no versions to do it. Both are pinned as tests that fail against the previous behaviour. **The frame is now channel-keyed** — `DiffQuery { channel, kind, a, b }`, room resolved from the subscription like a version fetch, reply keyed by that room until C50 moved it to the channel — rather than given a room-plus-zone gate of its own: the alternative would put a second answer to "which partitions may this reader see" beside the one a bound channel holds, which is the drift `zone_narrowing` exists to prevent, and C15 (#368) pinned that a fetch narrows by the **channel's** set rather than the actor's entitlement. It also names the rule the surface already implied: a request that serves a room's *content* is channel-keyed, one that serves *names* or mutates may ride a room off the frame. **Both sides go through `project_served_state` before the engine sees them**, so a served change list is by construction `path::diff` of the two states that reader would itself have been handed — the causal frontier aside, which the two seams scrub differently and a change list does not carry (pinned directly against the fetch seam). Filtering the change list *after* the diff — the board's guess — was rejected: a mark change carries no path at all, and a filter is a second home for the redaction rule. `ClientSession::diff_query` refuses to frame on a channel the session does not hold, as `fetch_version` already does — the channel-keyed frame put it in that class, and the server answers such a frame with a connection-closing violation. `Hub::diff_versions`/`diff_branches` take the per-side redaction as an argument, so every call site states what it narrows by (a nudge, not an enforcement — an identity closure is still one keystroke, which is what the engine-pinning suites pass); no recipient is passed (a diff is never adopted as state, and a change list carries no frontier), so the scrub goes whole. **A second defect had to be fixed for the first to work:** `authz_room` had no `DiffQuery` arm, so a diff ran with *no acting schema* — fail-closed for the gate, but zone declarations live in the schema, so the zone projection would never have run. The room-keyed management frames are still in that blind spot, filed as C49; review also filed C50 (a `DiffResult` carried no channel, so with the answers now genuinely per-channel a result could not be attributed to the query that asked; since fixed), C51 (a branch whose durable base does not decode is reported as a branch that does not exist), C52 (an element the live walk does not reach survives both read projections — the shortfall reaches every state-serving seam, and its diff-visible face is an orphaned annotation) and C53 (a log-shared branch materializes truncated past a compaction floor). C32 does not reach this seam: each side derives its gate index from its own tree. **Two smaller alignments with the fetch came out of review**: a diff now records an `Action::VersionRead` when the change list goes out (the audit action's own definition already said "version-diff read", and the diff was the untraced way to read a version), and an archived side that does not decode refuses *without closing* — the engine always failed such a query, but through `internal`, which drops a live stream over an unreadable archive. Specs: `crates/server/tests/diff_projection.rs` (17 — both halves incl. a mark anchored in the hidden zone and a structural add, the readable half of each, positive controls on both branch diffs, the fetch-oracle equality in all three deny shapes, the channel-scope-not-entitlement rule, a query from a branch-subscribed channel, an element deny whose target has left the live room, and the unnarrowable room), the reworked `diff_query.rs` (8, incl. a diff on an unbound channel), two undecodable-side cases in `version_unreadable.rs`, a diff-audit case in `audit_query.rs`, and channel-keyed frames across `protocol_diff.rs`/`client_diff.rs`/FFI/wasm/Go/Python/JS. Design in DECISIONS (2026-07-29). → *Server / Diff*. **C25 — a member placed itself into any room's replica set by minting a node id, so C13's leadership and durability gates did not hold against a compromised member (crates/server) — DONE (#369).** C13 (#365) bound a peer link to a member and gated replication, leadership and durability on it. Every one of those gates asks the same question — is the sender in `replicas_for(room)` — and placement is HRW over the member set, a pure and publicly computable function, so **which rooms a node replicates follows from its node id**. Gossip's join path lets an unknown node introduce itself, and `add_members`/`merge_liveness` put it straight into the `Cluster` the ring is built from. So an admitted member ground an id HRW placed on the room it wanted, introduced itself under it, and was inside that room's replica set: it superseded the leader with a forged epoch, pushed ops in, and acked as one of five so majority-ack released a client `Accepted` for a write two nodes held. **The fix is that learning a member and placing rooms on it are two admissions.** The roster is what a node dials, probes and gossips about; the ring is built from the **adopted** members alone, and a gossip-learned member is *pending* — reachable, converging, advertised back, in no room's replica set and no room's quorum — until the cluster adopts it. **Adoption could not be a local check**: placement must be identical on every node or two nodes disagree about who leads a room permanently, and "verified" is per-node by construction. So the *evidence* is disseminated instead of the verdict — a node records only that it completed an identity-checked peer link to a member, and that claim rides the same anti-entropy as liveness, as one `verified` flag per `Message::Gossip` tuple, recorded **against the member the receiving link is bound to** and against nobody the payload names. A member is adopted once two already-adopted **trust units** have verified it. **Cold review overturned four things in the first cut, and each is the interesting part.** (1) *Adoption is derived, never accumulated.* Sticky adoption made the ring a function of history — a member adopted before its vouchers were reaped stays placed on the node that saw that order and is never placed on one that did not, permanently — so the adopted set is recomputed from `configured` + `verifiers` to a fixpoint on every change, and is a pure function of state. (2) *A verifier is a trust unit, and a trust unit is a host.* Counting node ids was the mint one level up: a certificate names a host (C13) and a host mints unlimited ids, so a member holding two of them raised the whole bar alone. The bar counts distinct verifier **hosts** and excludes the candidate's own — a member vouching for a sibling on its host vouches for itself. (3) *An inbound link verifies nothing*, however well its certificate names the member: the member chooses when to dial in, so the vouch is one it caused, and it could dial in under each ground id in turn. Verification is this node's own dial; an indirect (ping-req) confirmation is liveness only. (4) *A member is dialed at its own node id.* The recorded advertise address was first-write-wins over a field any peer may set, and the ring turned on it, so whoever advertised a joiner first decided what every later dial verified — two nodes that saw it in a different order placed 27 of 64 sample rooms differently, forever. A node id *is* an advertise address, so the second name is dropped and the dial address is a function of the id alone; that also closes the reply half's freedom at the root rather than by a filter. **The bar is a constant**, not a fraction of the cluster (clamping it made the bar node-local and let a reap *lower* it and retroactively adopt), with one exception keyed on configuration: a node configured with no peers has no cluster to be outvoted by and its ring is what it has itself reached — the single-node deployment. Configured members are adopted from birth. A majority was rejected (an honest node dials every member it knows, so the marginal verifier past the second only delays a joiner) and so was relaying another node's verifications (with no signatures on the wire, "A and B verified X" is free to write). On the reply half a claim counts only where the dial established the sender, so under `CRDTSYNC_CLUSTER_REQUIRE_PEER_IDENTITY=1` a plaintext member's vouches are dropped — sound outbound, and exactly what C13's pass 3 rejected inbound. **Cold review ran twice and pass 2 overturned two more, both about what a trust unit is:** the unit was the host as *raw text* while the certificate binding compares it semantically, so `evil.example` beside `evil.example.` (or an IPv6 literal beside its expanded form) was one certificate and two vouchers — one machine holding two units adopts any id **anywhere**, which is not the accepted residual; and the bootstrap exception read `configured.len() <= 1` off a set that shrinks, so a node given one seed fell to the single-node rule the moment that seed was reaped and began placing on one vouch while its peers held the constant. **Pass 3 found two more, one root: an advertise address had many spellings and nothing reduced them.** The dial URL was built from the authority verbatim while the host was read as everything before the first `:`, so `wss://a.example:1@b.example:9000` *read* as `a.example` and *dialed* `b.example` — a member certified for `a.example` alone ground such an id onto any room, bound its own certificate to it, and let honest nodes verify it by dialing the honest node it embedded (a full bypass under mTLS, not the bounded own-host residual). And ids were never canonicalized, so `x:9000` / `ws://x:9000` / `WS://x:9000` / `x:9000/anything` were four members one honest endpoint answers for: all truthfully verified, all adopted, none of them ever speaks, so a room whose replica set filled with them waited on acks that never came — targeted, since HRW is public. An advertise address is now an authority and nothing else, in one canonical spelling, enforced at config, at the gossip tuple and at a link's claim. A present-but-empty `CRDTSYNC_CLUSTER_PEERS` (a clustered node running the single-node bar of one) is refused at startup. **Pass 4 found three more, all in how an id is derived from configuration:** the bootstrap exception read a *de-duplicated set*, so a peer list of only this node's own address collapsed to one member and dropped the bar to a single vouch — the very state the empty-list refusal had been added to catch; `CRDTSYNC_NODE_ID` was the one door that did not canonicalize (a regression from the plain `trim()` it replaced), so a node written in a non-canonical spelling carried a **doppelgänger of itself** — adopted on two honest vouchers, placed on 468 of 2000 sample rooms it never answers for, unable to refute a suspicion of itself, dropping every follower-head report; and the port was not canonicalized, so `:9000`/`:09000`/`:009000` were three ring positions on one listener, which is the targeted write-stall pass 3 was supposed to have closed and which an operator reaches with no attacker by writing a leading zero. The bar now reads the list rather than the set, every door canonicalizes, the port is read as a number, and an address with no canonical form is refused where it is written. **Pass 10 caught a fix that was worse than the defect it closed, which is the reason the loop runs to a clean pass rather than to a fixed one.** Pass 9 measured a real window — a claim naming a member the receiver had not yet met was dropped with its tuple, so retention, and therefore the ring, depended on frame arrival order (959/2000 rooms). Holding the claim instead closed it and opened an unbounded map: those ids are on no roster, so `reap_dead` can never strike them and `rebuild_placement` never scans them, making them state nothing reclaims. One inbound 24.9 MB frame retained **376.5 MB — 15x the wire** — and stalled the single-threaded registry actor 591 ms; at the default 64 MiB frame cap that is 3.3 million claims per frame. It also made the tombstone refusal expire with `TOMBSTONE_RETENTION_TICKS`, letting a returned member be adopted pre-verified (980/2000 rooms). The window was the cheaper defect and is not even a loss, since a verifier re-advertises what it verified every round: a claim that outruns its subject is re-sent on the next round with that same verifier — O(cluster size) intervals, since a node gossips to one uniformly-random peer per interval and claims are never relayed (C39). Measured at n=33: mean 6.8 intervals to hold any voucher's claim, p95 23, max 86; mean 22.6 to adoption, max 160. Reverted, with the healing property pinned instead. The regression test written for the held-claim version was itself vacuous — it drove `merge_liveness`, which admits the unmet member from the same tuple, so it passed on the very commit it was written to catch. **Passes 5 and 6 found two more, both places where a rule was stated but not made structural:** the canonical form was checked to be a fixpoint over a list of spellings rather than by construction, and brackets — the IPv6 literal's syntax — were never required to *contain* an IPv6 literal, so `[10.0.0.1:9000]:9443` (an operator typo, no attacker) parsed as a host whose own text carries colons and yielded the id `10.0.0.1:9000:9443`: accepted at the config door, adopted from birth as a configured member, placed on every room HRW gave it, and dialable by nobody, so those rooms could reach no quorum — and written as `CRDTSYNC_NODE_ID` it has no canonical form at all, so peers drop it from gossip and refuse its `PeerAuth` and the node joins nothing while believing it has. And a verifier claim gated the *node* it was about but never the **sender** it was stored under, so a certified host opening a link per port banked an unbounded verifier entry per link against every member — keyed on ids no reap ever strikes, because reaping strikes members. Both bars are now the same one predicate on both sides. **Residual, pinned as a passing test and filed as C34:** a member that owns a host owns every id under it, answers at each ground id, and honest nodes verify one *truthfully* — so adoption bounds the mint to the member's own host and to reachable ids, and cannot bound it further; closing it needs the ring to weigh trust units rather than ids. With no certificates the bar is not a bar at all (a secret-holder binds a link to any id, C13's own residual) and what remains is reachability. Suites: `crates/server/tests/adoption.rs` (51); `gossip.rs` convergence tests now drive rounds until the ring converges as well as the roster; C13's residual test is removed from `peer_identity.rs`. Wire: `Message::Gossip`'s tuple gains a trailing flag byte (`MemberAdvert`). See DECISIONS (2026-07-28). → *Cluster*. @@ -99,7 +101,7 @@ _Derived from code + git; a convenience view, not the source of truth._ **C13 — every gate downstream of peer admission trusted any cluster member equally (crates/server + crates/core) — DONE (#365).** C10's cluster secret is one deployment-wide value, so it separated a member from a stranger and members from each other not at all — and three gates downstream of admission were written as though it did. **Leadership**: `gate_replica_frame` checked only that *this* node held the room, and superseded its own claim on any strictly higher epoch, so any admitted peer could push ops into any room the node replicated and strip the leader of any room it led. **Membership**: `apply_gossip` merged whatever a peer advertised, so an admitted peer planted an arbitrary address in every node's member set — which every node then dialed and handed the cluster secret to, and which joined the placement ring and counted toward each room's quorum. **Durability**: `FollowerHeads` named its own reporter and `catch_up_follower_reporting` overwrote *that* node's watermark, so one member credited a third with data it does not hold and `write_has_majority` released a client `Accepted` for a write no majority ever held. All three were reproduced against the pre-fix tree before anything was changed. **A link now says who it is**: `PeerAuth` (tag 52) carries the dialer's node id alongside the secret and the acceptor binds the connection to that member; `Conn::peer` went from `bool` to the bound `NodeId`, and the three handlers that decide against a member take the sender as an argument. **The binding is the member's advertise host**, checked against any *host* the verified mTLS client certificate names — its `dNSName`/`iPAddress` SANs only, never an e-mail or URI SAN, never the Common Name, never a wildcard — one rule, case-insensitive, port excluded (no certificate carries one), all names rather than the leading one (a node cert conventionally spells its host as a DNS name *and* as the IP literal an address may use), and the *same fact the dialer already verifies in the other direction* under C12 when it authenticates an acceptor against the address it dialed. Rejected: per-node shared secrets (closes the hole by closing the gossip-join path, O(N) rotation), a self-asserted id checked only for membership (any member can assert any member's id — it stops only the stranger the secret already stops), an explicit subject↔node map (a second source of truth, updated on every join, skewable per node), and matching the whole advertise address (no certificate can carry a port). Stated cost: members sharing a host share a trust unit. **Identity is declared like transport** — `CRDTSYNC_CLUSTER_REQUIRE_PEER_IDENTITY=1` refuses every certificate-less peer, and a node requiring it while verifying no client certificate or presenting none of its own refuses to start; C10's "an optional gate cannot close an ingest seam" does not transfer, since the secret remains the mandatory admission gate and identity only refines trust *among* admitted members. `request`-mode client-cert verification suffices, so ordinary clients still connect certless. **Gossip's rule is asymmetric on purpose**: the inbound half introduces only the sender, the reply half (from a node this one *dialed*, i.e. already in its member set) introduces freely — which is what keeps a joiner converging, at the cost of one extra round of dials. **Cold review ran nine times and seven passes overturned part of the last** — Copilot was quota-exhausted for the whole unit and cubic reported `skipping`, so local review was the only external gate. **P1** argued a frame from a member outside a room's replica set should be *fenced* rather than dropping the link, and found two real holes: the gossip rule must constrain the dial *address* too (constraining the node id alone left the same channel open one field over, since a newly-learned member's address is recorded verbatim and every node dials it), and *all* of a certificate's names must be weighed rather than the leading one. **P2** found P1's fence wrong — the drop *is* the repair path, since the steady path mirrors only fresh commits and the backlog is re-sent only by the redial catch-up, so a silently fenced frame is never re-driven and later frames stack on a gap the leader never learns about — and found the minting residual below, a re-bindable `PeerAuth`, and that an IP literal must be compared as an address rather than a string. **P3** overturned two of P2's own additions: a `FollowerHeads` member check that changed no watermark (`rooms_led_for` already filters by placement) but dropped every inbound link during a join window, and a startup refusal of plaintext members resting on a wrong premise — a member's advertised scheme describes *its own* listener, while the link carrying its identity into this node is the one it dials. **P4 found the binding itself still open**: the certificate reader returned e-mail and URI SANs and the Common Name alongside host names and the binding compared all of them, so a certificate legitimately issued for one node, carrying another node's host in its e-mail SAN or CN, spoke as that node. **P5** found P4's split had made a *presented* certificate that names no host fail **open** — it looked certless, so the claim stood unchecked where before the split it was compared and refused. **P6** found P5's IP rule one character short: the reader tested raw SAN text while the comparison normalizes first, so `10.0.0.6.` passed as a name and matched as an address. Both now go through one normalizer — a reader and a comparator that disagree about what a value *is* are a hole wherever they meet. **P7 raised no HIGH and confirmed the binding sound**, finding P6's un-gating of the self-certificate refusal one step too broad. **P8** found P7's carve-out had dropped a refusal nothing replaced — an advertise address naming no host reached neither check — which now belongs to the dial policy, where it is an address error rather than a certificate one; and drew out the recorded residual that the client-certificate posture must be uniform cluster-wide, since the condition reads this node's own listener as a proxy for whether its *peers* request a certificate, which a rolling change makes wrong in both directions. The non-verifying case therefore warns rather than refusing. **P9** raised no HIGH and found the hostless-address refusal scoped to this node alone while the docs promised it for every member — a member's hostless address was not even a permanent dial error, so it was redialed forever with no startup refusal; it now lives in the address parse every consumer shares. **Surfaced and fixed en route**: `actor_from_client_cert` read only DNS/email/URI SANs, so an **IP-address** SAN — what a cluster addressed by IP literal issues — yielded no identity and fell through to the Common Name; IP SANs now read as their textual form (the live two-node test failed on exactly this). **Two limits, each pinned as a passing test rather than implied.** Inside a room's replica set the epoch is still the only arbiter, because a genuinely promoted replica must be able to supersede a stale leader — the election's question (HRW+epoch → Raft), not a credential's. And placement is HRW over the member set, a pure and publicly computable function, while the join path lets an unknown node introduce itself — so a member can **mint** a node id that places it into a chosen room's replica set and reach these gates from inside it. Identity bounds the mint space to the member's own certified host, which is what makes the gates hold against a member *misreaching*; against a **compromised** member they need placement over verified members, which cannot be a local filter (placement must be identical on every node or the ring diverges). Filed as **C25**, a prerequisite rather than a follow-on, and stated plainly: this unit closes cross-member impersonation and third-party assertion, not a member that has been taken over. **Deployment** — the third deployment-facing change in this area, after C10's required `CRDTSYNC_CLUSTER_SECRET` and C12's `CRDTSYNC_CLUSTER_CA` + client-cert knobs. A cluster with no peer certificates runs unchanged. One already running peer mTLS (i.e. any C12 deployment that set `CRDTSYNC_CLUSTER_CLIENT_CERT` *and* `CRDTSYNC_TLS_CLIENT_CA`) must have each node's certificate name that node's advertise host **before** upgrading, because a presented certificate is decisive whether or not identity is required — such a node now refuses to start, which is the loud version of what its peers would otherwise do silently. To identify members: give each node a certificate whose `dNSName`/`iPAddress` SAN is its own advertise host, point `CRDTSYNC_CLUSTER_CLIENT_CERT`/`_KEY` at it, set `CRDTSYNC_TLS_CLIENT_CA` with `CRDTSYNC_TLS_CLIENT_AUTH=request` (the default `require` would lock out every certless application client), then set `CRDTSYNC_CLUSTER_REQUIRE_PEER_IDENTITY=1`. Keep the client-CA and peer-certificate settings uniform across the cluster: a node holding a mismatched certificate while verifying none itself only warns, and its peers refuse every link it opens. **Mutation-swept with measured counts** — 29 mutations, every one re-measured under a floor guard that refuses to report unless the run compiled, was not signal-killed, and cleared 24 of its 25 suites; all 29 cleared it, none returned zero. (The floor exists because an earlier reading came back `0` from a run killed under machine contention, and a second mutation silently failed to compile — both would otherwise have read as coverage holes.) *The binding*: admitting a certificate that does not name the claimed member fails 2, reading a *presented* certificate that names no host as certless 1, admitting a certless link under a required identity 1, admitting an empty claim 1, allowing a re-bind 1. *What names a host*: letting an e-mail/URI SAN name one 4, the Common Name 5, dropping the IP-SAN reader 6, letting a `dNSName` be an address 2, testing that without the binding's own normalization 2, keeping only the leading host name 2, folding an arbitrary run of root labels 1. *The comparison*: comparing an IP as a string 1, keeping the root label 3, preferring an address over a name for the session actor 1. *The gates*: letting the sender be this node 1, a member that does not replicate the room 5, removing the gossip filter 3, leaving its dial address unconstrained 1, leaving the head-report reporter unchecked 1. *The dial*: claiming no node id 8, an address naming no host 2. *The startup refusals*: no membership 1, no client-certificate verification 1, no identity of its own 1, a member naming no host 1, this node's own certificate naming the wrong host 2, gating that on the policy 1, refusing rather than warning where no certificate is asked for 1. **Three fixes were shipped unpinned and the sweep is what caught them** — the un-gated self-certificate refusal, its warn-vs-refuse condition, and the single-root-label fold. Each is now killed by exactly one test and that test is the one added in response, confirmed by capturing the killing test per mutation rather than inferred: `a_certificate_for_another_host_refuses_to_start_even_mid_rollout`, `a_certificate_for_another_host_still_serves_where_no_peer_asks_for_one`, `only_one_root_label_folds`. So the sweep audits the *suite*, not only the source: all three were production changes the tests appeared to exercise and did not. Specs: `crates/server/tests/peer_identity.rs` (34 — a reproduction for each of the three, the replica-set supersede pinned as the residual, the gossip narrowing including the hidden-injection and still-disseminating-a-death cases, the join path, the forged-report-releases-no-Accepted durability chain, per-connection identity, the required-identity refusal, every node-to-node frame still working, a second `PeerAuth` refused on an admitted link, plus a real mTLS cluster replicating end to end with a certless client, a certificate naming its host several ways still binding it, a cluster-signed certificate for another address reaching no peer plane, and the startup refusals), 14 binding/normalization tests in `dial.rs` and 11 certificate-identity tests in `tls.rs`, with `peer_auth.rs`/`gossip.rs`/`self_heal.rs` updated where the old contract was the defect. See DECISIONS (2026-07-27). → *Cluster*. -**C3 — a malformed `Tx.count` was an unbounded remote memory-retention primitive (crates/core) — DONE (#363).** `count` tells a receiver to hold a group's members until that many arrive, and nothing bounded it at the decode boundary or ever gave up on a partial group: `count == 0`, `count == u32::MAX`, more members held than the count admits, and members of one `(client, tx id)` disagreeing on the count all left members buffered for the life of the replica — persisted by `encode_state`, and (since C6) counted among the ids the replica holds. Completeness was also decided against `buffer[idxs[0]].tx.count`, whichever member landed first, so a group told it was smaller than it is committed at that size and left the rest holding a size their bucket had already spent. **The cap is now a protocol constant, not the "default 1000" ARCHITECTURE named and nothing enforced** — a receiver decides completeness from a declared count, so a node with its own cap would emit groups every peer refuses — enforced at *both* ends: the codec refuses `0` and anything past `MAX_TX_MEMBERS`, and `tag_atomic` leaves an oversized group **untagged** rather than tagged with a size whose refusal takes the whole framed batch with it (an oversized transaction would otherwise become dropped ops, not a non-atomic one). **The line two cold-review passes drew, and the substance of the unit: only a judgement on a member's own declared size is order-free.** The first cut untagged a whole *bucket* the moment it looked unreachable; whether it looks that way is a function of which members have landed, so `m1(3), m2(3), m3(2)` delivered `m1,m2,m3` released all three while `m1,m3,m2` released two and then held `m2` against a bucket already given up on — one op set, two states, and stranding on 4 of 6 orders where the defect it replaced stranded on 2. So a size outside the cap is **refused** where it arrives — at the codec, which fails the frame carrying it, and at the `apply` seam an in-process relay or SDK reaches without decoding, which drops the one op and lets its group-mates through; refusing holds nothing either way. And a bucket's size is what its members **agree** it is: a completeness test, not a release, so a bucket without unanimity is never complete and is held. **What no count rule reaches, stated rather than claimed:** a rewrite consistent across every member is, to a receiver, an honest group of the smaller size followed by a stray — it commits at that size, over whichever members arrived first. Two review passes caught this being sold as fixed, the second finding a third route to it (two copies of one op id under different envelopes, so the bucket reads whichever won the dedup — a hostile duplicate, or a destranded copy racing the tagged one). All three are pinned by tests asserting the two orders *disagree*, and filed as **C21**. **Everything left is undecidable locally** (a member that never arrives is indistinguishable from one still in flight), so `Document::evict_partial_transactions` is a way to give up rather than a rule: it untags every group still waiting, reports how many, and how long to wait first is the caller's policy since the core reads no clock. It untags rather than discards for C11's reason — the members are ops the replica holds and no peer will resend — and leaves C6/C9's id accounting untouched (an evicted member is applied or still buffered, never free to re-mint) and C1's readiness gate running (a member whose container is absent waits on alone). **The consequence is stated, not hidden: a replica that never evicts does not converge with one that does.** **Feeds C19, and is checked against it:** this adds a second permanent refusal to `Document::apply`, whose `bool` `Hub::ingest_records` discards — so a refused op is persisted, dedup-poisoned, fanned out and ACKed `Accepted`. No *new* reachable trigger: the wire path decodes first and the codec fails the frame before `apply` sees it, so only an in-process caller reaches this refusal, and every replica refuses the same op (a pure function of it), so state still converges. C19 owns the seam. Rejected: verifying the group id against `TxId::derive` of its members — with unanimity the wrong-size commit is already gone, so it buys nothing state-observable while newly depending on no seam ever renumbering a seq under a live tag. Suites: 18 cases in `crates/core/tests/transaction.rs` (each shape reproduced first, including a spliced snapshot for the presentation that never commits and a six-order sweep pinning that nothing is released early) plus a malformed-count arm in `crates/core/tests/convergence.rs` that fails at seed 0 against the bucket-wide rule and separately pins that a group rewritten smaller converges from every order once evicted. Mutation-swept across the whole workspace (`--no-fail-fast`): no `count` bound at the codec fails 2; no refusal at the `apply` seam 1; the count read off the first member rather than agreed 4; no cap at the mint 1; eviction made a no-op 10; eviction discarding instead of untagging 11; a committed member waved past the readiness gate (C1's rule) 8. The unmutated tree fails 0. Miri-clean. Filed rather than bundled: **C20** (nothing calls the eviction seam yet) and **C21** (a well-formed forged group still diverges before eviction). See DECISIONS (2026-07-27). → *Transactions / Scope Constraints*. +**C3 — a malformed `Tx.count` was an unbounded remote memory-retention primitive (crates/core) — DONE (#363).** `count` tells a receiver to hold a group's members until that many arrive, and nothing bounded it at the decode boundary or ever gave up on a partial group: `count == 0`, `count == u32::MAX`, more members held than the count admits, and members of one `(client, tx id)` disagreeing on the count all left members buffered for the life of the replica — persisted by `encode_state`, and (since C6) counted among the ids the replica holds. Completeness was also decided against `buffer[idxs[0]].tx.count`, whichever member landed first, so a group told it was smaller than it is committed at that size and left the rest holding a size their bucket had already spent. **The cap is now a protocol constant, not the "default 1000" ARCHITECTURE named and nothing enforced** — a receiver decides completeness from a declared count, so a node with its own cap would emit groups every peer refuses — enforced at *both* ends: the codec refuses `0` and anything past `MAX_TX_MEMBERS`, and `tag_atomic` leaves an oversized group **untagged** rather than tagged with a size whose refusal takes the whole framed batch with it (an oversized transaction would otherwise become dropped ops, not a non-atomic one). **The line two cold-review passes drew, and the substance of the unit: only a judgement on a member's own declared size is order-free.** The first cut untagged a whole *bucket* the moment it looked unreachable; whether it looks that way is a function of which members have landed, so `m1(3), m2(3), m3(2)` delivered `m1,m2,m3` released all three while `m1,m3,m2` released two and then held `m2` against a bucket already given up on — one op set, two states, and stranding on 4 of 6 orders where the defect it replaced stranded on 2. So a size outside the cap is **refused** where it arrives — at the codec, which fails the frame carrying it, and at the `apply` seam an in-process relay or SDK reaches without decoding, which drops the one op and lets its group-mates through; refusing holds nothing either way. And a bucket's size is what its members **agree** it is: a completeness test, not a release, so a bucket without unanimity is never complete and is held — the hold since reversed by **C47**, which spends such a bucket's key once C21's record makes that safe. **What no count rule reaches, stated rather than claimed:** a rewrite consistent across every member is, to a receiver, an honest group of the smaller size followed by a stray — it commits at that size, over whichever members arrived first. Two review passes caught this being sold as fixed, the second finding a third route to it (two copies of one op id under different envelopes, so the bucket reads whichever won the dedup — a hostile duplicate, or a destranded copy racing the tagged one). All three are pinned by tests asserting the two orders *disagree*, and filed as **C21**. **Everything left is undecidable locally** (a member that never arrives is indistinguishable from one still in flight), so `Document::evict_partial_transactions` is a way to give up rather than a rule: it untags every group still waiting, reports how many, and how long to wait first is the caller's policy since the core reads no clock. It untags rather than discards for C11's reason — the members are ops the replica holds and no peer will resend — and leaves C6/C9's id accounting untouched (an evicted member is applied or still buffered, never free to re-mint) and C1's readiness gate running (a member whose container is absent waits on alone). **The consequence is stated, not hidden: a replica that never evicts does not converge with one that does.** **Feeds C19, and is checked against it:** this adds a second permanent refusal to `Document::apply`, whose `bool` `Hub::ingest_records` discards — so a refused op is persisted, dedup-poisoned, fanned out and ACKed `Accepted`. No *new* reachable trigger: the wire path decodes first and the codec fails the frame before `apply` sees it, so only an in-process caller reaches this refusal, and every replica refuses the same op (a pure function of it), so state still converges. C19 owns the seam. Rejected: verifying the group id against `TxId::derive` of its members — with unanimity the wrong-size commit is already gone, so it buys nothing state-observable while newly depending on no seam ever renumbering a seq under a live tag. Suites: 18 cases in `crates/core/tests/transaction.rs` (each shape reproduced first, including a spliced snapshot for the presentation that never commits and a six-order sweep pinning that nothing is released early) plus a malformed-count arm in `crates/core/tests/convergence.rs` that fails at seed 0 against the bucket-wide rule and separately pins that a group rewritten smaller converges from every order once evicted. Mutation-swept across the whole workspace (`--no-fail-fast`): no `count` bound at the codec fails 2; no refusal at the `apply` seam 1; the count read off the first member rather than agreed 4; no cap at the mint 1; eviction made a no-op 10; eviction discarding instead of untagging 11; a committed member waved past the readiness gate (C1's rule) 8. The unmutated tree fails 0. Miri-clean. Filed rather than bundled: **C20** (nothing calls the eviction seam yet) and **C21** (a well-formed forged group still diverges before eviction). See DECISIONS (2026-07-27). → *Transactions / Scope Constraints*. **C12 — a TLS-terminated cluster could not dial its own peers, so the peer link was plaintext-only (crates/server) — DONE (#361).** Both outbound node-to-node dials hard-coded `ws://` — `runtime::connect_peer` and `gossip::relay_roundtrip`, the latter shared by the anti-entropy exchange and the SWIM ping-req — and the server had no client-side `rustls::ClientConfig` at all; `tls.rs` built `ServerConfig` only. So `CRDTSYNC_TLS_CERT` on a clustered node gave it a listener its own peers could not speak to and the node started, bound, and never converged; and because C10's cluster secret is a *bearer* credential, the one link that must not be plaintext could be nothing else. **The transport is now per member, declared by its advertise address** — `wss://host:port` terminates TLS, `ws://host:port` or a bare `host:port` does not, any other scheme is refused rather than folded into a hostname. Nothing else could decide it: whether *this* node terminates TLS says nothing about the member it dials, and the advertise address is the one per-member datum the cluster already agrees on and already gossips, so it needs no new wire field and cannot skew between nodes (the cost: a scheme change re-identifies the node, so a rolling migration is a rolling re-join through the existing reap path). **Trust anchors are explicit** (`CRDTSYNC_CLUSTER_CA`) with no platform store or bundled public roots standing in — the link carries write access to every room, so an ambient store would widen the issuer set to every CA on the host — and a `wss://` member with none configured is a startup error, not a dial that fails every round. **mTLS is the same handshake read both ways** (`CRDTSYNC_CLUSTER_CLIENT_CERT`/`_KEY`), deliberately its own configuration rather than defaulting to the listener's cert: a `serverAuth`-only server certificate would make that default work in a lab and fail in production. **A cluster may mix plaintext and TLS members** — a live replicated store cannot switch every node at one instant and forbidding it would make a rollout a flag-day restart — but never silently: a TLS node names its plaintext peers at startup, and `CRDTSYNC_CLUSTER_REQUIRE_TLS=1` declares the rollout finished and refuses them. **The silent non-convergence is now a startup error**: a node whose advertised transport disagrees with the one it terminates refuses to start in both directions, as does one that requires TLS of its peers while serving plaintext; a gossip-learned member is checked at its dial instead, and a peer never opened is named after a run of attempts (`DialStreak`, distinct from C10's `RefusalStreak` over links that opened and died). **Closes C10's harvest for TLS members** — the acceptor's certificate is verified before the first frame is written, so a squatter at a member's advertise address gets a failed handshake instead of the credential (pinned by asserting the impostor's socket never sees the secret's bytes). Cold review added the remaining defects: the dial had no timeout, so a far end that accepted the socket and went silent wedged the task owning that follower's link forever (now bounded like the accept loop's handshake); `require_peer_tls` was checked against peers only; a *permanent* dial failure (a transport or trust refusal, decided before a socket opens) was redialed — and the follower reported down — four times a second forever, so `DialError::is_permanent` now separates it from a peer that is merely down; and `CRDTSYNC_CLUSTER_REQUIRE_TLS` resolved anything unrecognised to *off*, the permissive setting, rather than erroring. Left filed, deliberately: the bearer-replay property itself (a challenge-response credential — now worth designing, and wanting a per-member identity) and C10's admission confirmation. Mutation-swept whole-crate: restoring the pre-fix `ws://` fails 10, dropping the rustls connector 4, folding an unknown scheme into the authority 3, dropping the trust-anchor requirement 2, dialing plaintext under a required-TLS deployment 2, skipping this node's own transport check 2, unbounding the dial 1, treating a configuration refusal as transient 1, admitting peer TLS material with no cluster 1; the unmutated crate fails 0. The sweep found the self-transport require-TLS test passing with its rule deleted (its peer was refused first) and that is closed; removing the redial loop's streak reset fails 0 and is recorded as untested-by-construction, since it is a warning *not* printed on a live node's stderr. Suites: `crates/server/tests/peer_tls.rs` (20), 14 unit tests in `dial.rs`, 4 for `DialStreak`, 3 for the require-TLS switch. See DECISIONS (2026-07-27). → *Cluster*. @@ -567,11 +569,13 @@ scalar / counter / register / element / map (#22–#27), list Fugue (#24), text **C134 — a sequence merge installs the winning claim as a detached deep clone (crates/core) — READY, no dependencies. Filed by C40 (#400), which widened it; the shape pre-dates it.** `List::merge` folds a peer's sequence in, and where the peer's node carries a composite the merge stores `on.value.deep_clone()` — a fresh handle, not the `Rc` the document's per-id registry holds for that element. An op addressed to that id afterwards applies to the registry's handle while the sequence renders the clone, so the two drift. The `Seated::Vacant` arm has done this since before C40 (a node the receiver has never seen is cloned in whole); C40 adds a second arm that does it, since a claim that outranks the incumbent now replaces it rather than folding the two together — which is the right *semantics* (two composites contending for one id are different elements, and folding their content together cross-contaminated them) but keeps the detachment. It is not reachable from `Document` today: there is no `Document::merge`, and the `Element::merge` cluster is entered only from the public API and `XmlElement::merge`, so nothing in the op-fold or snapshot path takes it. That is why it is filed rather than fixed here, and it is also the reason the fix is not local: `List` has no registry to resolve an id against, so either the caller supplies one or `merge` stops being a `List`-level operation. → *Core / List*. -**C47 — a minority `count` rewrite folds one op set to two states, because unanimity is judged over the members that have *arrived* (crates/core) — READY, no dependencies. Found by cold review during C21 (#372); pre-existing on `main` before it. Filed under a different id during that unit and renumbered into C21's reserved range, its first id having been claimed by a sibling unit in flight.** C3 bounds a rewrite with "a bucket whose members disagree names no group and is never complete", and C21 leaned on that. It bounds rather than closes, because a unanimous **subset** can reach its own declared count before the dissenting member lands. **Reproduction:** take an honest three-member group and rewrite **two** members to declare 2, leaving the third at 3 — every member legal on its own terms, `is_admissible` passes. Delivered `x,y,z` or `y,x,z` the pair completes at 2, commits, spends the key, and `z` lands as a stray of a resolved group: all three present. Delivered in any of the other four orders the dissenter is in the bucket from the start, `tx_declared_count` returns `None` forever, and **nothing lands at all**. Two distinct states over six orders, measured. **C21 (#372) widened it**, and this entry should not read as though it left it alone: spending the key on the minority's commit makes the stray land, which is *visible* on pools where the un-recorded replica's commit cancelled itself out and the stray stayed held. Measured over 389 byte-identical forged pools x 12 orders, the split rate goes 58 -> 85 — 28 pools read two ways here that read one way before, 1 the other way. Forged envelopes only, and eviction still collapses all 85 to a single reading. It is the cheapest of the four rewrite shapes for an attacker, since it rewrites a minority rather than every member. **The fix is one rule, and it is a reversal:** a bucket whose members disagree can never honestly complete, so it **spends its key** where a commit spends it — every order then lands all three and spends the key, and the no-duplicate convergence fuzz passes. What it costs is C3's decision that a disagreement means *hold*: `a_rewritten_first_member_count_does_not_commit_the_group_at_the_wrong_size`, `a_bucket_whose_members_disagree_on_the_size_never_completes` and `a_rewritten_count_holds_the_same_set_whatever_order_it_arrives_in` all pin the opposite and would invert, and ARCHITECTURE §Opt-In: Atomic's unanimity sentence changes with them. Distinct from **C46**, which needs per-op-id evidence; this one needs only the rule to change. ARCHITECTURE already carries the caveat that unanimity bounds rather than closes, so the docs do not overclaim in the meantime. → *Transactions*. +**C48 — a recipient served a destranded group never spends its key, so it holds a stray the author and a whole-delivery recipient merge (crates/core + crates/server) — READY, no dependencies. Found by cold review during C21 (#372); the one shape C21's record *opens* rather than closes. Filed under a different id during that unit and renumbered into C21's reserved range, its first id having been claimed by a sibling unit in flight.** The record is per-replica **evidence**, and ARCHITECTURE §Opt-In: Atomic says every filtering seam destrands the survivors of a group it cuts — so the recipient that most needs the key is the one with no evidence to earn it. **Reproduction:** author `A` mints atomic `T = {m1, m2}` and spends `(A,T)` at the mint; `R_whole` receives both tagged, commits, spends `(A,T)`; `R_cut` receives both untagged (what `Session::zone_filter` and `Document::project_zones` deliver a zone-scoped subscriber) and spends nothing. A forged stray — an unrelated op of `A` re-tagged `(T, 2)`, the same envelope rewrite C21 exists for — then merges at `A` and `R_whole` and is **held** at `R_cut`. On `main` all three held it alike, so this pair agreed before the record and disagrees after; it is the one cost C21 adds rather than removes. **C47 (#402) widened it through a second door, and one that needs no destranding seam at all:** a bucket whose members *disagree* now spends its key, and both projections (`project_zones`, `project_read_paths`) clear the record whole, so a recipient handed projected bytes re-buckets a member a verbatim recipient merges as a stray. Measured on an all-one-zone group, where nothing straddles the cut: the verbatim/projected pair read the same on C47's base commit and differently after it. Same class, same fix, same question to answer first — so the keys a projection must re-add are now the straddling ones **and** the disagreeing ones. The forged-stray shape needs the forged op, so it is attacker-triggered rather than ordinary traffic, and **eviction (C20) closes it** — which is another reason C20 is a correctness policy rather than cleanup. **The other way is for the destranding seams to spend the keys they cut**, which they know: `project_zones` computes `split_groups` before calling `destrand_split`, and `Session::zone_filter` has the same set per batch. Two things to answer first. A snapshot projection currently *clears* the record whole so a subscriber cannot count the groups a withheld partition resolved — re-adding the straddling keys reveals that those groups existed, which is a narrower leak but a real one, and the naive `extend(split)` also re-adds groups that fell **wholly** inside the withheld partition, which is exactly what the clear exists to hide. And the live wire seam has no channel to tell a receiver "this key is spent" without a protocol frame. → *Transactions*. + +**C147 — the direct fold's applied/refused report says nothing about an op that released other ops without applying itself (crates/ffi + crates/wasm + crates/core) — READY, no dependencies. Found by C47's (#402) falsification pass, measured.** ARCHITECTURE §The direct fold reports what it refused beside what it applied gives a bare byte-pipe caller two counts, on the ground that "did not apply" and "no replica will ever hold this" call for opposite responses. The FFI (`crates/ffi/src/lib.rs`) and wasm (`crates/wasm/src/lib.rs`) folds are `ops.iter().filter(|op| doc.apply(op)).count()`, so an op that answers `false` is counted as neither — and `Document::apply` has a growing set of branches that answer `false` **while mutating state**. C21 (#372) named one (the disagreeing-envelope branch spends two keys and drains). C47 (#402) did not add a branch — its contradicting arrival takes the existing tagged path and normally answers `true` for itself — but it made the *existing* one under-report: the ops a contradiction releases are counted nowhere, because the fold counts `apply`'s `true`s one op at a time. Measured: a two-op disagreeing batch folded as `iter().filter(|op| doc.apply(op)).count()` reports **1** where **2** applied. The rarer sub-case answers `false` outright — a bucket holds `m1(count 2)` and the dissenting arrival `m3(count 3)` targets a container this replica has not materialised, so it spends the key, releases `m1`, and is itself held. The commit loop's own `op.tx = None; self.hold(op)` re-hold is the same shape and predates both. So a caller folding a batch over a byte pipe is told a number that undercounts what its own fold made visible, with no way to tell that from a duplicate. **What it needs is a decision about what the seam reports**, not a patch to the count: either the fold answers what it *changed* rather than what each op did, or `apply`'s three outcomes (applied / held / refused) reach the SDK boundary as three rather than being collapsed to a bool at it. The server is unaffected — `Hub::ingest_records` ignores the return value and derives its broadcast set elsewhere. → *SDK / Transactions*. -**C48 — a recipient served a destranded group never spends its key, so it holds a stray the author and a whole-delivery recipient merge (crates/core + crates/server) — READY, no dependencies. Found by cold review during C21 (#372); the one shape C21's record *opens* rather than closes. Filed under a different id during that unit and renumbered into C21's reserved range, its first id having been claimed by a sibling unit in flight.** The record is per-replica **evidence**, and ARCHITECTURE §Opt-In: Atomic says every filtering seam destrands the survivors of a group it cuts — so the recipient that most needs the key is the one with no evidence to earn it. **Reproduction:** author `A` mints atomic `T = {m1, m2}` and spends `(A,T)` at the mint; `R_whole` receives both tagged, commits, spends `(A,T)`; `R_cut` receives both untagged (what `Session::zone_filter` and `Document::project_zones` deliver a zone-scoped subscriber) and spends nothing. A forged stray — an unrelated op of `A` re-tagged `(T, 2)`, the same envelope rewrite C21 exists for — then merges at `A` and `R_whole` and is **held** at `R_cut`. On `main` all three held it alike, so this pair agreed before the record and disagrees after; it is the one cost C21 adds rather than removes. It needs the forged op, so it is attacker-triggered rather than ordinary traffic, and **eviction (C20) closes it** — which is another reason C20 is a correctness policy rather than cleanup. **The other way is for the destranding seams to spend the keys they cut**, which they know: `project_zones` computes `split_groups` before calling `destrand_split`, and `Session::zone_filter` has the same set per batch. Two things to answer first. A snapshot projection currently *clears* the record whole so a subscriber cannot count the groups a withheld partition resolved — re-adding the straddling keys reveals that those groups existed, which is a narrower leak but a real one, and the naive `extend(split)` also re-adds groups that fell **wholly** inside the withheld partition, which is exactly what the clear exists to hide. And the live wire seam has no channel to tell a receiver "this key is spent" without a protocol frame. → *Transactions*. +**C148 — nothing observes the atomic *view* itself, so two rules that differ only in it are indistinguishable to the suite (crates/core) — READY, no dependencies. Found by C47's (#402) falsification pass, measured.** `atomic_transact`'s whole guarantee is that no peer ever sees a partial group, and every test of it reads *final* state — so it observes which members landed, never whether they landed together. C47 hit the consequence: on a group rewritten downward, "commit the bucket at the rewritten size" and "spend the key and release the members untagged" land the same members, spend the same key and leave eviction the same nothing, differing only in whether the release was one group commit or several independent merges. Deleting the unanimity test from `tx_declared_count` therefore survived the whole **fixture** suite — the randomized convergence fuzz still killed it, so the invariant was observable but nameless, which is the weaker and more common failure: a suite that reports a regression without localising it. It took a fixture built on a rewrite *larger* (`a_bucket_reads_no_size_off_the_member_it_happens_to_hold_first`) to name it, and that construction works for that one mutant rather than generalising. **What it needs is an observable**: a seam that reports state as of a point *inside* a fold, or a diff/update event stream a test can assert never renders a group partially. The reactivity seam (`Document::diff`, the SDK `Doc.on("update")`) is the closest thing that exists and is not exercised against group commits. Until then the atomicity guarantee is pinned only where a member's absence is *visible in the final read* — which is most of the interesting cases and not all of them, and the gap is invisible to a mutation sweep only because the sweep reddens on the members rather than on the view. → *Transactions*. -**C46 — a second envelope of one op id that the buffer holds nothing to contradict (crates/core) — READY, no dependencies. Found by cold review during C21 (#372); pre-existing on `main` before it. Filed under a different id during that unit and renumbered into C21's reserved range, its first id having been claimed by a sibling unit in flight.** C21's record spends a key on evidence the buffer is *holding*, so it closes a second envelope that **disagrees with the copy the buffer holds** — in group id or in declared size — and misses every shape where there is nothing tagged to disagree with. Two ways in, and the second needs no forgery at all. **(a) The envelopes name different groups.** Deliver `X` tagged `(T2, 2)` beside a genuine member of `T2`: that bucket completes and applies `X`. Deliver `X` tagged `(T1, 2)` afterwards and it is a plain duplicate of an id already applied — the buffer holds nothing under `T1` to contradict, so nothing is spent, and `T1`'s own genuine member waits on an id that will never join it. Deliver the two copies the other way round and the conflict is caught in the buffer, both keys are spent, and every member lands. **Why C21 stopped here.** Closing it means spending a key on evidence from an id merely *applied*, and that trades this law for another: a resend is ordinary traffic on any transport that retries, so a delivery that spent a key would make state a function of how often an op arrived — and a fabricated group id on such a duplicate grows the persisted record without bound. Both were measured on C21's first pass. **What it would take** is per-op-id evidence rather than per-key: recording which group a member was *committed under* (bounded by the dedup set, one entry per committed tagged op) so a later envelope naming a different group is answerable without trusting the arriving tag. That is a second persisted structure, and it wants the same explicit decision C21's record got. **(b) The other envelope carries no tag at all** — which is exactly what ARCHITECTURE says every filtering seam produces, so this half is reachable on honest traffic rather than by rewrite: an honest three-member group plus a destranded bare copy of one member folds to **2 observable states** over its 24 orders (measured, and identical on `main`), because a bare copy that lands first is applied with no group recorded and the remaining members hold on a count they can never reach. `buffered_tx` returns `None` for a held-but-untagged op and an applied one is not buffered at all, so neither reaches the conflict rule. Until then the members converge on eviction (**C20**) and only the spent-key record differs; the (a) half is pinned as `a_copy_absorbed_by_another_groups_bucket_strands_its_own_group_until_eviction`. → *Transactions*. +**C46 — a second envelope of one op id that the buffer holds nothing to contradict (crates/core) — READY, no dependencies. Found by cold review during C21 (#372); pre-existing on `main` before it. Filed under a different id during that unit and renumbered into C21's reserved range, its first id having been claimed by a sibling unit in flight.** C21's record spends a key on evidence the buffer is *holding*, so it closes a second envelope that **disagrees with the copy the buffer holds** — in group id or in declared size — and misses every shape where there is nothing tagged to disagree with. Two ways in, and the second needs no forgery at all. **(a) The envelopes name different groups.** Deliver `X` tagged `(T2, 2)` beside a genuine member of `T2`: that bucket completes and applies `X`. Deliver `X` tagged `(T1, 2)` afterwards and it is a plain duplicate of an id already applied — the buffer holds nothing under `T1` to contradict, so nothing is spent, and `T1`'s own genuine member waits on an id that will never join it. Deliver the two copies the other way round and the conflict is caught in the buffer, both keys are spent, and every member lands. **Why C21 stopped here.** Closing it means spending a key on evidence from an id merely *applied*, and that trades this law for another: a resend is ordinary traffic on any transport that retries, so a delivery that spent a key would make state a function of how often an op arrived — and a fabricated group id on such a duplicate grows the persisted record without bound. Both were measured on C21's first pass. **What it would take** is per-op-id evidence rather than per-key: recording which group a member was *committed under* (bounded by the dedup set, one entry per committed tagged op) so a later envelope naming a different group is answerable without trusting the arriving tag. That is a second persisted structure, and it wants the same explicit decision C21's record got. **(b) The other envelope carries no tag at all** — which is exactly what ARCHITECTURE says every filtering seam produces, so this half is reachable on honest traffic rather than by rewrite: an honest three-member group plus a destranded bare copy of one member folds to **2 observable states** over its 24 orders (measured, and identical on `main`), because a bare copy that lands first is applied with no group recorded and the remaining members hold on a count they can never reach. `buffered_tx` returns `None` for a held-but-untagged op and an applied one is not buffered at all, so neither reaches the conflict rule. **(c) A third door, opened by C47 (#402) and measured by its falsification pass.** A disagreement now releases a member *out of the buffer*, and `buffered_tx` answers `None` for an id the buffer no longer holds — so whether the conflict rule fires at all became the delivery order's. Three admissible envelopes reach it: `x` under `(T1,2)`, its group-mate `y` under `(T1,3)`, and `x` again under a second group `(T2,2)`. C47's base commit reads **one** state over all six orders; with the rule it reads **two**, split 2/4 — the two orders that contradict before the second envelope lands release `x` and never spend `T2`. With a stray added under `T2` it is content-visible: present in all 24 orders on the base, absent in 8 of 24 with the rule. Same fix as (a) — which group a member was released *under*, not merely that its key resolved — so it does not change what this unit needs, only how much it now buys. Pinned as `a_copy_released_by_a_disagreement_leaves_the_second_envelope_nothing_to_contradict`. Until then the members converge on eviction (**C20**) and only the spent-key record differs; the (a) half is pinned as `a_copy_absorbed_by_another_groups_bucket_strands_its_own_group_until_eviction`. → *Transactions*. **C20 — nothing calls the partial-transaction eviction seam (crates/server + crates/core) — READY, no dependencies. Found during C3 (#363), by grepping for callers.** `Document::evict_partial_transactions` is the give-up path for a group no arrival will complete, and C3 deliberately shipped the seam without a policy: the core reads no clock, so *how long to wait* is the caller's. Nothing calls it — not `crates/server`'s room replica, not `ClientSession`, not an SDK — so the retention C3 bounds for the shapes a local test can name is still open for the ones it cannot: send one member each of N groups, distinct `TxId`s, in-range unanimous counts, and every op is buffered for the life of the replica and re-persisted into every snapshot. There is no size cap on `Document::buffer` either. C2 (#390) multiplies the *other* unbounded set's growth rate: a commit now spends one `resolved_tx` key per zone partition it spans rather than one per commit, so the persisted record grows by up to `zones+1` entries per straddling commit — same retention question, a steeper slope. Wants a room-maintenance tick on the server side and a session-level equivalent — **C21 (#372) removed the prerequisite this entry was written around.** Eviction now spends the bucket key it gives up on, so a tick landing between two members of an *honest* group untags it and the member that follows is a stray of a spent key rather than the first member of a fresh group: measured over 400 pools x 11 tick placements, replicas ticking out of phase converge on what a reader sees, where `main` before it diverged on the visible value. So a bare period no longer needs an age input or an evict-only-what-was-already-waiting contract to be safe for what a reader sees — what remains is the period itself, whether the client evicts on resume, and three narrower residues: a replica that never evicts still does not converge with one that does; a single tick placement can leave a member unread until the next tick, so it is the repeating policy that carries the guarantee rather than one tick; and the buffer of members untagged but not yet ready still differs by tick placement, so snapshot *bytes* are not preserved even where the reading is. Note the sharper reason it is not optional: **a replica that never evicts does not converge with one that does** (ARCHITECTURE §Opt-In: Atomic), so this is a correctness policy, not a cleanup. → *Transactions / Sessions*. diff --git a/crates/core/src/doc.rs b/crates/core/src/doc.rs index 213b8e3b..3c80f14e 100644 --- a/crates/core/src/doc.rs +++ b/crates/core/src/doc.rs @@ -3223,8 +3223,9 @@ impl Document { /// `false` covers three unrelated situations, and only the last is permanent: /// the op is already applied or already held (a duplicate); it is admissible but /// not applicable yet, so it is buffered and replays once a create makes its - /// target reachable or its transaction group completes; or it is one - /// [`Op::is_admissible`] refuses, which no later arrival changes. A caller that + /// target reachable or its transaction group resolves — by completing, or by its + /// members disagreeing so its key is spent; or it is one [`Op::is_admissible`] + /// refuses, which no later arrival changes. A caller that /// must tell "not yet" from "never" — an ingest seam deciding what to log, dedup /// and acknowledge — asks the op, since the refusal is a function of the op /// alone and never of this document's state. @@ -3624,12 +3625,22 @@ impl Document { self.apply_now(&op); progressed = true; } + // Read once per pass and handed to both group rules, so spending a key + // and committing a group cannot disagree about what the buffer holds. A + // pass that spends one takes the next pass to commit, which is the drain's + // ordinary fixpoint rather than a rule: the map a spend leaves behind + // describes tags the buffer no longer carries, and re-reading it is + // cheaper to justify than reasoning about which of its entries survived. + let groups = self.tx_buckets(); + if self.resolve_disagreeing_tx(&groups) { + progressed = true; + } // One complete atomic transaction: apply every member in seq order, so a // member that targets a container an earlier member creates reaches it on // the first pass. Order is a shortcut, not the mechanism — a member that // is not ready is re-buffered untagged and lands on the drain's fixpoint // either way. - if let Some(mut members) = self.take_complete_tx() { + else if let Some(mut members) = self.take_complete_tx(groups) { // The bucket's key is spent by this commit — its members are leaving // the buffer and the count they met cannot be met a second time — so // record it before anything else can arrive under it. @@ -3751,8 +3762,10 @@ impl Document { /// not monotone — a container is installed, displaced, and re-installed as /// ops arrive — so a group-wide resolution gate would make commit a window /// arrival order decides, and the same ops would fold to different states. - fn take_complete_tx(&mut self) -> Option> { - let groups = self.tx_buckets(); + fn take_complete_tx( + &mut self, + groups: HashMap<(ClientId, TxId), Vec>, + ) -> Option> { // Lowest buffer position wins when more than one group is complete, so the // commit order is the buffer's, not the hash map's. Draining to a fixpoint // after every fold keeps a replica's own buffer down to at most one complete @@ -3805,6 +3818,34 @@ impl Document { } } + /// Spend the key of every bucket whose members disagree on the group's size, + /// releasing what it holds — a group no arrival can complete resolves where a + /// commit resolves. + /// + /// Holding a disagreement instead would be a decision the arrival order makes: + /// where a *subset* of a rewritten group agrees, it reaches its own declared + /// size before the dissenting member lands, commits, and leaves the dissenter a + /// stray of a resolved key — so whether the bucket ever holds the disagreement + /// at all depends on which members came first, and one op set folds to two + /// states. Resolving it makes both halves of that arrival space read the same + /// set: the members held are released untagged, and the key is spent against + /// the members still to come (ARCHITECTURE §Opt-In: Atomic). + fn resolve_disagreeing_tx(&mut self, groups: &HashMap<(ClientId, TxId), Vec>) -> bool { + let split: Vec<(ClientId, TxId)> = groups + .iter() + .filter(|(_, idxs)| self.tx_declared_count(idxs).is_none()) + .map(|(key, _)| *key) + .collect(); + let mut spent = false; + for key in split { + spent |= self.resolved_tx.insert(key); + } + if spent { + self.untag_resolved(); + } + spent + } + /// The buffer positions of every held transaction member, bucketed by the /// `(author, group id)` key a group is identified by. fn tx_buckets(&self) -> HashMap<(ClientId, TxId), Vec> { @@ -3818,8 +3859,9 @@ impl Document { } /// The size the members held at `idxs` declare, or `None` if they do not all - /// declare the same one — a bucket without unanimity names no group, so it is - /// never complete. + /// declare the same one — a bucket without unanimity names no group, so it can + /// never complete and [`resolve_disagreeing_tx`](Self::resolve_disagreeing_tx) + /// spends its key instead. /// /// The size is the group's, not that of whichever member the buffer holds /// first: read off one member, a rewritten envelope chooses when the group @@ -6116,7 +6158,7 @@ impl Op { /// than splitting over it. That is what makes rejecting it safe, and it is the /// line an ingest seam needs: an admissible op that does not apply yet is /// *waiting* — held until a create makes its target reachable or its - /// transaction group completes — so it is state worth logging, fanning out and + /// transaction group resolves — so it is state worth logging, fanning out and /// acknowledging, while an inadmissible one is worth none of those. Admissible /// says only that no rule forbids the op outright; it does not promise the op /// applies, and an already-applied op stays admissible. diff --git a/crates/core/tests/convergence.rs b/crates/core/tests/convergence.rs index ca872d36..a53839aa 100644 --- a/crates/core/tests/convergence.rs +++ b/crates/core/tests/convergence.rs @@ -665,13 +665,13 @@ fn atomic_groups_do_not_change_what_ops_merge_to() { /// sizes no arrival meets, and sizes a group's members do not share — and every /// replica must fold it to one state on the ops alone, before any eviction. /// -/// Whether a *bucket* looks unreachable is a property of which of its members -/// have landed, so nothing may be decided from it: a replica that released a -/// bucket the moment its members disagreed would release a different set from one -/// served the same ops in another order, and the member that arrived after the -/// release would be held against a size its bucket had already given up on. Only -/// a judgement on a member's own declared size is order-free. Eviction then has to -/// leave them converged too, on whatever each was still holding. +/// Whether a *bucket* ever holds a disagreement is a property of which of its +/// members have landed — a rewritten subset that agrees among itself reaches its +/// own declared size and commits before the dissenter arrives — so holding one is +/// a decision the delivery order makes. A disagreement therefore resolves its key +/// where a commit resolves it: every order releases the same members, and every +/// member still to come is a stray of a spent key. Eviction then has to leave them +/// converged too, on whatever each was still holding. #[test] fn rewritten_group_counts_converge_under_every_order() { let seeds = if cfg!(miri) { 1 } else { 60 }; @@ -774,9 +774,65 @@ fn rewritten_group_counts_converge_under_every_order() { "seed {seed}: shuffle {round} diverged after eviction" ); } + + // A rewrite that leaves a *unanimous remainder whose size is exactly what it + // now declares* is the only shape a subset can complete around: every member + // but the last is rewritten to declare one fewer, so the rewritten members + // are unanimous among themselves and reach their own size before the + // dissenter lands. Delivered the other way the bucket disagrees from the + // start. Both halves of that arrival space have to read the same, on the ops + // alone — and, since a disagreement now spends its key, eviction has to find + // nothing left to do. + let minority = rewrite_all_but_last(&pool); + assert!( + minority.iter().zip(&pool).any(|(f, o)| f.tx != o.tx), + "seed {seed}: the relay rewrote no minority" + ); + let held = converge_shuffled(&minority, 100, 1, &mut Rng::new(seed)); + let reference = converge_evicting(&minority, 100, &mut Rng::new(seed)); + assert_eq!( + held, reference, + "seed {seed}: a minority rewrite left something for eviction" + ); + for round in 0..shuffles { + assert_eq!( + converge_shuffled(&minority, 180 + round as u8, 1, &mut rng), + held, + "seed {seed}: shuffle {round} diverged on a minority rewrite" + ); + assert_eq!( + converge_evicting(&minority, 180 + round as u8, &mut rng), + reference, + "seed {seed}: shuffle {round} diverged after eviction" + ); + } } } +/// `pool` with every member of each multi-member group except the last one it +/// holds rewritten to declare one fewer — leaving a unanimous majority that can +/// reach its own declared size and a single dissenting member that cannot. +fn rewrite_all_but_last(pool: &[Op]) -> Vec { + let mut out = pool.to_vec(); + let mut last: std::collections::HashMap<(crdtsync_core::ClientId, crdtsync_core::TxId), usize> = + std::collections::HashMap::new(); + for (i, op) in out.iter().enumerate() { + if let Some(tx) = op.tx { + last.insert((op.id.client, tx.id), i); + } + } + for (i, op) in out.iter_mut().enumerate() { + let Some(tx) = op.tx else { continue }; + if tx.count > 1 && last[&(op.id.client, tx.id)] != i { + op.tx = Some(crdtsync_core::Tx { + count: tx.count - 1, + ..tx + }); + } + } + out +} + /// Deliver `ops` shuffled, then evict whatever the replica is still holding — /// the policy every deployment runs, and what makes a group nobody will complete /// converge rather than sit. diff --git a/crates/core/tests/transaction.rs b/crates/core/tests/transaction.rs index ee9fbdaf..3fb6dc62 100644 --- a/crates/core/tests/transaction.rs +++ b/crates/core/tests/transaction.rs @@ -854,9 +854,10 @@ fn a_group_built_over_a_reused_sequence_does_not_collide_with_the_first() { // releases a group, and `encode_state` carries the buffer, so the next replica // starts holding them too. The bound at the decode boundary keeps the // unreachable sizes off the wire and a member carrying one past `apply` is -// untagged on its own account; unanimity across a bucket keeps a group's size -// from being whichever member landed first; and eviction is the way out for a -// group that is merely never completed, whatever left it that way. +// untagged on its own account; a bucket whose members disagree names no size at +// all, so rather than hold it the receiver spends its key on the spot; and +// eviction is the way out for a group that is merely never completed — a +// unanimous one nobody finishes. /// `op` re-tagged as a member of group `id` declaring `count` members — the /// envelope rewrite a hostile peer or a relay can perform on a member in flight. @@ -1039,102 +1040,90 @@ fn a_member_declaring_a_size_outside_the_cap_is_refused_on_arrival() { } #[test] -fn a_rewritten_first_member_count_does_not_commit_the_group_at_the_wrong_size() { - // Three members; the one that lands first is rewritten to declare two. Read - // off that member, the size is met by the pair — committing a group two - // thirds of the way through, and leaving the third holding a size its bucket - // has already spent. The bucket has to agree on its size instead. - let mut a = doc(1); - let ops = a.atomic_transact(|tx| { - tx.register(b"x", Scalar::Int(1)); - tx.register(b"y", Scalar::Int(2)); - tx.register(b"z", Scalar::Int(3)); - }); +fn a_disagreeing_bucket_releases_what_it_holds_the_moment_it_disagrees() { + // Three members; the one that lands first is rewritten to declare two. The + // size is still not read off that member — the bucket names no size at all — + // but naming none is not a reason to wait: no arrival can complete a bucket + // whose members disagree, so the key is spent the moment the second member + // contradicts the first, and the pair is released untagged then and there + // rather than at eviction. + let (a, ops) = triple(); assert_eq!(ops.len(), 3); let id = tx_id(&ops[0]); let mut b = doc(2); - assert!(!b.apply(&retagged(&ops[0], id, 2))); - assert!(!b.apply(&ops[1])); - assert_eq!(reg(&b, b"x"), None, "the pair is not the group"); - assert_eq!(reg(&b, b"y"), None, "the pair is not the group"); + assert!( + !b.apply(&retagged(&ops[0], id, 2)), + "one member is a bucket" + ); + assert_eq!(reg(&b, b"x"), None, "a lone member still waits"); + assert!(b.apply(&ops[1]), "the contradiction releases the bucket"); + assert_eq!(reg(&b, b"x"), reg(&a, b"x")); + assert_eq!(reg(&b, b"y"), reg(&a, b"y")); - // The third arrives; the bucket still names no size, so nothing commits — - // and eviction is what releases all three. - assert!(!b.apply(&ops[2])); - assert_eq!(reg(&b, b"z"), None); - assert_eq!(b.evict_partial_transactions(), 1); - for key in [&b"x"[..], b"y", b"z"] { + // The key is spent, so the member still to come is a stray of a resolved + // group and merges standalone — leaving eviction nothing to evict. + assert!(b.apply(&ops[2])); + assert_eq!(reg(&b, b"z"), reg(&a, b"z")); + assert_eq!(b.evict_partial_transactions(), 0); +} + +#[test] +fn a_bucket_reads_no_size_off_the_member_it_happens_to_hold_first() { + // Spending a disagreeing bucket's key does not make it safe to read the size + // off one member: which member the buffer holds first is the delivery order's, + // so a size read there decides the group's fate by arrival order and not by + // what its members say. A two-member group with one member rewritten *larger* + // separates the two rules where a smaller rewrite cannot — read off the + // rewritten member the bucket is short of its size and waits, read off the + // honest one it is at its size and commits, so the two orders part company. + // The bucket names no size at all, both orders spend the key, and both land + // the pair. + let (a, ops) = pair(); + let id = tx_id(&ops[0]); + let forged = [retagged(&ops[0], id, 3), ops[1].clone()]; + + let b = one_state_in_every_order(&forged); + for key in [&b"x"[..], b"y"] { assert_eq!(reg(&b, key), reg(&a, key), "member {key:?} did not land"); } } #[test] -fn a_bucket_whose_members_disagree_on_the_size_never_completes() { +fn a_bucket_whose_members_disagree_on_the_size_spends_its_key() { // The disagreement arriving last is the other half: the bucket is at its - // declared size in members, and still names no group. - let mut a = doc(1); - let ops = a.atomic_transact(|tx| { - tx.register(b"x", Scalar::Int(1)); - tx.register(b"y", Scalar::Int(2)); - tx.register(b"z", Scalar::Int(3)); - }); + // declared size in members and names no group, so the arrival that + // contradicts is also the one that releases all three. + let (a, ops) = triple(); let id = tx_id(&ops[0]); let mut b = doc(2); - b.apply(&ops[0]); - b.apply(&ops[1]); - assert!(!b.apply(&retagged(&ops[2], id, 2))); - for key in [&b"x"[..], b"y", b"z"] { - assert_eq!( - reg(&b, key), - None, - "member {key:?} committed a size-3 bucket" - ); + assert!(!b.apply(&ops[0])); + assert!(!b.apply(&ops[1])); + for key in [&b"x"[..], b"y"] { + assert_eq!(reg(&b, key), None, "a unanimous bucket waits, whole"); } - assert_eq!(b.evict_partial_transactions(), 1); + assert!(b.apply(&retagged(&ops[2], id, 2))); for key in [&b"x"[..], b"y", b"z"] { - assert_eq!(reg(&b, key), reg(&a, key)); + assert_eq!(reg(&b, key), reg(&a, key), "member {key:?} did not land"); } + assert_eq!(b.evict_partial_transactions(), 0); } #[test] -fn a_rewritten_count_holds_the_same_set_whatever_order_it_arrives_in() { - // Whether a bucket looks unreachable depends on which of its members have - // landed, so nothing may be decided from that: two replicas served the same - // ops in different orders would then release different sets and diverge. - let mut a = doc(1); - let ops = a.atomic_transact(|tx| { - tx.register(b"x", Scalar::Int(1)); - tx.register(b"y", Scalar::Int(2)); - tx.register(b"z", Scalar::Int(3)); - }); +fn a_rewritten_count_lands_the_same_set_whatever_order_it_arrives_in() { + // Whether a bucket ever *holds* the disagreement depends on which of its + // members have landed, so holding is a decision the arrival order makes: two + // replicas served the same ops in different orders would release different + // sets and diverge. Spending the key on the disagreement is what removes the + // choice — every order lands the whole set, and folds to the same bytes. + let (a, ops) = triple(); let id = tx_id(&ops[0]); let forged = [ops[0].clone(), ops[1].clone(), retagged(&ops[2], id, 2)]; - for order in [ - [0, 1, 2], - [0, 2, 1], - [2, 0, 1], - [2, 1, 0], - [1, 2, 0], - [1, 0, 2], - ] { - let mut b = doc(2); - for i in order { - b.apply(&forged[i]); - } - for key in [&b"x"[..], b"y", b"z"] { - assert_eq!(reg(&b, key), None, "order {order:?} released {key:?} early"); - } - b.evict_partial_transactions(); - for key in [&b"x"[..], b"y", b"z"] { - assert_eq!( - reg(&b, key), - reg(&a, key), - "order {order:?} stranded {key:?}" - ); - } + let b = one_state_in_every_order(&forged); + for key in [&b"x"[..], b"y", b"z"] { + assert_eq!(reg(&b, key), reg(&a, key), "member {key:?} did not land"); } } @@ -1164,6 +1153,30 @@ fn a_group_rewritten_smaller_on_every_member_folds_one_state_in_every_order() { } } +#[test] +fn a_group_rewritten_smaller_on_a_minority_of_members_folds_one_state_in_every_order() { + // Two of three members rewritten to declare two, the third left declaring + // three: every member legal on its own terms, and the bucket disagreeing only + // once the dissenter is in it. Delivered `x,y,z` or `y,x,z` the unanimous pair + // reaches its own declared size first and commits, so the dissenter is a stray + // of a resolved key and all three land; delivered any of the other four ways + // the dissenter is in the bucket from the start. A bucket that disagrees is a + // group no honest arrival completes, so it resolves where a commit does — and + // the two halves of the arrival space read the same. + let (a, ops) = triple(); + let id = tx_id(&ops[0]); + let forged = [ + retagged(&ops[0], id, 2), + retagged(&ops[1], id, 2), + ops[2].clone(), + ]; + + let b = one_state_in_every_order(&forged); + for key in [&b"x"[..], b"y", b"z"] { + assert_eq!(reg(&b, key), reg(&a, key), "member {key:?} did not land"); + } +} + #[test] fn an_unrelated_op_retagged_into_a_group_folds_one_state_in_every_order() { // The same shape built the other way round: an honest two-member group plus a @@ -1300,6 +1313,73 @@ fn two_envelopes_naming_different_groups_spend_both_keys() { /// record closes: it needs the two copies to name different groups, it is /// order-dependent on `main` before the record exists, and the record narrows it /// from a split in what a replica reads to a split in which keys it has spent. +/// The **third** door into that same residue, and the one this unit opened. A +/// second envelope is answerable only where the buffer is still *holding* the +/// first copy, and a disagreement now takes a member out of the buffer — so +/// whether the conflict rule fires at all became the delivery order's. Delivered +/// with the contradiction ahead of the second envelope, the copy it would have +/// disagreed with is already released and the second group's key is never spent; +/// delivered the other way round both keys are spent. The members converge on +/// eviction, and which keys each replica has spent does not — which is exactly +/// C46's state, reached without a group ever completing. Closing it needs the same +/// per-op-id evidence C46 is filed for: which group a member was released *under*, +/// not merely that its key resolved. Pinned here so the door is a measured fact +/// rather than a claim. +#[test] +fn a_copy_released_by_a_disagreement_leaves_the_second_envelope_nothing_to_contradict() { + let (mut a, group) = pair(); + let t1 = tx_id(&group[0]); + let t2 = crdtsync_core::TxId(0x5eed_5eed_5eed_5eed); + // `x` under its own group, `y` contradicting it, and `x` again under a second + // group — every envelope admissible, none of them malformed. + let ops = [ + retagged(&group[0], t1, 2), + retagged(&group[1], t1, 3), + retagged(&group[0], t2, 2), + ]; + + // The contradiction first: it releases `x`, so the second envelope is a plain + // duplicate of an applied id and `t2` is never spent. + let mut released = doc(9); + for i in [0, 1, 2] { + released.apply(&ops[i]); + } + // The second envelope first: the buffer is still holding `x` under `t1`, the + // conflict rule fires, and both keys are spent. + let mut contradicted = doc(9); + for i in [0, 2, 1] { + contradicted.apply(&ops[i]); + } + assert_ne!( + released.encode_state(), + contradicted.encode_state(), + "the two orders no longer differ — the third door has been closed, so this \ + test and C46's filing need revisiting rather than deleting" + ); + // Both readings agree on what a reader sees: the split is in the spent-key + // record, and every member has landed either way. + for key in [&b"x"[..], b"y"] { + assert_eq!(reg(&released, key), reg(&a, key)); + assert_eq!(reg(&contradicted, key), reg(&a, key)); + } + + // And eviction leaves every order reading the same document, which is what + // keeps this a residue rather than a divergence a deployment sees. + let mut evicted: Option>> = None; + for order in orderings(ops.len()) { + let mut d = doc(9); + for &i in &order { + d.apply(&ops[i]); + } + d.evict_partial_transactions(); + let read: Vec> = [&b"x"[..], b"y"].iter().map(|k| reg(&d, k)).collect(); + match &evicted { + None => evicted = Some(read), + Some(first) => assert_eq!(&read, first, "order {order:?} read differently"), + } + } +} + /// Filed as its own unit (KANBAN C46). #[test] fn a_copy_absorbed_by_another_groups_bucket_strands_its_own_group_until_eviction() { diff --git a/crates/wasm/src/lib.rs b/crates/wasm/src/lib.rs index 5f57afa2..a9426210 100644 --- a/crates/wasm/src/lib.rs +++ b/crates/wasm/src/lib.rs @@ -682,7 +682,7 @@ impl WasmDocument { /// /// **The two zeros mean opposite things.** An op that did not apply may be a /// duplicate, or be *waiting* — buffered until a create makes its target - /// reachable or its transaction group completes, which a later arrival does, + /// reachable or its transaction group resolves, which a later arrival does, /// including one later in this same batch, which `applied` does not count — /// while a refused op is a bug in whoever wrote it, and no arrival lifts it. /// The refusal is the stamp conditions [`Op::is_admissible`] names, the codec diff --git a/sdks/go/crdtsync/handle.go b/sdks/go/crdtsync/handle.go index 58a49b27..8f400a5f 100644 --- a/sdks/go/crdtsync/handle.go +++ b/sdks/go/crdtsync/handle.go @@ -391,7 +391,7 @@ func (d *Doc) SetSchema(schema []byte) bool { // The two counts separate an op that did not apply yet from one that never will. // applied is what the fold took as the ops arrived; one it did not take may be a // duplicate, or be waiting — buffered until a create makes its target reachable or -// its transaction group completes, which a later update does, including one later +// its transaction group resolves, which a later update does, including one later // in this same batch (released that way, it is not counted). refused is what no // replica will ever hold, which is a bug in whoever wrote it: a peer reached // offline, directly, or over a byte pipe the app carries itself has no server diff --git a/sdks/js/src/doc.ts b/sdks/js/src/doc.ts index 90df79bd..efecaf2e 100644 --- a/sdks/js/src/doc.ts +++ b/sdks/js/src/doc.ts @@ -145,7 +145,7 @@ export class Doc { * The outcome separates an op that did not apply *yet* from one that never * will. `applied` counts what the fold took as the ops arrived; one it did not * take may be a duplicate, or be waiting — buffered until a create makes its - * target reachable or its transaction group completes, which a later update + * target reachable or its transaction group resolves, which a later update * does, including one later in this same batch (released that way, it is not * counted). `refused` counts what no replica will ever hold, which is a bug in * whoever wrote it: a peer reached offline, directly, or over a byte pipe the diff --git a/sdks/python/crdtsync/__init__.py b/sdks/python/crdtsync/__init__.py index 71c2d7a9..3c371f2d 100644 --- a/sdks/python/crdtsync/__init__.py +++ b/sdks/python/crdtsync/__init__.py @@ -3168,7 +3168,7 @@ def apply_update(self, ops: bytes) -> ApplyOutcome: The outcome separates an op that did not apply *yet* from one that never will. ``applied`` counts what the fold took as the ops arrived; one it did not take may be a duplicate, or be waiting — buffered until a create makes - its target reachable or its transaction group completes, which a later + its target reachable or its transaction group resolves, which a later update does, including one later in this same batch (released that way, it is not counted). ``refused`` counts what no replica will ever hold, which is a bug in whoever wrote it: a peer reached offline, directly, or over a