Skip to content

feat(raft): serve a linearizable read from the quorum-contact lease - #346

Closed
EnRaiha wants to merge 1 commit into
mainfrom
fix/p2-leader-lease
Closed

EnRaiha wants to merge 1 commit into
mainfrom
fix/p2-leader-lease

Conversation

@EnRaiha

@EnRaiha EnRaiha commented Sep 19, 2026

Copy link
Copy Markdown
Contributor

Why

A linearizable read paid a ReadIndex quorum round even when the leader had just proven a majority was answering it. The check-quorum window answers the same question: a majority acknowledged this leader at last_quorum_contact, and any successor needs a majority too, so no other node can have moved the log past the commit index before an election timeout elapses.

Part of #165 (P2 — cluster consensus safety).

What

  • RaftNode::leader_lease_index(now) (nodedb-raft/src/node/quorum_contact.rs): the commit index while the node is Leader and now − last_quorum_contact < election_timeout_max. Both instants are this node's monotonic clock, so no skew bound applies; the window closes at the same threshold quorum_contact_lost demotes at, so a lease-answered read never outlives leader state.
  • MultiRaft::leader_lease_index(group_id) mirrors the existing per-group accessors.
  • confirm_read_index (read_index_wait.rs) serves locally when the lease is live and falls back to the probe otherwise. Refusal semantics unchanged (NotLeader / Timeout paths).

Validation

  • cargo test -p nodedb-raft --lib quorum_contact — 13 passed (lease live after election; lapses at the step-down threshold; no lease off the leader path).
  • cargo test -p nodedb-cluster --lib read_index — 4 passed.
  • cargo check -p nodedb-cluster --all-targets — clean; repository preflight passes.

Notes

  • Production beneficiary: MultiRaftReadGate::confirm_leader (the pgwire pre-dispatch read barrier).
  • No cluster-level fast-path test: a leader-hosted MultiRaft needs heavy setup; the raft-level unit is the evidence.

A linearizable read paid a ReadIndex quorum round even when the leader had
just proven a majority was answering it. The check-quorum window answers the
same question: a majority acknowledged this leader at last_quorum_contact,
and any successor needs a majority too, so no other node can have moved the
log past the commit index before an election timeout elapses.
leader_lease_index exposes that window; confirm_read_index serves locally
inside it and falls back to the probe otherwise. Both instants are monotonic
local time, and the window closes at the same threshold quorum_contact_lost
demotes at, so a lease-answered read never outlives leader state.
@EnRaiha EnRaiha added the area:cluster-raft Raft, replication, consensus safety label Sep 19, 2026
Copilot AI lite review requested due to automatic review settings September 19, 2026 07:46
@EnRaiha EnRaiha added the area:cluster-raft Raft, replication, consensus safety label Sep 19, 2026

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@farhan-syah

Copy link
Copy Markdown
Member

Closing: the lease reuses the step-down threshold as a read-safety bound, and that is unsafe.

last_quorum_contact is stamped at ack-receipt time (refresh_quorum_contact) and the lease holds for election_timeout_max. A follower starts its election timer at heartbeat arrival with a timeout in [min, max], so it can elect a new leader and commit inside the old leader's window; the old leader then serves a "linearizable" read from a stale commit index.

A correct lease is measured from the heartbeat send time and bounded by election_timeout_min with a clock-drift factor (Raft thesis §6.4.1), and needs a current-term-commit guard. Resubmit with that model and a test that drives the follower timeout inside the window.

@farhan-syah
farhan-syah deleted the fix/p2-leader-lease branch September 23, 2026 01:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area:cluster-raft Raft, replication, consensus safety

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants