Repository navigation
Conversation
This was referenced Sep 20, 2026
EnRaiha
added this pull request to stack #358
September 20, 2026 11:14
EnRaiha
removed this pull request from stack #358
September 21, 2026 02:36
…a failure
A full per-tenant bridge queue returned Error::Dispatch, which the Calvin
scheduler logged at ERROR and completed as a failed transaction, releasing its
locks. During the startup rebuild the epoch is re-driven, so a full queue spun:
one ERROR and one failed transaction per attempt, the queue never drained, CPU
pinned, and pgwire starved.
- add Error::DispatchCapacityBusy { tenant_id, inflight, cap }: the request was
not enqueued and nothing was applied, so it is retryable by type
- the bridge dispatcher returns it for the per-tenant cap
- the scheduler demotes the log to debug, counts it (dispatch_busy_count), and
paces the catch-up drain with a bounded jittered backoff (busy_backoff, 5 ms
doubling to a 250 ms cap) instead of spinning
- classify the variant like a dispatch failure and map it to BUSY for clients
Tests: full_queue_is_a_retryable_capacity_condition (new; failed before the
fix), full_queue_returns_error (unchanged), busy_backoff_is_bounded_and_jittered.
Issue 352.
The first change classified a full per-tenant queue as retryable and paced the epoch dispatch, but four other sites still logged ERROR and completed the work as failed: the CalvinResolve dispatch, the commit-resolution dispatch, the write-version record dispatch, and the second (active) epoch dispatch. On a large WAL replay those sites kept the catch-up loop busy and the readiness gate never opened. - note_dispatch_busy(&mut self, ..): counts the event and arms the bounded drain backoff; used by the epoch, CalvinResolve and commit-resolution sites - note_dispatch_busy_shared(&self, ..): counts only (no arming) for the one-way write-version record dispatch, which holds only a shared borrow - each site demotes the log to debug on the retryable class Issue 352.
…startup The readiness gate failed startup when the metadata group applied no entry for 30 s. With a large apply backlog the group needs minutes, so a slow boot became a restart loop (2026-09-20 incident; issue 352). Raise the bound to 300 s and keep failing a group that never applies. Follow-up: expose the bound as configuration.
The configurable [server] data_group_recovery_timeout_ms (600 s in production) was replaced by a hard-coded 60 s. On a backlogged recovery the data groups need longer: group 2 reached 261 of 262 committed entries inside 60 s and startup aborted into a restart loop (2026-09-20 deploy attempts of the HEAD+fixes builds). Restore the previous operator value as the constant; a follow-up should expose it as configuration again.
The house preflight requires the issue number to live in the PR, not in the code, and this branch touches the file.
The Calvin determinism gate forbids unmarked Instant::now() in the write path. The backoff deadline and the drain gate read the wall clock for pacing only — scheduler observability, never Calvin WAL data. Comment-only change.
…into modules The scheduler and dispatch files grew past the 500-line house limit. Move the capacity-busy accounting and backoff into core/busy.rs and the active dependent-read dispatch into core/active_dispatch.rs, and route the static dispatch busy path through the shared helper instead of duplicating it.
Each removed doc line repeated the variant's #[error] text, so the message is now the single source for what the variant says. The enum had also grown past its base line count; this returns it below the unchanged-file bound.
EnRaiha
force-pushed
the
fix/calvin-backpressure
branch
from
September 21, 2026 14:00
7700837 to
cbae048
Compare
Member
|
Closing. On a capacity-busy dispatch, |
5 of 8 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
A full per-tenant bridge queue returned
Error::Dispatch; the Calvin schedulerlogged it at ERROR and completed the transaction as failed, releasing its locks.
During the startup rebuild the epoch is re-driven, so a full queue became a spin:
33,435
dispatch failedERRORs in 60 s, CPU pinned, pgwire starved. Itreproduced on the production server with a build derived from main: the node came
up and never served.
What
df9c3b818Error::DispatchCapacityBusy { tenant_id, inflight, cap }— retryable by type. The dispatcher returns it for the per-tenant cap; the epoch dispatch demotes the log to debug, counts it (dispatch_busy_count), and paces the catch-up drain withbusy_backoff(5 ms doubling to a 250 ms cap, deterministic jitter)ae60781c4note_dispatch_busy/note_dispatch_busy_shared)c5a61d025RAFT_READY_STALL_TIMEOUT30 s → 300 s: with a large apply backlog the metadata group needs minutes, and 30 s turned a slow boot into a restart loopeb4f0b295DATA_GROUP_RECOVERY_TIMEOUT60 s → 600 s: restores the value the production config carried before the bound became a hard-coded constantdbfccd46cThe gateway maps the new variant to BUSY, so a client still sees a retryable
overload rather than an internal failure, and the tracker entry is cancelled on
the deferred path.
How to test
full_queue_is_a_retryable_capacity_condition— new; fails on main (the errorwas
Error::Dispatch), passes with the fixfull_queue_returns_error— unchanged: a full queue is still an errorbusy_backoff_is_bounded_and_jittered— bound ≤ 257 ms, jitter varies withvShard and epoch
Validation
storm-class ERRORs across the monitored window,
SELECT 1returns, theknowledge-graph store is intact (70,978 rows).
storm and no service.
the previous hard-coded bounds aborted the start with
metadata group applied no entry for 30sanddata raft group recovery timeout after 60s.Notes
this branch splits the busy accounting into
core/busy.rsand the activedispatch into
core/active_dispatch.rs(dispatch.rs543 → 387,scheduler.rs508 → 459 non-test lines), and drops variant docs inerror/types.rsthat restated their display messages, returning it below itsbase count (601 → 586). The enum itself cannot go under 500 without changing
the public API, so C5 treats it as the pre-existing over-limit file it is:
no growth.
Fixes #352
Fixes #353
Tradeoffs
Tradeoff:
error/types.rscannot reach 500 non-test lines while its enum stays one type, so C5 treats it as a pre-existing over-limit file that must not grow; the refactor keeps it below its base instead.