Skip to content

test(e2e): release the stale block once its ingress window has reopened - #272

Merged
spalladino merged 2 commits into
mainfrom
spl/fix-stale-block-ingress-window
Sep 18, 2026
Merged

spalladino merged 2 commits into
mainfrom
spl/fix-stale-block-ingress-window

Conversation

@spalladino

Copy link
Copy Markdown
Collaborator

Context

multi-node/block-production/cross_chain_messages.parallel.test.ts → builds multiple blocks per slot with L1 to L2 messages has been failing since its reorg phase landed with #243. It surfaced on #221, which is unrelated to it — that PR touches no end-to-end code and is simply the first one since #243 whose diff invalidated the e2e test cache. #243's own final head merged with ci cancelled, so this phase was never validated.

It is not a flake, and cannot be handled as one. src/multi-node/.*\.test\.ts already matches an owned entry in .test_patterns.yml, so ci3/run_test_cmd already gave it its one automatic retry; it failed both times, with different proposer/validator index draws and the same numbers to within 10ms. flake_error_threshold is read only by the dashboard, never by run_test_cmd, so no entry would change the outcome.

Root cause

The failure is TimeoutError: Timeout awaiting validator N compares the stale block's Inbox prefix.

reorgWithReplacement re-mines the replaced suffix from the fork point, which leaves the L1 head stamped behind where it was — one 12s L1 slot, in both failing runs. The e2e clock follows L1, so protocol time walks backwards with it.

The held block is signed at the very start of its build frame (secondsIntoBuildFrame: 0.652, indexWithinCheckpoint: 0), and for its slot that instant is getCheckpointProposalReceiveStart. The release therefore landed ~11.73s before the window it had to be inside. proposal_validator.ts gates ingress on both ends of that window, so every peer dropped the proposal as too early:

Penalizing peer for invalid slot number 20
  {"nowMs":1789768393272,"windowStartSeconds":1789768405,"windowDeadlineSeconds":1789768471.5}

The named validator never ran the Inbox-prefix comparison at all, so the retryUntil polling for it timed out after 144s. The product side behaved correctly throughout — every node detected the rolling-hash divergence, rolled back 5→4 messages, and the proposer pruned its stale block and abandoned the slot with inbox_prefix_reorged.

The gate could not have caught this. remainingIngressBudgetMs measured only the distance to the window's upper bound, so it reported a healthy 78233 ms computed off the already-rewound clock, and expect(ingressBudgetMs).toBeGreaterThan(0) passed while the release was in fact too early.

Change

  • HoldContext.remainingIngressBudgetMs becomes ingressWindow(), returning { opensInMs, closesInMs } from a single reading of the clock — a proposal released now is acceptable exactly when opensInMs <= 0 < closesInMs, and the two ends cannot be sampled either side of a moving clock. The schedule already exposed getProposalReceiveStartSeconds(); the gate just never surfaced it.
  • The reorg case waits for the window to reopen before releasing. Interval mining stamps the next L1 blocks forward again, so the wait resolves on its own and lands back at roughly the same point in the build frame it was at before the reorg — which is why canStartAnotherBlock() still holds. The bound is wall-clock (retryUntil measures with a real timer), so a chain that stopped advancing fails at the wait rather than downstream at the assertion it silently breaks.
  • The release-time checks now all sample one now.

The p2p lower bound is correct product behaviour and is untouched; this is a harness defect only.

Testing

Red/green on the gate: the new reports how long the ingress window is still to open case fails with ctx.ingressWindow is not a function before the change and passes after — 25/25 in checkpoint_proposal_job_test_gate.test.ts. Full yarn-project bootstrap, build, format-check and lint are green.

The e2e case itself was not rerun locally; CI is the verification for it.

The cross-chain reorg case released a signed block the instant it had
replaced the L1 suffix. Replacing the suffix re-mines it from the fork
point, which leaves the L1 head stamped behind where it was, and the e2e
clock follows L1, so protocol time walked backwards with it. The held
block is signed at the very start of its build frame, which is the same
instant its slot's proposal receive window opens, so the release landed
~12s before the window it needed to be inside. Every peer dropped the
proposal at p2p ingress for a slot that had not opened yet, the named
validator never compared its Inbox prefix, and the test timed out
awaiting a comparison that could no longer happen.

The gate could not have caught this: `remainingIngressBudgetMs` measured
only the distance to the window's upper bound, so it reported a healthy
budget off the already-rewound clock. It becomes `ingressWindow()`,
which returns both ends from a single reading of the clock, and the
reorg case waits for the window to reopen before releasing. Interval
mining stamps the next L1 blocks forward again, so the wait resolves on
its own and lands back at the same point in the build frame; its bound
is wall-clock, so a chain that stopped advancing fails at the wait
rather than at the assertion it silently breaks.

The e2e itself was not rerun locally; CI is the verification for it.
@greptile-apps

greptile-apps Bot commented Sep 18, 2026

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

The PR appears safe to merge; no actionable correctness, security, or repository-rule violations were identified.

Summary

This PR corrects the stale-block reorg test harness so it waits until the checkpoint proposal receive window has reopened before releasing the held proposal.

  • Replaces a deadline-only ingress budget with a two-sided ingress-window calculation based on one clock sample.
  • Adds unit coverage for receive-window opening calculations.
  • Makes the multi-node reorg scenario wait for the lower ingress bound and verify both bounds plus remaining block-building capacity before release.
Diagram
sequenceDiagram
  participant Test as Reorg test
  participant L1 as Interval-mined L1
  participant Gate as Proposal hold gate
  participant Peer as Validator ingress

  Gate->>Gate: Hold signed stale block
  Test->>L1: Replace canonical suffix
  L1-->>Gate: Protocol clock moves backward
  loop Until receive window reopens
    Test->>Gate: ingressWindow()
    Gate-->>Test: opensInMs, closesInMs
    L1->>L1: Mine next interval block
  end
  Test->>Gate: Verify window and next subslot
  Test->>Gate: Release stale block
  Gate->>Peer: Gossip signed proposal
  Peer->>Peer: Validate slot and Inbox prefix
  Peer-->>Test: Report Inbox-prefix mismatch
Loading

Reviews (1) · Last reviewed commit: "test(e2e): release the stale block once ..."

Review corrections, all documentation or diagnostics; the timing logic is
unchanged.

The wait's comment claimed the clock only climbs back because interval
mining stamps the next L1 blocks forward, and that a chain which stopped
advancing would fail at the wait. Both are wrong: the shared provider is
an offset on the real clock, so it resumes advancing on its own and the
wait would succeed from wall-clock passage alone. The bound is there so a
clock that never arrives names the window rather than failing downstream.
The comment also stated the rewind as a certainty; it has now been
measured at both nothing and a full L1 slot for the same reorg depth, so
it says that instead.

`ingressWindow`'s contract said a proposal is acceptable *exactly* when
`opensInMs <= 0 < closesInMs`. Peers judge at receive time and widen both
ends by their clock-disparity tolerance, so that is a sufficient
condition, not the exact one. The unit case now pins `closesInMs` at both
ends, including the inclusive deadline boundary that makes the difference.

The wait also logs the window and the reorg depth before it starts, so a
timeout shows how far the clock had actually moved, and `HoldContext`
grows a `now()` rather than having callers reach through to the event's
schedule for the shared sample.
@spalladino
spalladino merged commit 42c3f18 into main Sep 18, 2026
6 checks passed
@spalladino
spalladino deleted the spl/fix-stale-block-ingress-window branch September 18, 2026 22:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant