Skip to content

test(sequencer): add preconf rpc and store chaos e2e suites - #52

Merged
pratikspatil024 merged 5 commits into
mainfrom
ppatil/sequencer-rpc-chaos-e2e
Sep 16, 2026
Merged

pratikspatil024 merged 5 commits into
mainfrom
ppatil/sequencer-rpc-chaos-e2e

Conversation

@pratikspatil024

@pratikspatil024 pratikspatil024 commented Sep 11, 2026 •

Copy link
Copy Markdown
Member

Summary

Stacked on #47. That suite covers the publisher side — whether validators write to the store. This adds the two things it does not: what an RPC client actually sees, and whether the store can hurt the chain. Both are ported down from John's manual campaign in #tmp_seqstore-testing-and-qa to a size CI can afford.

Additive by construction: both scripts source sequencer_test_utils.sh rather than editing it, so #47 stays conflict-free and the three suites run as separate workflow steps against one enclave — the way the sibling kurtosis leg already sequences smoke / rpc / validator / downtime.

sequencer_rpc_test.sh

check where it comes from
A preconfirmed receipt is marked preconfirmation: true and carries blockHash: null the bh: null check; that pairing is the client's only signal before canonicalisation, so a preconfirmed receipt with a block hash would let a client treat speculative state as final
Every sampled transaction reaches canonical with a real hash and status 0x1 the 452,275-receipt re-fetch, scaled to 120
eth_sendRawTransactionSync is a registered method it was absent from a deployed private RPC — -32601 breaks any client built on the sync path
eth_getCode(multicall3, "pending") succeeds 25/25 intermittent failure there becomes random estimateGas failures for every geth-based client
bor_getInvalidPreconfBlocks answers its contract and rejects a reversed range the 0x4f4df 0x4f4e1 query
…and reports where the audit's coverage of that range stops (pendingFrom), plus the 1024-height request cap 0xPolygon/bor#2388 — both response shapes accepted, see Rollout notes

Transaction hashes come from the pending block, not the load generator's stdout. That keeps the suite off polycli's log format, and it is the stronger test: the pending view is the consumer's own speculative state, so anything in it ought to be servable as a preconfirmation.

sequencer_chaos_test.sh

Five episodes through one harness — broker stop (the store service #47 never touches, and the one whose loss turned preconfirmations back into ordinary mining during manual testing), ingress latency, gateway pause (frozen with sockets open, so consumers block rather than fail fast), a seeded random fault, and an oversized-record burst.

Every episode holds the fault, asserts the chain kept building, repairs, and then requires publishing to resume on its own — a store fault that permanently de-registered a publisher would leave a healthy-looking chain with every preconfirmation silently gone.

A 1 Hz head sampler runs across all nine-ish nodes for the whole session, and the closing assertion is the invariant from the campaign: no node ever reported one hash for a height and later a different one. The store is not supposed to be able to reorganise the chain, so a rewrite here is a finding rather than reorg noise.

Fault injection reuses this repo's existing gaiadocker/iproute2 + tc netem mechanism, so CI pulls nothing new. Chaos gates PRs here because that is already the norm — the stateless suite runs test_extreme_network_latency_recovery, test_milestone_settlement_latency_resilience and test_producer_restart_recovery on every PR.

Throughput is reported, never asserted

The campaign measured 25–50% throughput cost from a store outage or latency, against a design brief that says the store cannot affect block production. A threshold here either fails today or freezes whichever number happens to be true into CI, so the numbers are printed for a human to read and the gate is on liveness and correctness.

Executed tests

  • shfmt -d clean at 3.13.1, matching the checks gate, with the repo's .editorconfig.
  • shellcheck -S warning -x clean on all three scripts. At -S info the only remaining notes are SC1091 source-path resolution, which feat: sequence-store publisher e2e suite #47's existing scripts share.
  • bash -n clean; the workflow parses as YAML with the steps in the intended order.

Run against a live devnet

Both suites ran on a local kurtosis devnet: bor and heimdall-v2 built from source, the store built from sequence-store rather than pulled (the ghcr grant this repo still needs only blocks pulling the image, so it does not block local validation). l1_backend: anvil and a source-built polycli, since there is no darwin asset for the pinned release — so the L1 layer and tooling are not bit-identical to CI.

The first run failed, and four assertions turned out not to be measuring what they claimed. All four are fixed in this PR; the detail is in the commit message. In short: the preconfirmation check was measuring its own poll latency and reported 6% on a node that was preconfirming perfectly; the oversized-record burst was sending nothing at all, because --mode s writes to storage at ~20k gas per word and every submission failed the node's tx fee cap; nothing verified that the burst had reached the chain, which is what let that hide; and count_txs_to_sink called a helper that is not in scope in the utils, so its guard reported a burst that had landed as missing.

After the fixes, both suites were re-run end to end on a fresh enclave (the earlier one had a permanently disabled publisher — see below — which fails every episode's recovery assertion):

sampled=120 preconfirmed=82 already-canonical=38 unmarked=0 never-served=0
Preconfirmation coverage 100% of 82 transactions caught while speculative
matched=120 mismatched=0 missing=0
Method is registered (rejected the invalid payload with code -32000)
All 25 pending reads succeeded
Ledger returned 0 record(s) ... Reversed range rejected
auditedThrough=0xc2 unauditedThrough=0x7f
broker stop         13 blocks/45s   4/4 publishers live
ingress latency     12 blocks/45s   4/4 publishers live
gateway pause       11 blocks/45s   4/4 publishers live
random store fault  11 blocks/45s   4/4 publishers live
large-calldata      203 of 120 mined, 0 new fences, all 4 publishers still live
No node rewrote a height it had reported (414 samples)
All sequence-store chaos tests passed (seed 4242)

already-canonical moved 34 → 38 between the two devnets while coverage stayed at 100%: the excluded bucket absorbs timing variance instead of it surfacing as a coverage number that drifts run to run. Under the original version that variance was the measurement.

The clean run needed a bor that no branch in CI built from — until now. It required Jerry's size cap (0xPolygon/bor#2400, cd7ea97f58) cherry-picked locally on top of the publisher branch, and that cherry-pick was deliberately not part of this PR: it is his commit from cffls/preconf-record-size-cap, and stacking it here would have misattributed it and widened the scope. So the claim was precisely that these suites pass against the fully-fixed pairing — store with 0xPolygon/sequence-store#12, bor with #2400.

That pairing is now what this leg builds. #2400 merged into cffls/sequence-publisher as b56ff74a4, which is the ref this workflow checks bor out at, and store #12 merged on 2026-09-10 so the pinned dev tag carries it. Nothing here changes; the local cherry-pick is simply no longer needed, and the burst episode is expected to pass rather than fail (see the finding below, now historical).

The burst episode is validated in both directions. With bor and the burst held constant and only the store image changing:

ingress image burst self-fences episode
without 0xPolygon/sequence-store#12 120KB × 240 mined 8, MESSAGE_TOO_LARGE fails
with #12 120KB × 240 mined 0 passes

That also corrected the default: 32KB lands 400 transactions, puts ~1.3MB in a block, and never reaches the path. The threshold is sharp and lowering the payload disarms the episode silently — assert_burst_landed cannot detect it, because those transactions do land, they are merely too small to coalesce past 1MB. The comment records the measurements so the next person tuning this can tell.

bor_getPreconfAuditStatus returning unauditedThrough=0x7f was worth noting at the time: 127 is exactly the last pre-Rio block, and it came back as 127 on both devnets independently, so it was a property of 0xPolygon/bor#2388 rather than an artifact of one chain's history.

That method no longer exists. Review on #2388 folded the audit's coverage into bor_getInvalidPreconfBlocks as a pendingFrom field and removed both bor_getPreconfAuditStatus and the unauditedThrough mark, on the grounds that a caller should not need a second call to interpret an empty result. So this suite no longer probes that method; test_invalid_preconf_blocks_contract reads pendingFrom instead, and the devnet figures above are kept as the record of what the earlier revision reported, not as a claim about the current one.

Two things the tooling does not support, found by checking rather than assuming:

  • polycli has no --sync-txs — not on the pinned v0.1.103 and not on main, so the manual runs used a local build. The sync path therefore gets a registration check rather than a latency measurement. Worth revisiting if that flag lands upstream.
  • --calldata requires --mode contract-call plus an address, which the burst now uses with an unallocated account — calldata to an account with no code is valid, costs only the calldata, and needs no deployment step.

One product finding, now closed out

A bor without the size cap (0xPolygon/bor#2400) talking to a store with #12 permanently disables its own publisher:

ERROR Sequencer publishing disabled: store rejected entry status=ACK_STATUS_MALFORMED

sequencer_publish_state then sits at 4 (failed) indefinitely, entries frozen, while the node keeps producing blocks — observed stuck for about an hour.

This is expected pre-cap, not a defect to chase. 0xPolygon/bor#2400 removes the trigger: rejectOversized is the only thing in the store that emits MALFORMED, and a bor that caps its records never produces one. Terminal is then the right response, for the reason #2400's own comment gives — a MALFORMED after the cap means the store's max.message.bytes is below bor's record cap, so the publisher cannot produce anything that would be accepted.

It mattered for this PR because this repo builds bor from cffls/sequence-publisher, which did not carry #2400 at the time — so the burst episode was expected to fail here, and I left it as a visible, explained failure rather than weakening an assertion to hide it. #2400 has since landed on that branch, so the trigger is gone and the episode should pass. The negative-control row in the table above stays as the record of what the pre-cap pairing did.

One residual worth an alert rather than a code change: bor hardcodes maxMessageBytes = 1 << 20 to mirror the store's topic setting, so the two are coupled by a comment rather than by configuration. A topic created or altered with a smaller limit would take every publisher terminal — correctly, but silently, since the node keeps producing and only sequencer_publish_state shows it.

Rollout notes

CI-only; no node or network impact. timeout-minutes goes 60 → 75 for the three suites (functional ~25m, most of it reaching Rio at block 128; rpc ~5m; chaos ~10m), and the diagnostics dump now triggers on any of the three failing.

Merge order: this sits on cffls/sequencer-e2e, so #47 lands first.

The ledger response shape is mid-change, and the suite accepts both. bor_getInvalidPreconfBlocks returns a bare array of {number, reason} on the bor this leg currently builds, and an object of {invalid, pendingFrom} once 0xPolygon/bor#2388 lands. test_invalid_preconf_blocks_contract detects which it got and checks each against its own contract — heights-and-reasons for the array, hex heights plus a present-but-possibly-null pendingFrom and the 1024-height request cap for the object. The object branch goes live on its own the moment the workflow's bor ref carries #2388, with no change needed here. A missing pendingFrom is a failure while a null one is the legitimate "fully audited", since telling those apart is the whole reason the field exists. No CI is pinned to unmerged branches either way.

The publisher suite covers whether validators write to the store. These cover
what a client sees and whether the store can hurt the chain, both ported down
from the manual campaign in #tmp_seqstore-testing-and-qa to a size CI affords.

Both scripts source the existing utils rather than editing them, so the
functional suite is untouched and the three run as separate workflow steps
against one enclave, the way the sibling kurtosis leg already sequences its
suites.

rpc suite: a preconfirmed receipt is marked preconfirmation:true and carries a
null blockHash, which is the client's only signal before canonicalisation;
every sampled transaction then has to reach the canonical chain with a real
hash and a success status, the 452k-receipt re-fetch scaled to 120. Plus
eth_sendRawTransactionSync being registered at all - it was missing from a
deployed private RPC, which returns -32601 and breaks any client built on the
sync path - the multicall3 pending read looped 25 times, since intermittent
failure there becomes random estimateGas failures for geth-based clients, and
bor_getInvalidPreconfBlocks answering an array of numbered, reasoned records
and rejecting a reversed range.

Hashes come from the pending block, not the load generator's output: the
pending view is the consumer's own speculative state, so anything in it should
be servable as a preconfirmation, and it keeps the suite off polycli's log
format.

chaos suite: five episodes through one harness - broker stop, ingress latency,
gateway pause, a seeded random fault, and an oversized-record burst. Each
holds the fault, asserts the chain kept building, repairs, and requires
publishing to resume on its own; a store fault that permanently de-registered
a publisher would leave a healthy-looking chain with every preconfirmation
silently gone. A 1Hz head sampler runs across the whole session and the
closing assertion is that no node ever reported one hash for a height and
later a different one, which is the property a store fault must not be able
to break.

Throughput under fault is printed, never asserted. The campaign measured
25-50% cost from store outage or latency, against a design brief that says the
store cannot affect block production. A threshold either fails today or
freezes whichever number is currently true into CI, so the numbers are output
for a human.

bor_getPreconfAuditStatus skips on -32601 rather than failing, the same shape
the suite already uses for an absent polycli, so it lands now and goes live
once the audit ships.

Two notes on what the tooling actually supports: polycli has no --sync-txs on
the pinned release (or on main), so the sync path gets a registration check
rather than a latency measurement; and --calldata needs contract-call mode
with a deployed address, so the oversized-record burst uses store mode with
--store-data-size.

Timeout goes to 75 minutes for the three suites, and the diagnostics dump now
triggers on any of them failing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@pratikspatil024
pratikspatil024 marked this pull request as ready for review September 11, 2026 11:58
…hing

Ran both suites against a local kurtosis devnet with the store built from
source. Four of the assertions were not measuring what they claimed, and two
of them I had already reported as passing.

The preconfirmation check was measuring its own latency. It read one pending
block of 143 transactions, then walked it with a request per hash; by the
tenth, the rest had canonicalised, so preconfirmation was absent and they
scored as misses - 6% coverage on a node that was preconfirming correctly.
Probed directly, 8 of 8 pending transactions were preconfirmed at 0ms with a
null blockHash. A transaction that reached the canonical chain before the
harness looked says nothing either way, so those are now excluded rather than
counted against coverage, and each receipt is classified: preconfirmed,
already-canonical, unmarked, or never-served. A speculative receipt with no
preconfirmation flag is now a hard failure, which is the case that would let a
client treat unconfirmed state as final. Coverage went from a spurious 6% to
86 of 86 transactions caught while still speculative.

The oversized-record burst sent nothing at all. polycli's store mode writes
the payload into contract storage at roughly 20k gas per word, so a 32KB
transaction wants ~24M gas and fails the node's tx fee cap - 400 submissions,
400 rejections, tps 0. The episode passed anyway, and the entry count it
reported as evidence came from the background transfer load. It now sends
calldata instead, where zero bytes cost 4 gas each, addressed to an
unallocated account so no deployment step is needed.

Nothing checked that the burst reached the chain, which is what let that hide.
assert_burst_landed counts transactions addressed to the sink across the
burst's block range and fails when the window is empty.

count_txs_to_sink called rpc_call, which is local to sequencer_rpc_test.sh and
not in scope in the utils. Every call failed, the count came back zero, and
the new guard reported a burst that had demonstrably landed as missing. It
uses rpc_post. Shell has no import graph, so neither bash -n nor shellcheck
can see a cross-file function reference; only running it does.

The burst sizes are now measured rather than guessed. Against an ingress
without the fix, on max.message.bytes=1048576: 32KB x 400 mined puts ~1.3MB in
a block and never fences, while 120KB x 240 fences eight times with
MESSAGE_TOO_LARGE. The same 120KB burst against an ingress with the fix stays
clean, so the episode detects the defect and clears the fix with bor and the
burst held constant. 120KB also sits just under the txpool's 128KB ceiling.
The default was 32KB, which provably never reached the path. Lowering it
disarms the episode silently, because those transactions do land and
assert_burst_landed cannot tell they were too small to coalesce past 1MB - the
comment says so.

The burst now checks publishers per node instead of a summed entry count,
which hides one dead publisher. That is not hypothetical: where the store
rejects an oversized entry as MALFORMED rather than self-fencing, a producer
without the bor-side size cap disables its own publishing and does not
re-enable it, so the chain keeps building while that node silently stops
preconfirming. A later episode's recovery check is what surfaced it; the burst
should catch its own damage.

Burst concurrency is configurable - each sender holds its own payload, and the
default killed the run on a memory-constrained host.

Also verified live: bor_getPreconfAuditStatus reported auditedThrough 0xf8
with unauditedThrough 0x7f, and 127 is exactly the last pre-Rio block, so the
unheld-window path behaves as intended against a real store. 120 of 120
sampled transactions canonicalised with no mismatches. The multicall3 pending
read succeeded 25 of 25, so that intermittent failure did not reproduce here
and is not claimed fixed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Review on 0xPolygon/bor#2388 folded the audit's coverage into
bor_getInvalidPreconfBlocks as a pendingFrom field and removed
bor_getPreconfAuditStatus along with the unauditedThrough mark, so the
method this suite probed no longer exists and the bare array of
{number, reason} becomes an object of {invalid, pendingFrom}.

Detect which shape came back and check each against its own contract:
numbers and reasons for the array, hex heights plus a present
pendingFrom and the 1024-height request cap for the object. The bor this
leg builds still returns the array, so both have to be accepted for the
object branch to go live on its own when the ref moves.

A missing pendingFrom fails while a null one passes: null is the
legitimate "fully audited", and telling the two apart is the whole
reason the field exists.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Base automatically changed from cffls/sequencer-e2e to main September 15, 2026 16:00
pratikspatil024 and others added 2 commits September 15, 2026 21:33
Empty commit. The leg's path filter only fires for main, so this PR's
suites have never run in CI; #47 merging is what makes them eligible.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The burst passed 120KB of calldata as one argv entry: 245762 hex
characters. Linux caps a single argument at MAX_ARG_STRLEN (32 pages,
131072 bytes) where macOS allows far more, so the episode ran on a
laptop and could not run in CI, which failed it with "Argument list too
long".

60KB encodes to 122882 characters, inside the limit, and a check on the
encoded length now fails with a reason if anyone raises it again rather
than letting polycli discover it a CI run later. Generating the payload
with printf's width instead of seq drops a 245760-element argument list
on the way there too.

The smaller payload should still reach the coalescing cap: a record is
capped by bytes before the 64-transaction count, so 18 pending 60KB
transactions fill 1MB where 33 were needed at 32KB. That is reasoning
rather than a measurement — with bor's own cap in place the assertion
passes whether or not the cap engaged — and the note on the constant
says so.

assert_burst_landed did its job here: it caught that nothing reached the
chain and failed instead of asserting over an empty window.

Also corrects the episode's comment, which still described the store
mode this replaced with contract-call calldata.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@pratikspatil024
pratikspatil024 merged commit 59448bd into main Sep 16, 2026
13 checks passed
@pratikspatil024
pratikspatil024 deleted the ppatil/sequencer-rpc-chaos-e2e branch September 16, 2026 04:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants