Skip to content

feat: sequence-store publisher e2e suite - #47

Merged
cffls merged 9 commits into
mainfrom
cffls/sequencer-e2e
Sep 15, 2026
Merged

cffls merged 9 commits into
mainfrom
cffls/sequencer-e2e

Conversation

@cffls

@cffls cffls commented Aug 24, 2026

Copy link
Copy Markdown
Contributor

Summary

Adds the sequence-store & publisher e2e integration: a kurtosis devnet config with the in-enclave sequence store (4 publishing validators + 1 RPC on bor:local, heimdall-v2, 4s blocks), a test suite porting the store-outage drill validated on a lab devnet against bor PR 0xPolygon/bor#2355, and a workflow running the suite in this repo's own CI. Consumable from bor CI exactly like the stateless suite: copy configs/kurtosis-sequencer-e2e.yml, run tests/sequencer_tests/.

The tests assert: all publishers live post-Rio (sequencer_publish_state=1, entries flowing); a 300s store outage costs zero block production (≥60 blocks at 4s); recovery floors the backfill at milestone finality (Sequencer backfill jumping finalized heights in the producer log, reconcile_forwardjump ≥ 1, reconcile_supersede == 0), and publishing resumes.

Notes:

  • The store image pulls from the private ghcr.io/0xpolygon/sequence-store package (Actions read access granted to this repo and bor); SEQUENCE_STORE_TAG is latest until a release tag is pinned. A build-from-source fallback is documented in the workflow.
  • kurtosis-pos is pinned to main-branch commit 0cdede09 (the sequence-store launcher landed after v1.3.4); re-pin at the next release.
  • This repo's workflow builds bor from the sequencing branch until the publisher merges to develop; triggers are path-scoped to this suite's files + workflow_dispatch so the feature-branch build never gates unrelated PRs.

Executed tests

The drill this suite ports ran end-to-end on a kurtosis devnet (lab host): 300s store outage at heights 151..226, 75/75 blocks produced during the outage, recovery forward-jumped the finalized debt in one reconcile (floor=232 > debt 152..229, zero rebuilt), zero revocations/reorders/generation mismatches in the store probe. Locally: all YAML parses, both scripts pass bash -n; the CI job itself has not yet run (needs this PR's workflow on a runner).

Rollout notes

CI-only change; no node or network impact. Requires the ghcr package access already granted. Follow-ups tracked in comments: pin SEQUENCE_STORE_TAG, re-pin kurtosis-pos to a release tag, switch the bor ref to develop after 0xPolygon/bor#2355 merges.

🤖 Generated with Claude Code

Kurtosis devnet config with the in-enclave sequence store (4 publishing
validators + 1 RPC, bor built from the sequencing branch until the
publisher merges), a test suite porting the store-outage drill —
publishers live post-Rio, 300s outage with stall-free production,
recovery floored at milestone finality (forward jump, zero
supersessions) — and a path-scoped workflow running it in this repo's
CI. Consumable from bor CI like the stateless suite: copy
configs/kurtosis-sequencer-e2e.yml, run tests/sequencer_tests/.

The store image pulls from ghcr.io/0xpolygon/sequence-store (private;
Actions read access granted to this repo and bor), with a documented
build-from-source fallback.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Comment on lines +42 to +66
runs-on: ubuntu-latest
timeout-minutes: 20
steps:
- name: Checkout bor
uses: actions/checkout@v5
with:
repository: 0xPolygon/bor
# The sequencer publisher lives on the sequencing branch
# (PR #2355); switch this to develop once it merges there.
ref: sequencing

- name: Build bor docker image
run: docker build -t bor:local --file Dockerfile .

- name: Save bor docker image
run: docker save bor:local | gzip > bor-image.tar.gz

- name: Upload bor docker image
uses: actions/upload-artifact@v4
with:
name: bor-image
path: bor-image.tar.gz
retention-days: 1

build-heimdall-v2:
Comment on lines +67 to +89
runs-on: ubuntu-latest
timeout-minutes: 20
steps:
- name: Checkout heimdall-v2
uses: actions/checkout@v5
with:
repository: 0xPolygon/heimdall-v2
ref: develop

- name: Build heimdall-v2 docker image
run: docker build -t heimdall-v2:local --file Dockerfile .

- name: Save heimdall-v2 docker image
run: docker save heimdall-v2:local | gzip > heimdall-v2-image.tar.gz

- name: Upload heimdall-v2 docker image
uses: actions/upload-artifact@v4
with:
name: heimdall-v2-image
path: heimdall-v2-image.tar.gz
retention-days: 1

e2e-tests:
cffls and others added 5 commits August 24, 2026 15:28
The sequence-store package has no latest tag — its docker workflow
publishes dev from the default branch and version tags on v* releases;
pull dev until one is cut. shfmt 3.13.1 formatting applied per the
checks workflow.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The first run built bor from sequencing, which does not carry the
publisher yet (0xPolygon/bor#2355 is unmerged) — no sequencer metrics
exist on such a build and test 1 reads 0/4 forever. Build from the PR
branch until it merges (then sequencing, then develop), and on a test-1
failure print each validator's sequencer series count and publish
state, telling a sequencer-less image apart from a wrong-state
publisher.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The sequencer e2e leg was failing in the "Pre kurtosis run" setup step
(foundry download) at the old 0cdede0 main pin, so the sequencer suite
never ran. v1.4.2 is the first kurtosis-pos release tag that carries the
sequence-store launcher (fully contains 0cdede0) and ships the
setup-action fix bor's own kurtosis legs already run with. Bumping to it
unblocks the run now that the publisher branch it builds bor from is
rebased onto develop.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Brings the sibling kurtosis legs up to their v1.4.1 pin and refreshed
Actions from #49/#50/#51 so they stop failing the stale-v1.3.4 setup on
this branch. The sequencer leg keeps its own v1.4.2 pin.
The v1.4.1 pin's shared setup action installs foundry via
foundry-toolchain@c7450ba (v1.8.0), which fails the "Pre kurtosis run"
step with a corrupt foundryup download (`cannot execute binary file`).
v1.4.2 bumps that action to v1.9.1 (908c5403) — the same fix bor develop
landed in #2377. This puts all three pos-workflows kurtosis legs
(e2e, stateless, sequencer) on v1.4.2.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@cffls
cffls marked this pull request as ready for review August 27, 2026 20:06
cffls and others added 3 commits August 28, 2026 09:39
Extends the sequencer e2e suite with two scenarios proven on local
kurtosis devnets that the current three tests don't cover:

- pending-state readable on every validator: guards the regression where
  a signer outside the active producer set stopped refreshing its pending
  snapshot, so eth_getBalance/eth_call against "pending" returned
  "missing trie node / layer stale" on non-producing nodes.
- producer takeover: stops the active producer's bor and requires the
  chain to keep advancing (a store window on a displaced parent used to
  wedge the seal barrier), then restarts it and requires a clean rejoin.
  Adoption/supersede counters are logged as evidence rather than asserted,
  since they depend on whether a dangling window existed at the cut.

Also asserts reconcile displaced-records stay 0 in steady state (one owner
per height; a displacement is a revoked preconfirmation).

Scenarios needing a load generator (consumer preconf receipts,
dual-publisher contention) or a bigger topology (multi-broker / multi-
gateway chaos) are left for a follow-up that adds those to the CI rig.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…ered preconfs

The suite only checked producer self-reported counters (supersede,
displaced). Add a check that reads the seqstore-auditor's evidence — an
independent follower of the store log that classifies every superseded
generation — and fails on any revocation that dropped preconfirmations or
reordered transactions. This catches a producer that drops or reorders
content even when its own metrics disagree, and a supersession storm (the
pre-follow-model churn regression).

Known limitation: "dropped" and "reordered" are per-transaction, so they
only bite when blocks carry load. This rig has no transaction generator,
so today the check asserts the absence of violations rather than provokes
them. Adding a load step (so windows carry txs) turns it into a real
revocation/reorder check, especially across the takeover cut — tracked as
the follow-up alongside the load-dependent scenarios noted earlier.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
The auditor's dropped/reordered signals are per-transaction, so they only
mean anything when windows carry content. Add a background load generator
(polycli transfer mode, the devnet's prefunded account, sent to the rpc
node so no validator we stop is the entry point) that runs through the
outage, recovery, and takeover phases, and verify transactions actually
land before proceeding. The takeover now revokes real preconfirmations if
it regresses, which the auditor check catches as dropped transactions.

The sequencer workflow gains the polycli install step the sibling kurtosis
legs already have. When polycli is unavailable the load is skipped, not
failed, so the suite still runs without the load-dependent coverage.

Reordering is only fully provoked by independent senders; a single loader's
nonce order limits it, so that signal stays best-effort.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@cffls
cffls merged commit 8092c2e into main Sep 15, 2026
13 checks passed
@pratikspatil024
pratikspatil024 deleted the cffls/sequencer-e2e branch September 15, 2026 16:00
pratikspatil024 added a commit that referenced this pull request Sep 15, 2026
Empty commit. The leg's path filter only fires for main, so this PR's
suites have never run in CI; #47 merging is what makes them eligible.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
pratikspatil024 added a commit that referenced this pull request Sep 16, 2026
* test(sequencer): add preconf rpc and store chaos e2e suites

The publisher suite covers whether validators write to the store. These cover
what a client sees and whether the store can hurt the chain, both ported down
from the manual campaign in #tmp_seqstore-testing-and-qa to a size CI affords.

Both scripts source the existing utils rather than editing them, so the
functional suite is untouched and the three run as separate workflow steps
against one enclave, the way the sibling kurtosis leg already sequences its
suites.

rpc suite: a preconfirmed receipt is marked preconfirmation:true and carries a
null blockHash, which is the client's only signal before canonicalisation;
every sampled transaction then has to reach the canonical chain with a real
hash and a success status, the 452k-receipt re-fetch scaled to 120. Plus
eth_sendRawTransactionSync being registered at all - it was missing from a
deployed private RPC, which returns -32601 and breaks any client built on the
sync path - the multicall3 pending read looped 25 times, since intermittent
failure there becomes random estimateGas failures for geth-based clients, and
bor_getInvalidPreconfBlocks answering an array of numbered, reasoned records
and rejecting a reversed range.

Hashes come from the pending block, not the load generator's output: the
pending view is the consumer's own speculative state, so anything in it should
be servable as a preconfirmation, and it keeps the suite off polycli's log
format.

chaos suite: five episodes through one harness - broker stop, ingress latency,
gateway pause, a seeded random fault, and an oversized-record burst. Each
holds the fault, asserts the chain kept building, repairs, and requires
publishing to resume on its own; a store fault that permanently de-registered
a publisher would leave a healthy-looking chain with every preconfirmation
silently gone. A 1Hz head sampler runs across the whole session and the
closing assertion is that no node ever reported one hash for a height and
later a different one, which is the property a store fault must not be able
to break.

Throughput under fault is printed, never asserted. The campaign measured
25-50% cost from store outage or latency, against a design brief that says the
store cannot affect block production. A threshold either fails today or
freezes whichever number is currently true into CI, so the numbers are output
for a human.

bor_getPreconfAuditStatus skips on -32601 rather than failing, the same shape
the suite already uses for an absent polycli, so it lands now and goes live
once the audit ships.

Two notes on what the tooling actually supports: polycli has no --sync-txs on
the pinned release (or on main), so the sync path gets a registration check
rather than a latency measurement; and --calldata needs contract-call mode
with a deployed address, so the oversized-record burst uses store mode with
--store-data-size.

Timeout goes to 75 minutes for the three suites, and the diagnostics dump now
triggers on any of them failing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(sequencer): fix four assertions that passed without testing anything

Ran both suites against a local kurtosis devnet with the store built from
source. Four of the assertions were not measuring what they claimed, and two
of them I had already reported as passing.

The preconfirmation check was measuring its own latency. It read one pending
block of 143 transactions, then walked it with a request per hash; by the
tenth, the rest had canonicalised, so preconfirmation was absent and they
scored as misses - 6% coverage on a node that was preconfirming correctly.
Probed directly, 8 of 8 pending transactions were preconfirmed at 0ms with a
null blockHash. A transaction that reached the canonical chain before the
harness looked says nothing either way, so those are now excluded rather than
counted against coverage, and each receipt is classified: preconfirmed,
already-canonical, unmarked, or never-served. A speculative receipt with no
preconfirmation flag is now a hard failure, which is the case that would let a
client treat unconfirmed state as final. Coverage went from a spurious 6% to
86 of 86 transactions caught while still speculative.

The oversized-record burst sent nothing at all. polycli's store mode writes
the payload into contract storage at roughly 20k gas per word, so a 32KB
transaction wants ~24M gas and fails the node's tx fee cap - 400 submissions,
400 rejections, tps 0. The episode passed anyway, and the entry count it
reported as evidence came from the background transfer load. It now sends
calldata instead, where zero bytes cost 4 gas each, addressed to an
unallocated account so no deployment step is needed.

Nothing checked that the burst reached the chain, which is what let that hide.
assert_burst_landed counts transactions addressed to the sink across the
burst's block range and fails when the window is empty.

count_txs_to_sink called rpc_call, which is local to sequencer_rpc_test.sh and
not in scope in the utils. Every call failed, the count came back zero, and
the new guard reported a burst that had demonstrably landed as missing. It
uses rpc_post. Shell has no import graph, so neither bash -n nor shellcheck
can see a cross-file function reference; only running it does.

The burst sizes are now measured rather than guessed. Against an ingress
without the fix, on max.message.bytes=1048576: 32KB x 400 mined puts ~1.3MB in
a block and never fences, while 120KB x 240 fences eight times with
MESSAGE_TOO_LARGE. The same 120KB burst against an ingress with the fix stays
clean, so the episode detects the defect and clears the fix with bor and the
burst held constant. 120KB also sits just under the txpool's 128KB ceiling.
The default was 32KB, which provably never reached the path. Lowering it
disarms the episode silently, because those transactions do land and
assert_burst_landed cannot tell they were too small to coalesce past 1MB - the
comment says so.

The burst now checks publishers per node instead of a summed entry count,
which hides one dead publisher. That is not hypothetical: where the store
rejects an oversized entry as MALFORMED rather than self-fencing, a producer
without the bor-side size cap disables its own publishing and does not
re-enable it, so the chain keeps building while that node silently stops
preconfirming. A later episode's recovery check is what surfaced it; the burst
should catch its own damage.

Burst concurrency is configurable - each sender holds its own payload, and the
default killed the run on a memory-constrained host.

Also verified live: bor_getPreconfAuditStatus reported auditedThrough 0xf8
with unauditedThrough 0x7f, and 127 is exactly the last pre-Rio block, so the
unheld-window path behaves as intended against a real store. 120 of 120
sampled transactions canonicalised with no mismatches. The multicall3 pending
read succeeded 25 of 25, so that intermittent failure did not reproduce here
and is not claimed fixed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(sequencer): track the reshaped invalid-preconf response

Review on 0xPolygon/bor#2388 folded the audit's coverage into
bor_getInvalidPreconfBlocks as a pendingFrom field and removed
bor_getPreconfAuditStatus along with the unauditedThrough mark, so the
method this suite probed no longer exists and the bare array of
{number, reason} becomes an object of {invalid, pendingFrom}.

Detect which shape came back and check each against its own contract:
numbers and reasons for the array, hex heights plus a present
pendingFrom and the 1024-height request cap for the object. The bor this
leg builds still returns the array, so both have to be accepted for the
object branch to go live on its own when the ref moves.

A missing pendingFrom fails while a null one passes: null is the
legitimate "fully audited", and telling the two apart is the whole
reason the field exists.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* ci(sequencer): trigger the e2e leg now that #47 has landed

Empty commit. The leg's path filter only fires for main, so this PR's
suites have never run in CI; #47 merging is what makes them eligible.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test(sequencer): size the burst payload for Linux argv

The burst passed 120KB of calldata as one argv entry: 245762 hex
characters. Linux caps a single argument at MAX_ARG_STRLEN (32 pages,
131072 bytes) where macOS allows far more, so the episode ran on a
laptop and could not run in CI, which failed it with "Argument list too
long".

60KB encodes to 122882 characters, inside the limit, and a check on the
encoded length now fails with a reason if anyone raises it again rather
than letting polycli discover it a CI run later. Generating the payload
with printf's width instead of seq drops a 245760-element argument list
on the way there too.

The smaller payload should still reach the coalescing cap: a record is
capped by bytes before the 64-transaction count, so 18 pending 60KB
transactions fill 1MB where 33 were needed at 32KB. That is reasoning
rather than a measurement — with bor's own cap in place the assertion
passes whether or not the cap engaged — and the note on the constant
says so.

assert_burst_landed did its job here: it caught that nothing reached the
chain and failed instead of asserting over an empty window.

Also corrects the episode's comment, which still described the store
mode this replaced with contract-call calldata.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants