You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The adaptive job's race step ("Adaptive — race tests (memlock raised, up-switch required)" in .github/workflows/ci.yml) runs go test -race ... -skip '^(TestRampH1Sync|TestRampH1Async)$' ./adaptive/.... These two tests are the only CI coverage at the scale celeris#662 is about: 2048 connections promoted and reverted on a GitHub-hosted runner.
The workflow rule is that every quarantine has an open issue and is removed when that issue closes. The entry was there for celeris#657 and celeris#662. #657 was closed by #681 and #687. #662 closes when PR #674 merges. After that, this issue is the one tracking the entry.
The r6/fix re-gate, 2026-09-26, under load. Other lanes' native go test load was on the host, and the containers' docker check could not see it. Main failed TestRampH1Async in 3 of 4 containers and TestRampH1Sync in 2 of 4, with 6,615 read: reset and 1 dial: reset in 4,354,688 requests. Per ramp it served 176k-1.06M requests, against 1.06M-1.28M in the r6 gate. Two of the four containers are the P7 ramp pair's main arm, which PR fix(engine): keep serving on a paused listener for 1.5 s with TCP_DEFER_ACCEPT cleared, so a switch or PauseAccept no longer resets clients that had not yet sent a request (celeris#662, #675) #674 marks contaminated and no longer cites (r6/round4/30-CORRECTIONS.md §6): p7-a01 (both FAIL, 2,580 errors, 219,754 Sync requests) and p7-a04 (both PASS, 403 errors). The other two are full ./adaptive G5 containers that ran in the same loaded window: g5-r1 (Sync PASS, Async FAIL) and g5-r2 (both FAIL).
So main's ramp failures have been seen only in that loaded run, and every failing ramp served at most 458,167 requests. Whether the load causes the failures is not established. What is established is that in the one run with no known load, main passed every ramp run.
r6 gate (2026-09-20, 53b52f1): both tests passed in 6 of 6 containers, with 0 errors in 13,260,792 requests. The outgoing epoll held 0 of 2048 in every high phase.
r6/fix re-gate (2026-09-26, 1d90b5d, under the load described above): both tests passed in 4 of 4 containers. There were 2 dial: reset errors in 3,612,871 requests: handshakes still in flight when a listener closed, which closing a listen socket cannot avoid.
Review round 4 (2026-09-27, c12d1b8, the same product code and ./adaptive tests as the head 7aeb2ec): 1 full ./adaptive container with other containers beside it. Both tests passed, with 0 errors in 1,455,255 requests.
None of this has run on a GitHub-hosted runner yet. That runner is the shape in which celeris#686's resets were seen, which is why the quarantine stays in place until the lift rule is met.
What it takes to lift
The lift rule is at least 6 GitHub-hosted runs of identical bytes, all green:
Every run must pass both tests. Record each ramp phase's epoll= / io_uring= placement, the err= census, and each ramp's ramp complete: ok= throughput from the log.
Delete the -skip, and add both names to the step's tally: exactly one top-level RUN and one PASS each, and no SKIP line. That keeps them from going quiet again.
If a run fails, classify the failure before re-quarantining: placement (#657's class), a dial-time reset at a listener close (inherent), or something new. Classify it against the r6 gate, not the re-gate. In the r6 gate main passed every ramp run, so "main fails this too" is not a known baseline. Put the ramp's throughput beside each failure, because each of main's local failures came with throughput well below the gate's. Keep this issue open until the tally is in place.
Evidence (local, under evidence/celeris-662/):
r6/round4/r43/RAMP-BASELINE.txt (python3 r6/round4/r43/tools/ramp-baseline.py): every container above, per arm, with the two runs kept apart and each ramp's throughput.
r6/ship/RAMP-RESET-CENSUS.txt: the per-arm error census of both runs. Its r6-gate header says 2026-09-26, but that gate ran on 2026-09-20 (r6/round4/30-CORRECTIONS.md §3).
r6/fix/G5-TALLY.txt
r6/gate/02-G5-TALLY.txt
Edited 2026-09-27 (PR #674, review round 4.3). Main's baseline used to be one pooled figure ("failed in 3 of 4 containers ... 2 of 4 ... 6,615 read: reset"). That figure came from the loaded re-gate alone and included P7's main arm, which #674 marks contaminated. It left out the r6 gate, where main passed every ramp run. The two runs are now shown separately.
What is quarantined
The
adaptivejob's race step ("Adaptive — race tests (memlock raised, up-switch required)" in.github/workflows/ci.yml) runsgo test -race ... -skip '^(TestRampH1Sync|TestRampH1Async)$' ./adaptive/.... These two tests are the only CI coverage at the scale celeris#662 is about: 2048 connections promoted and reverted on a GitHub-hosted runner.The workflow rule is that every quarantine has an open issue and is removed when that issue closes. The entry was there for celeris#657 and celeris#662. #657 was closed by #681 and #687. #662 closes when PR #674 merges. After that, this issue is the one tracking the entry.
What failed before
io_uring=1316-1318of 2048 after 20 s, with err=0 across 938,818 requests. Nothing was dropped; the adaptive: forced switches leave keep-alive connections behind — async conns miss the 1.2 s flap window in every run, and a one-worker io_uring strands sync conns until their requests fail #657 transplant just never finished../adaptive), so 12 of 12 ramp runs passed. Those runs had 1,556read: reseterrors in 13,885,363 requests (83 to 495 per container). That is the celeris#686 class, the accept hand-over that fix(engine): keep serving on a paused listener for 1.5 s with TCP_DEFER_ACCEPT cleared, so a switch or PauseAccept no longer resets clients that had not yet sent a request (celeris#662, #675) #674 removes, and on its own it failed no test. Each ramp served 1.18M-1.28M requests (Sync) and 1.06M-1.12M (Async). Isolation: each log checks that no container ran at its start. Native host load was not recorded.go testload was on the host, and the containers' docker check could not see it. Main failedTestRampH1Asyncin 3 of 4 containers andTestRampH1Syncin 2 of 4, with 6,615read: resetand 1dial: resetin 4,354,688 requests. Per ramp it served 176k-1.06M requests, against 1.06M-1.28M in the r6 gate. Two of the four containers are the P7 ramp pair's main arm, which PR fix(engine): keep serving on a paused listener for 1.5 s with TCP_DEFER_ACCEPT cleared, so a switch or PauseAccept no longer resets clients that had not yet sent a request (celeris#662, #675) #674 marks contaminated and no longer cites (r6/round4/30-CORRECTIONS.md§6): p7-a01 (both FAIL, 2,580 errors, 219,754 Sync requests) and p7-a04 (both PASS, 403 errors). The other two are full./adaptiveG5 containers that ran in the same loaded window: g5-r1 (Sync PASS, Async FAIL) and g5-r2 (both FAIL).What #674 shows locally
53b52f1): both tests passed in 6 of 6 containers, with 0 errors in 13,260,792 requests. The outgoing epoll held 0 of 2048 in every high phase.1d90b5d, under the load described above): both tests passed in 4 of 4 containers. There were 2dial: reseterrors in 3,612,871 requests: handshakes still in flight when a listener closed, which closing a listen socket cannot avoid.c12d1b8, the same product code and./adaptivetests as the head7aeb2ec): 1 full./adaptivecontainer with other containers beside it. Both tests passed, with 0 errors in 1,455,255 requests.None of this has run on a GitHub-hosted runner yet. That runner is the shape in which celeris#686's resets were seen, which is why the quarantine stays in place until the lift rule is met.
What it takes to lift
The lift rule is at least 6 GitHub-hosted runs of identical bytes, all green:
-skipremoved. Raise memlock and setCELERIS_REQUIRE_UPSWITCH=1, as the step already does. Do this 6 times on GitHub-hosted runners, for example on a draft PR that changes only that line, re-run 5 times.epoll=/io_uring=placement, theerr=census, and each ramp'sramp complete: ok=throughput from the log.-skip, and add both names to the step's tally: exactly one top-level RUN and one PASS each, and no SKIP line. That keeps them from going quiet again.If a run fails, classify the failure before re-quarantining: placement (#657's class), a dial-time reset at a listener close (inherent), or something new. Classify it against the r6 gate, not the re-gate. In the r6 gate main passed every ramp run, so "main fails this too" is not a known baseline. Put the ramp's throughput beside each failure, because each of main's local failures came with throughput well below the gate's. Keep this issue open until the tally is in place.
Evidence (local, under
evidence/celeris-662/):r6/round4/r43/RAMP-BASELINE.txt(python3 r6/round4/r43/tools/ramp-baseline.py): every container above, per arm, with the two runs kept apart and each ramp's throughput.r6/ship/RAMP-RESET-CENSUS.txt: the per-arm error census of both runs. Its r6-gate header says 2026-09-26, but that gate ran on 2026-09-20 (r6/round4/30-CORRECTIONS.md§3).r6/fix/G5-TALLY.txtr6/gate/02-G5-TALLY.txtEdited 2026-09-27 (PR #674, review round 4.3). Main's baseline used to be one pooled figure ("failed in 3 of 4 containers ... 2 of 4 ... 6,615
read: reset"). That figure came from the loaded re-gate alone and included P7's main arm, which #674 marks contaminated. It left out the r6 gate, where main passed every ramp run. The two runs are now shown separately.