Summary
When adaptive reverts from io_uring to epoll, epoll's ResumeAccept wakes its parked loops and returns without waiting for any of them to re-create a listener. io_uring's PauseAccept then closes io_uring's listener. If the first woken loop is scheduled late, the SO_REUSEPORT group has no listener for tens of microseconds, and dials arriving in that window are refused (ECONNREFUSED).
This is on main and independent of #657's fix. It was measured as a side result of the pre-registered campaign on PR #681 (evidence: evidence/celeris-657/impl/pr2-r2/measure/, REPORT.txt and results.json, analysis by tools/analyze.py M).
Measured on main (b79888e)
Shape: round 1's switch-stall shape (5k conn/s churn, 32 keep-alives, two promote/revert cycles), 8 MiB memlock (io_uring 1 worker, epoll 4 loops), 104 runs, 208 reverts.
- Refusals: 7 of 104 runs refused at least one churn dial within ±100 ms of a revert, 8 dials across 208 reverts (3.4% of reverts).
- The gap, per revert: an in-process trace plus an independent 1 ms netlink LISTEN sampler.
- The first epoll
listen() returns a median 67 µs after the resume call (p95 158, max 254).
- io_uring's listener close returns a median 125 µs after the pause call.
- In 8 of 208 reverts a listener-free gap of 15-91 µs opened.
- Every refused dial ended inside the gap. No refused revert had an epoll listener up before io_uring's close.
- All refusals happened with at least one epoll loop parked at
ResumeAccept. That was the case in every revert in this shape.
The fix is known, as a measured control
A control variant in which ResumeAccept waits until at least one epoll listener is in LISTEN before returning removed every refusal: 0 of 80 reverts, against 12 of 416 for base and the PR combined. A positive control that delays every epoll re-listen by 20 ms refused in 24 of 24 reverts, so the instrument sees the gap when it exists.
Proposed: make ResumeAccept (or performSwitch's revert order) ensure the incoming engine has a listener in LISTEN before the outgoing engine closes its own. Bound the wait, and count a timeout. Add a failing-first test that forces the late wake-up (the POS injection) and asserts 0 refused dials.
Found by the refused-dials measurement on PR #681. Related: #657, #662.
Summary
When adaptive reverts from io_uring to epoll, epoll's
ResumeAcceptwakes its parked loops and returns without waiting for any of them to re-create a listener. io_uring'sPauseAcceptthen closes io_uring's listener. If the first woken loop is scheduled late, the SO_REUSEPORT group has no listener for tens of microseconds, and dials arriving in that window are refused (ECONNREFUSED).This is on main and independent of #657's fix. It was measured as a side result of the pre-registered campaign on PR #681 (evidence:
evidence/celeris-657/impl/pr2-r2/measure/, REPORT.txt and results.json, analysis bytools/analyze.py M).Measured on main (b79888e)
Shape: round 1's switch-stall shape (5k conn/s churn, 32 keep-alives, two promote/revert cycles), 8 MiB memlock (io_uring 1 worker, epoll 4 loops), 104 runs, 208 reverts.
listen()returns a median 67 µs after the resume call (p95 158, max 254).ResumeAccept. That was the case in every revert in this shape.The fix is known, as a measured control
A control variant in which
ResumeAcceptwaits until at least one epoll listener is in LISTEN before returning removed every refusal: 0 of 80 reverts, against 12 of 416 for base and the PR combined. A positive control that delays every epoll re-listen by 20 ms refused in 24 of 24 reverts, so the instrument sees the gap when it exists.Proposed: make
ResumeAccept(orperformSwitch's revert order) ensure the incoming engine has a listener in LISTEN before the outgoing engine closes its own. Bound the wait, and count a timeout. Add a failing-first test that forces the late wake-up (the POS injection) and asserts 0 refused dials.Found by the refused-dials measurement on PR #681. Related: #657, #662.