Found while verifying #749 (the WebSocket backpressure oracle rework), on the configuration CI does not run: TestBackpressureInboundSequenceIntegrity at its default size (96 connections x 4 bursts x 16,000 frames, no WS484_* env) under -race.
What happens
On GitHub's runners (kernel 6.17, 4 vCPU, 8 MiB memlock), with #749's oracle at ad88696, run 36357012593, 5 processes per arch:
|
x86 |
arm64 |
| epoll |
FAILED 3/5 |
0/5 |
| io_uring, io_uring/multishot_recv |
0/5, 0/5 |
0/5, 0/5 |
Each failure is one connection whose close wait reached the oracle's 240 s cap while bytes were still arriving (the longest silence in the wait was 1.1-2.0 s). Its WSO-GIVEUP record, in all three:
- Server socket:
inq=0, so the engine had read everything that arrived. But it advertised rcv_wnd=2048, with rcv_mss=2048, rcv_ssthresh 137-211 KB and rcv_space 90-144 KB.
- Client socket: 245-273 KB still unsent (
notsent), snd_wnd=2048, and rwnd_limited_ms 237,602-244,984, so it was receive-window-limited for almost the whole 240 s.
- The client's Close frame sits behind those bytes.
So the client-to-server direction moved about 2 KiB per ACK for four minutes, while the engine read promptly. The x86 epoll subtest took 257-289 s in 4 of the 5 processes (42 s in the fifth).
Without -race, the same configuration passed 10/10 (run 36357515331; the epoll subtest took 7-35 s on x86).
main's oracle at the same configuration (698bed6, run 36358390264) fails too: 5 of 10 processes (x86 2/5, arm64 3/5), through its fixed 30 s close budget, and one x86 epoll subtest took 277 s. It prints no state, so whether those failures are this crawl is not known. It also takes every connection whose flood write timed out out of its frame count (150-290 such bursts per subtest in that run). #749 keeps those connections in every verdict and prints the record, which is how this became visible.
Not known
- Whose doing it is. The window the server advertises is the kernel's choice, but it depends on how the socket's receive memory is used. Candidates:
- Why only x86 epoll.
Next
- At the give-up, capture
/proc/net/sockstat (TCP mem, against tcp_mem) and ss -tmi for both sockets (skmem: rb, r, d drops).
- Rerun with a larger
tcp_mem, or with 48 connections, to see whether the crawl follows memory pressure.
- If the engine is not involved, decide whether the default size under
-race should shrink (as -short already does), and say so in the test.
CI is not affected: it runs the test with WS484_* (16 x 2 x 1,000), where this has not appeared.
Evidence: evidence/celeris-b2/b2b/11-fix/02-dispatch/ (V-C4-inbound-default-race-36357012593, V-D4-inbound-default-norace-36357515331, BASE-main-default-race-36358390264, and their tally-*.txt).
Refs #749, #633
Found while verifying #749 (the WebSocket backpressure oracle rework), on the configuration CI does not run:
TestBackpressureInboundSequenceIntegrityat its default size (96 connections x 4 bursts x 16,000 frames, noWS484_*env) under-race.What happens
On GitHub's runners (kernel 6.17, 4 vCPU, 8 MiB memlock), with #749's oracle at
ad88696, run 36357012593, 5 processes per arch:Each failure is one connection whose close wait reached the oracle's 240 s cap while bytes were still arriving (the longest silence in the wait was 1.1-2.0 s). Its
WSO-GIVEUPrecord, in all three:inq=0, so the engine had read everything that arrived. But it advertisedrcv_wnd=2048, withrcv_mss=2048,rcv_ssthresh137-211 KB andrcv_space90-144 KB.notsent),snd_wnd=2048, andrwnd_limited_ms237,602-244,984, so it was receive-window-limited for almost the whole 240 s.So the client-to-server direction moved about 2 KiB per ACK for four minutes, while the engine read promptly. The x86 epoll subtest took 257-289 s in 4 of the 5 processes (42 s in the fifth).
Without
-race, the same configuration passed 10/10 (run 36357515331; the epoll subtest took 7-35 s on x86).Before #749
main's oracle at the same configuration (
698bed6, run 36358390264) fails too: 5 of 10 processes (x86 2/5, arm64 3/5), through its fixed 30 s close budget, and one x86 epoll subtest took 277 s. It prints no state, so whether those failures are this crawl is not known. It also takes every connection whose flood write timed out out of its frame count (150-290 such bursts per subtest in that run). #749 keeps those connections in every verdict and prints the record, which is how this became visible.Not known
-raceon 4 vCPUs.Next
/proc/net/sockstat(TCPmem, againsttcp_mem) andss -tmifor both sockets (skmem:rb,r,ddrops).tcp_mem, or with 48 connections, to see whether the crawl follows memory pressure.-raceshould shrink (as-shortalready does), and say so in the test.CI is not affected: it runs the test with
WS484_*(16 x 2 x 1,000), where this has not appeared.Evidence:
evidence/celeris-b2/b2b/11-fix/02-dispatch/(V-C4-inbound-default-race-36357012593,V-D4-inbound-default-norace-36357515331,BASE-main-default-race-36358390264, and theirtally-*.txt).Refs #749, #633