test(websocket): the backpressure oracles wait on progress and read while they wait, judge every give-up with both ends' timeline, and fail a connection the engine stops reading (celeris#633, celeris#623, celeris#611, celeris#607 class) - #749
Conversation
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Important Review skippedReview was skipped as selected files did not have any reviewable changes. ⚙️ Run configurationConfiguration used: Repository: goceleris/celeris/.coderabbit.yaml Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: goceleris/celeris/.coderabbit.yaml Review profile: CHILL Plan: Advanced Run ID: 📒 Files selected for processing (4)
Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 9 remain after this review. 📝 WalkthroughWalkthroughLinux WebSocket backpressure tests now share bounded wait and connection-diagnostic support. The inbound-sequence test distinguishes delivered frames from frames not delivered after echo writes fail. The pause-cancel test tracks client writes and close handshakes and asserts handler-error and receive-stall results. ChangesWebSocket test oracles
Priority: ⬇️ Low Estimated code review effort: 4 (Complex) | ~60 minutes Change: Other · Severity of issue fixed: Low Merge Risk: ⚪ Minimal · up to This change strengthens the Linux WebSocket backpressure tests and does not modify product code. No concrete regressions remain, so it is ready to merge after normal CI. 🚥 Pre-merge checks | ✅ 4✅ Passed checks (4 passed)
Comment |
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
…, with waits that give up on no progress and both ends' timeline (celeris#633, celeris#623) TestBackpressurePauseDoesNotCancelInflightSend finished its last frame and its Close frame with fixed 15 s and 10 s deadlines, waited a fixed 10 s for the server's close, never read while it waited, and counted a connection that never finished its writes (clientCloseFail) without judging it. celeris#633 measured what those deadlines judged on GitHub's runners: a server that was behind but moving (457 of 459 close-timeouts were still receiving bytes), a client that stalled itself because its full receive queue makes Linux 6.17 discard the server's ACKs, and server errors that were the echo of the client's own give-up (1,165 of 1,165 after that client's close). The client now writes in 1 s slices, empties its receive queue without blocking after every slice that times out, and gives up only after 90 s with no progress (no byte written, no shrink of its send queue) or at 240 s; the close wait is re-armed by every byte and gives up after 30 s of silence or at 240 s. SO_RCVBUF stays 32 KiB and the flood never reads, so the echo SEND still blocks, which is what the test is for. In the same change, clientCloseFail is asserted (celeris#623), and a server error is judged only when it came before its own client closed. Every give-up prints a WSO-GIVEUP record: the connection's milestones and a 1 s timeline of both ends (client socket; handler phase and echoes; chanReader depth, spill, pause state and the pauses and resumes it applied; the server socket read through the engine's descriptor: bytes read by the engine, bytes accepted by the kernel and ACKed, the handler bytes the engine has not handed to the kernel), and the shape it stalled in. The client joins its handler through ?cid=<its port> and Conn.Query.
…p, and judge the echo failure with its connection (celeris#611) TestBackpressureInboundSequenceIntegrity's handler returned when its echo write failed, leaving frames the engine had delivered unread, and the test then reported them lost (99.8-99.9% of the "missing" frames were still queued at handler exit); the echo failure itself was logged and never judged. The handler now stops echoing and keeps reading to the end of the stream, so the frame count measures delivery; frames read after the echo failed are counted separately, and a mismatch names the connection and both populations. An echo write or read error before its own client closed fails the test with the error and the connection's address; after the client's close it is that client's teardown and is printed, not judged. Its client gets the same wait/read pattern as the #482 oracle's (celeris#633): a flood write that times out is the backpressure the test creates and ends the burst, instead of taking the connection out of the frame-count verdict; the last frame and the Close frame are written on progress with drains between slices; the close wait is re-armed per byte. A client that cannot finish, and a close that arrives as an RST, now fail the test with the connection's timeline, as clientCloseFail and clientRST were counted and not judged. writeAll, which only the celeris_closeprobe-tagged close-drain test still uses, moves to a file under that tag, so the default build does not carry an unused helper.
…e every server error on a connection that closed cleanly (celeris#607 class, celeris#623) The progress-based client of the previous commits empties its receive queue between write slices. That is what reopens a client window a blocked server SEND waits on, so it also gets through a connection whose engine chained its RECV behind that SEND (celeris#607). Measured in CI's shape on GitHub's runners with #607 re-introduced (flushSendLink's Detached early-return deleted), 5 processes per arch: the #482 oracle before this PR counted such connections as clientCloseFail (19 to 55 of 96, in 6 of 10 processes) and never judged them; this PR's previous head finished every one of them (0 of 10 processes failed) while the engine held a chained recv for up to 6.5 s. So the rig now watches the server side of every connection every 250 ms for its whole life. A connection whose server socket holds unread bytes, whose chanReader is neither paused nor holding anything, and whose handler has not exited is one the engine should be reading. If the engine reads nothing from it for 5 s the test fails with that stretch's state at both ends (WSO-STALL), whatever the client does afterwards. Bytes read, pauses and resumes only grow, so equal readings at a stretch's two ends prove it; a watcher run more than 1 s after its tick was due means the process was starved, and every open stretch starts over. The healthy maximum measured in CI's shape was 2.3 s (epoll) and 1.7 s (io_uring). After its flood the #482 client now holds its backpressure idle for 8 s, so a stretch that begins late in the flood still reaches the limit. On io_uring both oracles also assert RecvLinkedArms == 0: every connection on these servers is a detached WebSocket, and a chain on one is #607's own mechanism. Server errors: an error is excused only on a connection whose client gave up or was reset (which already fails the test) and only after that client closed. An error on a connection whose client read the server's close is judged whenever the handler noted it: that client's queue was empty and its close sent a FIN, and an engine error is noted only when the handler next returns from a read or write. Also: - every wait also ends at the test binary's deadline less 60 s, and a subtest that starts with less than 30 s of that left fails at once, so a wedge under CI's -timeout=300s ends as each subtest's verdict and WSO-GIVEUP record, not a goroutine dump, while a slow run that would have finished in time still does; - the inbound verdict line counts bursts a flood write deadline ended (floodDeadlines, floodDeadlineConns), and a connection whose close wait gave up or was reset leaves the frame count, as a failed write already did: the frames still in its client's send queue were never the engine's to deliver, and the give-up fails the test with its record; - the timeline's engine-read and pending-write bytes count from the connection's first byte net of the upgrade, whose sizes the client reads from its own socket, not from a baseline sampled after the 101 was queued; a drain that finds the connection reset or closed ends the wait as "rst" or "eof"; - PauseCancel's protoErr and otherWriteErr keep main's meaning (every error); the judged ones are protoErrJudged and otherWriteErrJudged; the waits line adds recvStallMax, watchResets and watchMaxLag, and a "longest unread stretch" line shows how close a run came to the limit.
1a6a60f to
1ee81fc
Compare
Summary
clientCloseFailis asserted, in the same change as the client fix, never before it. A connection that never finishes its last frame or its Close frame fails the test.WSO-STALL). On io_uring both oracles also assert that no recv was chained behind a send. With io_uring: a linked recv never starts because its chained SEND blocks on a detached WebSocket peer's closed window #607 re-introduced, this head fails 10 of 10 processes, and each check alone fails all 10; the previous head passed 10 of 10.WSO-GIVEUPrecord: the connection's timeline from both ends, one sample per second of waiting, and the shape it stalled in. Every wait also ends before the test binary's-timeout, so a wedge ends as a verdict, not a goroutine dump.What changes
Client: waits on progress (both tests)
The #633 investigation measured what the old fixed deadlines judged on GitHub's runners (kernel 6.17):
tcp_sequence()), its kernel then drops the server's end-of-window segments before processing the ACK they carry. The client never learns that the server reopened its window.The new client,
backpressure_oracle_linux_test.go(shared):rstoreof.SO_RCVBUFstays at 32 KiB and the flood still never reads, so the echo SEND still blocks. That is what the io_uring: WS recv-pause cancels by raw fd, killing the conn's in-flight SEND #482 test is for.The #607 class: what the drain hides, and how it is judged now
#607 was a recv the io_uring engine chained behind a send (
IOSQE_IO_LINK) on a detached connection. The send waited on the client's closed window, so the connection could not read at all. The drain above is exactly what reopens that window. So a client that drains gets through such a connection, and never gives up. Measured below: with #607 re-introduced, this PR's previous head passed 10 of 10 processes.Two checks now judge it directly.
WSO-STALL), both engines, both tests. Every 250 ms, for each connection's whole life, the rig reads the server side. A connection is one the engine should be reading when its server socket holds unread bytes, its chanReader is neither paused nor holding anything, and its handler has not exited. If the engine then reads nothing from it for 5 s, the test fails with that stretch's state at both ends, whatever the client does afterwards.watchResetson thewaits:line). It never happened in any run below; the longest watcher lag was 0.73 s.longest unread stretch, so a healthy run shows how close it came.RecvLinkedArms == 0, io_uring, both tests. Every connection on these servers is a detached WebSocket, detached by its upgrade before the engine flushes anything on it. A chained recv on one is io_uring: a linked recv never starts because its chained SEND blocks on a detached WebSocket peer's closed window #607's own mechanism.TestFlushSendLinkNeverChainsOnDetachedConnguards the function; this guards every path that reaches it.The timeline in a
WSO-GIVEUPrecordThe client sends its local port as
?cid=, and the handler reads it back withConn.Query. That joins each connection to its handler. Each sample reads:ReadMessage,WriteMessage, exited), frames echoed, time since the last echo.So the record shows the engine's recv and send progress, its resumes and its pending write bytes, the same way on epoll and io_uring. It names the shape it finds:
celeris#672 shape: paused with nothing buffered;celeris#705 shape: paused, channel empty, chunks in the spill;Server errors
An error is excused only on a connection whose client gave up or was reset, which already fails the test, and only after that client closed. Every other error is judged, whenever the handler noted it:
The excused ones are printed as "not judged, its client gave up".
The test binary's
-timeoutCI runs both oracles and the chanReader tests in one binary under
-timeout=300s, and a wedge costs at least 90 s per subtest. So every wait also ends at the binary's deadline less 60 s (reasonbudget, with itsWSO-GIVEUPrecord). A subtest that starts with less than 30 s of that left fails at once and says so, instead of starting a flood it cannot finish. A slow run that would have finished in time still does: the budget never ends a wait before the binary itself would have.TestBackpressurePauseDoesNotCancelInflightSendclientCloseFailis asserted (The WebSocket backpressure oracle counts clientCloseFail, prints it, and never asserts it: 66 of 96 connections failed and the test passed #623).protoErrandotherWriteErrkeep main's meaning: every error. The judged ones areprotoErrJudgedandotherWriteErrJudged; the excused areserverErrsExcused;recvStallscountsWSO-STALLrecords.waits:line:closeOver10s: the close waits the old 10 s deadline would have failed;closeSilentOver10s: those with a silent stretch over 10 s;recvStallMax,watchResets,watchMaxLag: the watch.TestBackpressureInboundSequenceIntegrity(#611)echoErrJudged,serverErrsExcused).floodDeadlinesandfloodDeadlineConnson the verdict line count the bursts a write deadline ended, so a run shows whether it took that path at all.clientCloseFailandclientRSTwere counted and never judged. Both are now asserted, with a record for each.writeAll, now used only by theceleris_closeprobe-tagged close-drain test, moves to a file under that tag.Negative controls
Every arm is a build with one change. The tally counts only
--- PASS/FAIL/SKIPlines, from the run artifacts, and one process is one observation. Unless a row says otherwise, the runs are on GitHub's runners in CI's shape:-race, 8 MiB memlock, the CI step's regex^(TestBackpressure|TestChanReader)and-timeout=300s, itsWS484_*env, count 1.#607 re-introduced (#623's own acceptance step)
m607 deletes
flushSendLink's Detached early-return (engine/iouring/worker.go), so a detached connection chains its RECV behind its SEND again. #623 asked for exactly this: run the oracle against a build with the #607 defect and record the tally, so the new assertion is known to fire on it. m607 changes io_uring only; the epoll subtest passed in every m607 process below.698bed6(the oracle before this PR)clientCloseFail19-55 of 96 connections in 6 of 10 processes, never judged.1a6a60f(this PR's previous head)LINKBLOCK blockedMaxMs4,000-6,492 in 9 of 10).1ee81fc(this head)WSO-STALLin 10/10 processes (511 records; per process the longest stretch was 6.0 s to 24.0 s), andRecvLinkedArms > 0in 10/10. Each check alone fails every process.The same arm on this round's earlier heads, with the same two checks: 36359217411 (
44ee852, on698bed6) failed 10/10 and 36356482425 (ad88696) 20/20, both checks firing in every process.So asserting
clientCloseFailwith the old client would have caught #607 in 6 of 10 processes; the new client alone catches it in none; the new client with the watch and the linked-recv check catches it in all.#672 re-introduced
m672 re-introduces #672: it deletes
requestPause's stale-pause re-check.On this round's final test code (
44ee852+ m672, the CI step's regex,-timeout=300s, 5 processes per arch, unless the row says otherwise):WS482_BP=16 WS484_BP=16WSO-GIVEUPrecords, and the inbound io_uring/multishot_recv subtest failed in 2/10 (one per arch) through its close wait, with a record and "handlers still running 20s after the clients finished". All 107 records areceleris#672 shape. Every process ended with its own verdicts inside the 300 s (the longest binary took 192 s): no goroutine dump.BP256)1ee81fc+ m672,WS482_BP=16, PauseCancel only,-timeout=150s, 3 processes per archbudget, the restidle; allceleris#672 shape), and the io_uring subtest then failed at once, saying how little of the-timeoutwas left. Every binary ended at 90-91 s of its 150 s, with its verdicts.So the inbound oracle's give-up assertions have now fired on an engine wedge, and a wedge in CI's shape ends as verdicts.
Round 0's arms (
WS482_BP=16, which makes the pause fire often; PauseCancel only; 5 processes per arch per arm):WSO-GIVEUPrecords, allceleris#672 shape: paused, channel empty, no spill, the kernel holding the unread bytes.closeTimeout; 4 of them (arm64 shards 1 and 2, x86 shards 1 and 4) also through theprotoErrassertion, the client-teardown RST echo this PR now excuses. 99 connections were counted asclientCloseFailand never judged. The old oracle prints no state, so that they were wedged is inferred from the new oracle's records on the same mutant, not shown. On x86 shard 3 the subtest PASSED with 3 of them.The io_uring subtest passed in every m672 process on both oracles, so m672 does not wedge io_uring in this workload. Under the old oracle, the laptop io_uring subtest still counted unjudged
clientCloseFailin two of five rounds (1 and 4 connections).An engine that loses a frame, and an echo failure (#611)
mdrop1 makes
chanReader.Appenddrop the bytes of exactly one whole frame, once per process, on the connection whose frames carry connection index 0, at the first frame boundary past 64 KiB: an engine that loses one delivered frame. inj611 is a test mutant: connection 0's echo write fails once it has echoed 200 frames, and the connection itself stays healthy. Both together,^TestBackpressureInboundSequenceIntegrity$, CI'sWS484_*env:44ee852+ mdrop1 + inj611conn 0 (127.0.0.1:<port>): frame count mismatch: sent=2000 delivered=1999, never delivered 1 (of the delivered, 1798 were read after the handler's echo failed), and a sequence gap. Every subtest judged the injected echo failure with its address:1 handler error(s) (1 echo write, 0 read).Round 0's inj611 arms (laptop Docker,
--cpus 4, 8 MiB memlock,-race, CI'sWS484_*env, 5 processes each):echoWriteErrors=1was printed and never judged.127.0.0.1:<port>(round 0's wording: "1 handler error(s) before their own client closed (1 echo write, 0 read)"). framesIn = framesSent (32000). 1,799 frames were read after the echo failed.The laptop series also repeated the m672 arms, 5 rounds each:
celeris#672 shapeWS482_BP=16Evidence:
evidence/celeris-b2/b2b/11-fix/(this round:01-m607/,02-dispatch/,branches.txt,scripts/mutate2.py,scripts/tally2.py) andevidence/celeris-b2/b2b/01-negative-control/(round 0).Verification (celeris-stress, tallied from the artifacts)
Which head each run used
1ee81fc: the branch rebased onto main0e239b1.git range-diffshows all three commits as=against13ca047/3421840/44ee852on698bed6. Main's new commits (fix(iouring): leave a claimed async conn to its claim when a reap is retried or lands (celeris#758) #765, fix(iouring): read an async conn's header deadline before its dispatch goroutine runs, under detachMu (celeris#722) #743, fix(server): a Start that never serves releases the caller's listener, the CPU monitor and the settle re-opener (#737) #747) change io_uring's async-mode reaps and header deadline andStart's failure path; these tests run the sync engines and aStartthat succeeds.tmp/b2b-r1-*, never PRs). Its runs used the heads below. Between them only the watch's limits, the budget arithmetic and the inbound frame count's exclusions moved; each row says which it had.5c7a324and1a6a60f: the client, not the watch.CI's shape, no mutant
^(TestBackpressure|TestChanReader),-race, 8 MiB memlock,-timeout=300s, CI'sWS484_*env, count 1, GitHub's runners (kernel 6.17):WSO-STALL1ee81fc(this head)44ee852(this head's tree on698bed6)d7f1122(this head's tree, except that an inbound close give-up stayed in the frame count)ad88696(asd7f1122, but the budget was half the time left)580cffe(asad88696, with a 6 s quiet; cancelled after 28 of 40 shards when the quiet was lengthened)watchResets0 everywhere).1ee81fc's run: no write wait ran past 13 s without progress (against 90 s); close waits took up to 12.8 s, the longest silence in one 12.8 s (the ~12 s close of io_uring: one connection in ~4 of 24 runs never completes the WebSocket Close handshake #566, against the 30 s give-up); PauseCancel subtests took 21 s at the median and 30 s at most; the longest binary took 58 s of its 300 s.The inbound test at its default size (no
WS484_*)96 connections x 4 bursts x 16,000 frames, both arches, 5 processes per arch,
-timeout=30m, headad88696(its inbound test is this head's, except that a close give-up now also leaves the frame count):-racefloodDeadlines(bursts a write deadline ended)inq=0) but advertised a 2 KiB window, and the client still held 245-273 KB unsent, receive-window-limited for 238-245 s. Filed as WebSocket inbound oracle at its default size under -race: on x86 epoll one connection's client-to-server stream crawls at a 2 KiB server window for minutes, and its close misses the 240 s cap (3 of 5 processes) #783, with what is not known yet (kernel memory pressure is one candidate).698bed6), through its fixed 30 s close budget and with no state printed. It also took 150-290 bursts per subtest out of its frame count.-racethe healthy epoll engine left a connection unread for up to 3.9 s (the watch's limit is 5 s). Without-race, 1.2 s.Round 0's runs (the client, before the watch)
V1, V2, the round-0 negative controls and the H3 diagnostic ran on
5c7a324(on main3fe9620); V4-V13 on1a6a60f(the same commits rebased onto698bed6, range-diff=).Every run is
^TestBackpressurewith CI's inputs:-race, 8 MiB memlock,WS484_*, count 1, on GitHub's runners (kernel 6.17). Everything is tallied from the run artifacts.5c7a324,1a6a60f5c7a324+ baked coveragecomplete, and no runner was lost. Per arch:clientCloseFail,closeTimeout,clientRST, judged server errors,ecanceled, sequence gaps, parse errors, overflows.-covermode. It is the PR head withgo tool cover -mode=atomicbaked intomiddleware/websocket's 15 non-test files (tmp/b2b-633-cov), the epoll: a rare close-handshake stall with the receive queue already drained, 1 in 73 (split from #607) #633 investigation's method. With main's oracle, the investigation's coverage reproducer (C1) failed 10/10 processes per arch.The whole websocket suite (round 0), base
3fe9620against head5c7a324, laptop Docker,-race, CI env:Pre-registration:
evidence/celeris-b2/b2b/03-stress/10-PREREGISTRATION.md, hashed before any of these runs finished.#633
#633 stays open.
The original symptom is an epoll close handshake without coverage that did not complete within the old absolute 10 s wait (1 in 73 processes on the laptop, once in CI).
The rule. I pre-registered that #633 is closed by this PR only if the no-coverage runs show 0 such waits in at least 218 processes per arch. That is what excludes 1 in 73 at p < 0.05 on each arch.
The runs have them. 220 processes per arch:
io_uring had none.
Every one of them was still receiving bytes. The longest silence in any close wait was 6.9 s, and each ended in a clean close. The old oracle would have failed them. The new one passes them and logs the latency (
closeOver10sandcloseSilentOver10sin thewaits:line). So the phenomenon is not gone.Conclusion. #633's epoll close-timeout is the slow-drain artifact the investigation called H1, and it happens without coverage and without instrumentation too. It is not a stall:
closeSilentOver10sis 0 in all 440 processes. A real stall now arrives as a give-up with both ends' state, and an engine that stops reading a connection as aWSO-STALL.Did it become more frequent after 9f4d89b? No rise shown.
evidence/celeris-b2/b2b/10-slowclose-bisect/). Failed processes:9f4d89b49d2726)bb231d5)9f4d89balready fails the old 10 s wait on x86 epoll in about 3% of iterations.On this evidence #633 can be closed by the maintainer; this PR does not close it.
#716 item 2: the engine-side stall on
pausedMuThis PR does not change it. I measured it, with the decision rule written and hashed before the first run (
evidence/celeris-b2/b2b/05-716-stall/10-PREREGISTRATION.md).How. A throwaway branch (
tmp/b2b-716-stall, main + instrumentation) records these inrequestPauseandresumeIfDrained:pausedMu, timed exactly (the TryLock fast path is left alone otherwise);ResumeRecvcall made under the lock.It ran on GitHub runners (kernel 6.17, 8 MiB memlock), without
-race. Workloads: both backpressure tests, the inbound one at its full default size, on both engines. Sync and async mode, 5 processes per arch per mode. Runs 36339356834 (sync) and 36339367373 (async).The engine side. This is the event-loop thread, or in async mode the goroutine that runs the WebSocket data callback.
requestPausecallspausedMuheldRule (pre-registered): a fix is needed if any cell has p99.9 ≥ 1 ms, a single wait ≥ 10 ms, or ≥ 1% of worker time. None met it, so no v1.6.0 fix.
One pre-registered prediction failed: "max < 1 ms". The measured max is 3.84 ms.
The exposure is real, and #716 argued it right. This was reported, not judged. The handler holds
pausedMuacrossResumeRecvfor as long as scheduling takes, and that hold has a heavy tail in this in-process load (96 handlers and 96 clients on 4 vCPUs):Its p99 is at most 16 µs everywhere, except arm64 io_uring PauseCancel (at most 262 µs sync, 524 µs async).
A
requestPausethat lands inside such a hold stalls the loop for the rest of it. It rarely lands there, because the engine takes a connection'spausedMuonly when that same connection crosses high water. #716's design (store the desired state under the lock, then queue and Signal after unlocking) removes the exposure if a workload ever shows it.Numbers:
evidence/celeris-b2/b2b/05-716-stall/20-analysis.txtand21-predictions-vs-outcome.txt. Also posted on #716.The io_uring/epoll asymmetry of the lost-window-update stall
Question. Before this PR, uninstrumented main lost the client's window update on io_uring and never on epoll:
clientCloseFailin 18/100 against 0/100 processes (#633). Why?Method. A diagnostic arm reruns the new oracle with its drain switched off. That is
tmp/b2b-h3-diag,CELERIS_B2B_NODRAIN=1: the old client never read while it waited. Run 36340735013: GitHub, CI shape, 20 processes per arch. It adds, for every connection, both sockets' segment statistics at the end of the flood (H3SEG, per socket), and the deltas of/proc/net/netstat's TcpExt counters per subtest (per network namespace, see below).It reproduces the asymmetry, now with the timeline:
All 21 are the lost-window-update shape, and the engine is healthy in every one:
ReadMessage, nothing buffered, never paused (pauses=0);pendingWrite=0);The client sits at a zero window both ways, with 57-58 KB unread. At the give-up its last processed ACK was 97 s old, which is older than the 90 s wait: it had stopped processing ACKs before the flood ended.
The difference is segment size. io_uring's echo leaves in smaller segments, and the connections that stall are the ones with the smallest.
The kernel's receive-memory counters move with it. These are
/proc/net/netstatdeltas per subtest, and that file counts the whole network namespace: client and server share one loopback kernel there, and the server's receive queues are flooded at the same time. So they cannot be attributed to the client's sockets. Per subtest, io_uring against epoll, summed over 20 processes:PruneCalledTCPRcvCollapsedBeyondWindowProposed mechanism, not isolated. More, smaller segments exhaust the client's 32 KiB receive memory before its advertised window is used up. The kernel prunes and collapses the queue, and then retracts its window. With its queue non-empty, the 6.17
tcp_sequence()check then drops the server's next end-of-window segment beforetcp_ack()runs, so the ACK that would reopen the client's send window is lost for good. The per-socket evidence shows the two ends of that chain: the stalled connections are the ones with the smallest server segments, and each ends at a zero window both ways with 57-58 KB unread. The steps between are not isolated. The counters are consistent with them, but they are namespace-wide:BeyondWindowcounts such drops on any socket, the server's included.Verdict. So the asymmetry goes with the size of io_uring's sends, meeting a Linux ≥ 6.17 receiver behaviour; no engine state was found wedged in any of the 21. This PR removes it from the oracle, because the client reads while it waits (see V1: 0 give-ups).
Not measured here:
The kernel side has an upstream fix (0e24d17bd966) that is in no release tag yet.
Evidence:
evidence/celeris-b2/b2b/04-h3-diag/20-analysis.txt.Not in this PR
./middleware/websocketfromtest-coverage.yml(ci: set up CodeRabbit, Codecov, CodSpeed (celeris#690) #699). That waits for this to land and for a coverage run with 0 give-ups and 0WSO-STALL(the coverage arm above predates the watch and the 8 s quiet, which were not measured under coverage).-raceon x86 epoll (WebSocket inbound oracle at its default size under -race: on x86 epoll one connection's client-to-server stream crawls at a 2 KiB server window for minutes, and its close misses the 240 s cap (3 of 5 processes) #783, above).go vet -tags celeris_closeprobealready fails on main (readerPaused redeclared). fix(websocket): never strand a chunk that spills after the handler drained the channel (celeris#705); fail the start helpers fast with Start's error (celeris#706) #730 fixes that.Closes #623
Closes #611
Refs #633, #607, #716, #699, #783