Skip to content

test(websocket): the backpressure oracles wait on progress and read while they wait, judge every give-up with both ends' timeline, and fail a connection the engine stops reading (celeris#633, celeris#623, celeris#611, celeris#607 class) - #749

Merged
FumingPower3925 merged 4 commits into
mainfrom
test/celeris-633-ws-backpressure-oracle
Sep 28, 2026

Conversation

@FumingPower3925

@FumingPower3925 FumingPower3925 commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Summary

What changes

Client: waits on progress (both tests)

The #633 investigation measured what the old fixed deadlines judged on GitHub's runners (kernel 6.17):

  • The server was behind, not stuck. 457 of 459 coverage close-timeouts were still receiving bytes when the 10 s ran out.
  • The client stalled itself. It never read while it waited, so its receive queue stayed full. Since Linux 6.17 (tcp_sequence()), its kernel then drops the server's end-of-window segments before processing the ACK they carry. The client never learns that the server reopened its window.
  • The counted server errors were the client's own give-up echoing back. A socket closed with unread data sends an RST. 1,165 of 1,165 such errors came after their own client's close.

The new client, backpressure_oracle_linux_test.go (shared):

  • It writes its last frame and its Close frame in 1 s slices.
  • After every slice that times out, it empties its receive queue without blocking. A drain that finds the connection reset or closed ends the wait as rst or eof.
  • It gives up only after 90 s with no progress (nothing written, and its send queue did not shrink), or at 240 s in total.
  • The close wait is re-armed by every byte received. It gives up after 30 s of silence, or at 240 s.
  • SO_RCVBUF stays at 32 KiB and the flood still never reads, so the echo SEND still blocks. That is what the io_uring: WS recv-pause cancels by raw fd, killing the conn's in-flight SEND #482 test is for.

The #607 class: what the drain hides, and how it is judged now

#607 was a recv the io_uring engine chained behind a send (IOSQE_IO_LINK) on a detached connection. The send waited on the client's closed window, so the connection could not read at all. The drain above is exactly what reopens that window. So a client that drains gets through such a connection, and never gives up. Measured below: with #607 re-introduced, this PR's previous head passed 10 of 10 processes.

Two checks now judge it directly.

  • The watch (WSO-STALL), both engines, both tests. Every 250 ms, for each connection's whole life, the rig reads the server side. A connection is one the engine should be reading when its server socket holds unread bytes, its chanReader is neither paused nor holding anything, and its handler has not exited. If the engine then reads nothing from it for 5 s, the test fails with that stretch's state at both ends, whatever the client does afterwards.
    • The proof is two readings, not the samples between them. Bytes read, pauses and resumes only ever grow, so equal readings at both ends mean the engine read nothing in between while nothing asked it to stop.
    • A process starved of CPU could fake that. So if the watcher itself runs more than 1 s after its tick was due, every open stretch starts over (watchResets on the waits: line). It never happened in any run below; the longest watcher lag was 0.73 s.
    • The healthy ceiling, measured in CI's shape on GitHub's runners over 218 subtest runs per engine: 2.3 s on epoll, 1.7 s on io_uring. On epoll the longest ones come while other connections keep the loop busy (the flood, and the drains after the quiet); they are logged, not judged.
    • The io_uring: WS recv-pause cancels by raw fd, killing the conn's in-flight SEND #482 client now holds its backpressure idle for 8 s after the flood, reading and writing nothing. A healthy engine reads what the socket holds, or pauses. An engine that stopped reading the connection behind its blocked echo cannot, so that stretch outlasts 5 s however late in the flood it began. (With 6 s of quiet, the longest io_uring: a linked recv never starts because its chained SEND blocks on a detached WebSocket peer's closed window #607 stretch in four x86 processes was only 4.5 s.)
    • Every subtest logs its longest unread stretch, so a healthy run shows how close it came.
  • RecvLinkedArms == 0, io_uring, both tests. Every connection on these servers is a detached WebSocket, detached by its upgrade before the engine flushes anything on it. A chained recv on one is io_uring: a linked recv never starts because its chained SEND blocks on a detached WebSocket peer's closed window #607's own mechanism. TestFlushSendLinkNeverChainsOnDetachedConn guards the function; this guards every path that reaches it.

The timeline in a WSO-GIVEUP record

The client sends its local port as ?cid=, and the handler reads it back with Conn.Query. That joins each connection to its handler. Each sample reads:

  • The client socket: send and receive queue, both windows, RTO backoff.
  • The handler: where it is (ReadMessage, WriteMessage, exited), frames echoed, time since the last echo.
  • The chanReader: depth, spill, whether it holds the engine paused, and the pauses and resumes it applied. The callbacks are wrapped to count them.
  • The server socket, read through the engine's own descriptor with read-only calls: received and still unread bytes, bytes the kernel accepted and the peer ACKed, windows, retransmissions.
  • Derived from these: the bytes the engine has read, and the handler's bytes the engine has not yet handed to the kernel. Both count from the connection's first byte, net of the upgrade request and the 101, whose sizes the client reads from its own socket once the handshake is done. (A baseline sampled when the handler started could miss a 101 still queued.)

So the record shows the engine's recv and send progress, its resumes and its pending write bytes, the same way on epoll and io_uring. It names the shape it finds:

  • celeris#672 shape: paused with nothing buffered;
  • celeris#705 shape: paused, channel empty, chunks in the spill;
  • socket readable and nothing delivered;
  • handler bytes never handed to the kernel;
  • lost window update;
  • server behind.

Server errors

An error is excused only on a connection whose client gave up or was reset, which already fails the test, and only after that client closed. Every other error is judged, whenever the handler noted it:

  • A client that read the server's close (EOF) had an empty receive queue, and its close sent a FIN, so nothing it did explains a server error.
  • A handler notes an engine error only when its next read or write returns, so the time it was noted says nothing about its cause.

The excused ones are printed as "not judged, its client gave up".

The test binary's -timeout

CI runs both oracles and the chanReader tests in one binary under -timeout=300s, and a wedge costs at least 90 s per subtest. So every wait also ends at the binary's deadline less 60 s (reason budget, with its WSO-GIVEUP record). A subtest that starts with less than 30 s of that left fails at once and says so, instead of starting a flood it cannot finish. A slow run that would have finished in time still does: the budget never ends a wait before the binary itself would have.

TestBackpressurePauseDoesNotCancelInflightSend

TestBackpressureInboundSequenceIntegrity (#611)

  • After an echo write fails, the handler stops echoing and keeps reading. Frames read after the failure are counted separately.
  • A frame-count mismatch now names the connection's address, and splits the frames into never delivered and read after the echo failed.
  • Echo and read errors are judged by the rule above, with the error and the address (echoErrJudged, serverErrsExcused).
  • The client gets the same wait/read pattern, and the watch.
  • clientCloseFail and clientRST were counted and never judged. Both are now asserted, with a record for each.
  • writeAll, now used only by the celeris_closeprobe-tagged close-drain test, moves to a file under that tag.

Negative controls

Every arm is a build with one change. The tally counts only --- PASS/FAIL/SKIP lines, from the run artifacts, and one process is one observation. Unless a row says otherwise, the runs are on GitHub's runners in CI's shape: -race, 8 MiB memlock, the CI step's regex ^(TestBackpressure|TestChanReader) and -timeout=300s, its WS484_* env, count 1.

#607 re-introduced (#623's own acceptance step)

m607 deletes flushSendLink's Detached early-return (engine/iouring/worker.go), so a detached connection chains its RECV behind its SEND again. #623 asked for exactly this: run the oracle against a build with the #607 defect and record the tally, so the new assertion is known to fire on it. m607 changes io_uring only; the epoll subtest passed in every m607 process below.

build + m607 run io_uring subtest what the oracle saw
main 698bed6 (the oracle before this PR) 36353889731 PASSED 10/10 clientCloseFail 19-55 of 96 connections in 6 of 10 processes, never judged.
1a6a60f (this PR's previous head) 36353899414 PASSED 10/10 Every connection finished: the drain got it through, while the engine held a chained recv for up to 6.5 s (LINKBLOCK blockedMaxMs 4,000-6,492 in 9 of 10).
1ee81fc (this head) 36359543720 FAILED 10/10 WSO-STALL in 10/10 processes (511 records; per process the longest stretch was 6.0 s to 24.0 s), and RecvLinkedArms > 0 in 10/10. Each check alone fails every process.

The same arm on this round's earlier heads, with the same two checks: 36359217411 (44ee852, on 698bed6) failed 10/10 and 36356482425 (ad88696) 20/20, both checks firing in every process.

So asserting clientCloseFail with the old client would have caught #607 in 6 of 10 processes; the new client alone catches it in none; the new client with the watch and the linked-recv check catches it in all.

#672 re-introduced

m672 re-introduces #672: it deletes requestPause's stale-pause re-check.

On this round's final test code (44ee852 + m672, the CI step's regex, -timeout=300s, 5 processes per arch, unless the row says otherwise):

run env result
36358641405 CI's, plus WS482_BP=16 WS484_BP=16 FAILED 10/10. PauseCancel epoll failed in 10/10 with 105 WSO-GIVEUP records, and the inbound io_uring/multishot_recv subtest failed in 2/10 (one per arch) through its close wait, with a record and "handlers still running 20s after the clients finished". All 107 records are celeris#672 shape. Every process ended with its own verdicts inside the 300 s (the longest binary took 192 s): no goroutine dump.
36358019325 CI's (BP 256) passed 10/10: at the default buffer m672 does not wedge in 10 processes, which is why round 0 and the row above use 16.
36359976973 1ee81fc + m672, WS482_BP=16, PauseCancel only, -timeout=150s, 3 processes per arch FAILED 6/6, the budget's own control. The epoll waits gave up at the budget (61 of 69 records say budget, the rest idle; all celeris#672 shape), and the io_uring subtest then failed at once, saying how little of the -timeout was left. Every binary ended at 90-91 s of its 150 s, with its verdicts.

So the inbound oracle's give-up assertions have now fired on an engine wedge, and a wedge in CI's shape ends as verdicts.

Round 0's arms (WS482_BP=16, which makes the pause fire often; PauseCancel only; 5 processes per arch per arm):

arm build epoll subtest what the oracle saw
new oracle + m672 36342287418 FAILED 10/10 (x86 5/5, arm64 5/5) 112 WSO-GIVEUP records, all celeris#672 shape: paused, channel empty, no spill, the kernel holding the unread bytes.
old oracle (main) + m672 36342301050 FAILED 9/10 All 9 through closeTimeout; 4 of them (arm64 shards 1 and 2, x86 shards 1 and 4) also through the protoErr assertion, the client-teardown RST echo this PR now excuses. 99 connections were counted as clientCloseFail and never judged. The old oracle prints no state, so that they were wedged is inferred from the new oracle's records on the same mutant, not shown. On x86 shard 3 the subtest PASSED with 3 of them.

The io_uring subtest passed in every m672 process on both oracles, so m672 does not wedge io_uring in this workload. Under the old oracle, the laptop io_uring subtest still counted unjudged clientCloseFail in two of five rounds (1 and 4 connections).

An engine that loses a frame, and an echo failure (#611)

mdrop1 makes chanReader.Append drop the bytes of exactly one whole frame, once per process, on the connection whose frames carry connection index 0, at the first frame boundary past 64 KiB: an engine that loses one delivered frame. inj611 is a test mutant: connection 0's echo write fails once it has echoed 200 frames, and the connection itself stays healthy. Both together, ^TestBackpressureInboundSequenceIntegrity$, CI's WS484_* env:

run build result
36358416945 44ee852 + mdrop1 + inj611 FAILED 10/10 (x86 5/5, arm64 5/5). In every process the epoll subtest (the first to carry connection 0 past 64 KiB; mdrop1 fires once per process) printed conn 0 (127.0.0.1:<port>): frame count mismatch: sent=2000 delivered=1999, never delivered 1 (of the delivered, 1798 were read after the handler's echo failed), and a sequence gap. Every subtest judged the injected echo failure with its address: 1 handler error(s) (1 echo write, 0 read).

Round 0's inj611 arms (laptop Docker, --cpus 4, 8 MiB memlock, -race, CI's WS484_* env, 5 processes each):

arm result
old inbound oracle + inj611 PASSED 5/5. 578-800 of connection 0's frames were never read (framesIn 30201 of 30779-31001). The handler had returned, the server closed, the client's write failed, and the client was left out of the frame count. echoWriteErrors=1 was printed and never judged.
new inbound oracle + inj611 FAILED 5/5, naming the error and 127.0.0.1:<port> (round 0's wording: "1 handler error(s) before their own client closed (1 echo write, 0 read)"). framesIn = framesSent (32000). 1,799 frames were read after the echo failed.

The laptop series also repeated the m672 arms, 5 rounds each:

arm result
new oracle + m672 FAILED 5/5, with 26 records, all celeris#672 shape
old oracle + m672 FAILED 4/5. Round 4 PASSED, with 3 epoll connections that never finished their writes.
new oracle, no mutant, WS482_BP=16 PASSED 5/5, with 0 give-ups

Evidence: evidence/celeris-b2/b2b/11-fix/ (this round: 01-m607/, 02-dispatch/, branches.txt, scripts/mutate2.py, scripts/tally2.py) and evidence/celeris-b2/b2b/01-negative-control/ (round 0).

Verification (celeris-stress, tallied from the artifacts)

Which head each run used

CI's shape, no mutant

^(TestBackpressure|TestChanReader), -race, 8 MiB memlock, -timeout=300s, CI's WS484_* env, count 1, GitHub's runners (kernel 6.17):

run head processes (x86 + arm64) failed WSO-STALL give-ups
36359349677 1ee81fc (this head) 10 + 10 0 0 0
36358829347 44ee852 (this head's tree on 698bed6) 10 + 10 0 0 0
36357803529 d7f1122 (this head's tree, except that an inbound close give-up stayed in the frame count) 10 + 10 0 0 0
36356474960 ad88696 (as d7f1122, but the budget was half the time left) 20 + 20 0 0 0
36355864370 580cffe (as ad88696, with a 6 s quiet; cancelled after 28 of 40 shards when the quiet was lengthened) 20 + 8 0 0 0
  • The watch's healthy margin, over 218 PauseCancel subtest runs per engine in CI's shape (these runs, plus 90 per engine from two earlier candidates with a 3 s limit): the longest unread stretch was 2.3 s on epoll and 1.7 s on io_uring, against the 5 s limit. At CI's inbound size no stretch appeared at all. The watcher never lagged by more than 0.73 s, so no stretch was ever reset (watchResets 0 everywhere).
  • The waits in 1ee81fc's run: no write wait ran past 13 s without progress (against 90 s); close waits took up to 12.8 s, the longest silence in one 12.8 s (the ~12 s close of io_uring: one connection in ~4 of 24 runs never completes the WebSocket Close handshake #566, against the 30 s give-up); PauseCancel subtests took 21 s at the median and 30 s at most; the longest binary took 58 s of its 300 s.
  • The 8 s quiet adds about 8 s to each PauseCancel subtest.

The inbound test at its default size (no WS484_*)

96 connections x 4 bursts x 16,000 frames, both arches, 5 processes per arch, -timeout=30m, head ad88696 (its inbound test is this head's, except that a close give-up now also leaves the frame count):

run -race failed floodDeadlines (bursts a write deadline ended)
36357515331 no 0/10 x86: epoll 79 in 60 connections, io_uring 61 in 22; arm64: 5 in 3; so the restored judging ran, and passed.
36357012593 yes 3/10, all x86 epoll in every process: 803-1,418 bursts per arch and engine

Round 0's runs (the client, before the watch)

V1, V2, the round-0 negative controls and the H3 diagnostic ran on 5c7a324 (on main 3fe9620); V4-V13 on 1a6a60f (the same commits rebased onto 698bed6, range-diff =).

Every run is ^TestBackpressure with CI's inputs: -race, 8 MiB memlock, WS484_*, count 1, on GitHub's runners (kernel 6.17). Everything is tallied from the run artifacts.

runs head processes failed give-ups
V1 36339704263 and V4-V13 (10 runs, below) 5c7a324, 1a6a60f 220 x86 + 220 arm64 0 / 440 0
V2 36339720476, coverage 5c7a324 + baked coverage 20 + 20 0 / 40 0
  • Without coverage. Every shard was complete, and no runner was lost. Per arch:
    • 21,120 PauseCancel connections per engine, all closed cleanly;
    • 3,520 inbound connections per engine variant;
    • 7.04 M frames sent and 7.04 M delivered per variant;
    • no counter above 0: clientCloseFail, closeTimeout, clientRST, judged server errors, ecanceled, sequence gaps, parse errors, overflows.
  • V4-V13 run ids: 36343721891, 36343736235, 36344778254, 36344794117, 36345879024, 36345892452, 36346940144, 36346955264, 36347961059, 36347977221.
  • The 95% upper bound on the per-process failure rate is 1.7% per arch.
  • The waits a slow server now gets:
    • frame completion up to 9.2 s;
    • the longest stretch with no progress 9.1 s, against a 90 s give-up;
    • close handshakes up to 47.6 s (x86 epoll), with the longest silence in any close wait 6.9 s.
  • V2 is the coverage stand-in, because the celeris-stress planner refuses -covermode. It is the PR head with go tool cover -mode=atomic baked into middleware/websocket's 15 non-test files (tmp/b2b-633-cov), the epoll: a rare close-handshake stall with the receive queue already drained, 1 in 73 (split from #607) #633 investigation's method. With main's oracle, the investigation's coverage reproducer (C1) failed 10/10 processes per arch.
    • It passes, even though coverage makes everything slow: frame completion took up to 47.7 s and close handshakes up to 170 s.
    • The longest silence in a close wait was 6.8 s, and the longest no-progress stretch 45.4 s, half the give-up.
    • 2,015 close waits went over 10 s, and the old oracle would have failed every one.
  • The cluster was not used. It was preferred, but the batch perf checkpoint held the cluster group from 19:15Z for about 4 h 50 min, with nine more rows queued behind it. The two queued cluster runs were withdrawn before dispatch, and the verification moved to GitHub's runners, the CI environment where epoll: a rare close-handshake stall with the receive queue already drained, 1 in 73 (split from #607) #633 was seen.

The whole websocket suite (round 0), base 3fe9620 against head 5c7a324, laptop Docker, -race, CI env:

memlock base (PASS/FAIL/SKIP) head (PASS/FAIL/SKIP) PASS→FAIL
8 MiB 236/0/2 236/0/2 0
128 MiB 236/0/2 236/0/2 0

Pre-registration: evidence/celeris-b2/b2b/03-stress/10-PREREGISTRATION.md, hashed before any of these runs finished.

#633

#633 stays open.

The original symptom is an epoll close handshake without coverage that did not complete within the old absolute 10 s wait (1 in 73 processes on the laptop, once in CI).

The rule. I pre-registered that #633 is closed by this PR only if the no-coverage runs show 0 such waits in at least 218 processes per arch. That is what excludes 1 in 73 at p < 0.05 on each arch.

The runs have them. 220 processes per arch:

arch epoll close waits over 10 s processes longest
x86 43 17/220 47.6 s
arm64 2 2/220 15.6 s

io_uring had none.

Every one of them was still receiving bytes. The longest silence in any close wait was 6.9 s, and each ended in a clean close. The old oracle would have failed them. The new one passes them and logs the latency (closeOver10s and closeSilentOver10s in the waits: line). So the phenomenon is not gone.

Conclusion. #633's epoll close-timeout is the slow-drain artifact the investigation called H1, and it happens without coverage and without instrumentation too. It is not a stall: closeSilentOver10s is 0 in all 440 processes. A real stall now arrives as a give-up with both ends' state, and an engine that stops reading a connection as a WSO-STALL.

Did it become more frequent after 9f4d89b? No rise shown.

  • Why I asked. With the old oracle at count 5, lane B-2a's runs failed x86 epoll in 4-11 of 20 processes. The investigation had 0/60 at count 1.
  • The test. A pre-registered A/B, old oracle, x86, 20 processes per commit, count 5 (evidence/celeris-b2/b2b/10-slowclose-bisect/). Failed processes:
commit failed p (two-sided Fisher) against 9f4d89b
9f4d89b 3/20
+ #674 (49d2726) 6/20 0.45
+ #671 (bb231d5) 4/20 1.0

On this evidence #633 can be closed by the maintainer; this PR does not close it.

#716 item 2: the engine-side stall on pausedMu

This PR does not change it. I measured it, with the decision rule written and hashed before the first run (evidence/celeris-b2/b2b/05-716-stall/10-PREREGISTRATION.md).

How. A throwaway branch (tmp/b2b-716-stall, main + instrumentation) records these in requestPause and resumeIfDrained:

  • every contended acquisition of pausedMu, timed exactly (the TryLock fast path is left alone otherwise);
  • every ResumeRecv call made under the lock.

It ran on GitHub runners (kernel 6.17, 8 MiB memlock), without -race. Workloads: both backpressure tests, the inbound one at its full default size, on both engines. Sync and async mode, 5 processes per arch per mode. Runs 36339356834 (sync) and 36339367373 (async).

The engine side. This is the event-loop thread, or in async mode the goroutine that runs the WebSocket data callback.

result
requestPause calls 19,806
found pausedMu held 3 (0.015%), all arm64 io_uring PauseCancel
those 3 waits 1.02 ms (sync), 3.84 ms and 0.41 ms (async)
p99.9 over all calls 0 in every cell
blocked time at most 2.1e-4 of worker time

Rule (pre-registered): a fix is needed if any cell has p99.9 ≥ 1 ms, a single wait ≥ 10 ms, or ≥ 1% of worker time. None met it, so no v1.6.0 fix.

One pre-registered prediction failed: "max < 1 ms". The measured max is 3.84 ms.

The exposure is real, and #716 argued it right. This was reported, not judged. The handler holds pausedMu across ResumeRecv for as long as scheduling takes, and that hold has a heavy tail in this in-process load (96 handlers and 96 clients on 4 vCPUs):

  • 262 ms max (arm64 epoll, inbound, sync);
  • 193 ms (x86 epoll, inbound, async);
  • 68 ms (x86 epoll, PauseCancel, sync).

Its p99 is at most 16 µs everywhere, except arm64 io_uring PauseCancel (at most 262 µs sync, 524 µs async).

A requestPause that lands inside such a hold stalls the loop for the rest of it. It rarely lands there, because the engine takes a connection's pausedMu only when that same connection crosses high water. #716's design (store the desired state under the lock, then queue and Signal after unlocking) removes the exposure if a workload ever shows it.

Numbers: evidence/celeris-b2/b2b/05-716-stall/20-analysis.txt and 21-predictions-vs-outcome.txt. Also posted on #716.

The io_uring/epoll asymmetry of the lost-window-update stall

Question. Before this PR, uninstrumented main lost the client's window update on io_uring and never on epoll: clientCloseFail in 18/100 against 0/100 processes (#633). Why?

Method. A diagnostic arm reruns the new oracle with its drain switched off. That is tmp/b2b-h3-diag, CELERIS_B2B_NODRAIN=1: the old client never read while it waited. Run 36340735013: GitHub, CI shape, 20 processes per arch. It adds, for every connection, both sockets' segment statistics at the end of the flood (H3SEG, per socket), and the deltas of /proc/net/netstat's TcpExt counters per subtest (per network namespace, see below).

It reproduces the asymmetry, now with the timeline:

give-ups processes
io_uring, x86 19 8/20
io_uring, arm64 2 2/20
epoll, both arches 0 0/40

All 21 are the lost-window-update shape, and the engine is healthy in every one:

  • handler idle in ReadMessage, nothing buffered, never paused (pauses=0);
  • every handler byte handed to the kernel (pendingWrite=0);
  • server receive queue empty and window open;
  • the server holds 1 unacked segment, at RTO backoff 8 at the give-up.

The client sits at a zero window both ways, with 57-58 KB unread. At the give-up its last processed ACK was 97 s old, which is older than the 90 s wait: it had stopped processing ACKs before the flood ended.

The difference is segment size. io_uring's echo leaves in smaller segments, and the connections that stall are the ones with the smallest.

x86 epoll io_uring (all) io_uring (stalled)
server bytes per data segment, p50 2,126 856 495 (min 232)
data segments at flood end, p50 31 58 118

The kernel's receive-memory counters move with it. These are /proc/net/netstat deltas per subtest, and that file counts the whole network namespace: client and server share one loopback kernel there, and the server's receive queues are flooded at the same time. So they cannot be attributed to the client's sockets. Per subtest, io_uring against epoll, summed over 20 processes:

counter x86 arm64
PruneCalled 2.6x (268k vs 102k) 9.2x
TCPRcvCollapsed 3.2x 10.4x
BeyondWindow 6.1x (404 vs 66) 3.0x

Proposed mechanism, not isolated. More, smaller segments exhaust the client's 32 KiB receive memory before its advertised window is used up. The kernel prunes and collapses the queue, and then retracts its window. With its queue non-empty, the 6.17 tcp_sequence() check then drops the server's next end-of-window segment before tcp_ack() runs, so the ACK that would reopen the client's send window is lost for good. The per-socket evidence shows the two ends of that chain: the stalled connections are the ones with the smallest server segments, and each ends at a zero window both ways with 57-58 KB unread. The steps between are not isolated. The counters are consistent with them, but they are namespace-wide: BeyondWindow counts such drops on any socket, the server's included.

Verdict. So the asymmetry goes with the size of io_uring's sends, meeting a Linux ≥ 6.17 receiver behaviour; no engine state was found wedged in any of the 21. This PR removes it from the oracle, because the client reads while it waits (see V1: 0 give-ups).

Not measured here:

  • why io_uring's detached-connection sends are smaller;
  • whether a real client that writes a large burst without reading, against io_uring, can hit the same kernel stall.

The kernel side has an upstream fix (0e24d17bd966) that is in no release tag yet.

Evidence: evidence/celeris-b2/b2b/04-h3-diag/20-analysis.txt.

Not in this PR

Closes #623
Closes #611
Refs #633, #607, #716, #699, #783

@FumingPower3925 FumingPower3925 added this to the v1.6.0 milestone Sep 27, 2026
@FumingPower3925 FumingPower3925 added bug Something isn't working testing Testing infrastructure and helpers middleware Middleware implementation labels Sep 27, 2026
@coderabbitai

coderabbitai Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Important

Review skipped

Review was skipped as selected files did not have any reviewable changes.

⚙️ Run configuration

Configuration used: Repository: goceleris/celeris/.coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: 5bd21c4a-2806-40dd-8094-eebf5fb16938

📥 Commits

Reviewing files that changed from the base of the PR and between 1ee81fc and d7c3e42.

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: goceleris/celeris/.coderabbit.yaml

Review profile: CHILL

Plan: Advanced

Run ID: 38efbe1c-4b15-4729-a782-f204096ce0a2

📥 Commits

Reviewing files that changed from the base of the PR and between 0e239b1 and 1ee81fc.

📒 Files selected for processing (4)
  • middleware/websocket/backpressure_oracle_linux_test.go
  • middleware/websocket/inbound_sequence_linux_test.go
  • middleware/websocket/pause_cancel_linux_test.go
  • middleware/websocket/writeall_closeprobe_linux_test.go

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 9 remain after this review.


📝 Walkthrough

Walkthrough

Linux WebSocket backpressure tests now share bounded wait and connection-diagnostic support. The inbound-sequence test distinguishes delivered frames from frames not delivered after echo writes fail. The pause-cancel test tracks client writes and close handshakes and asserts handler-error and receive-stall results.

Changes

WebSocket test oracles

Layer / File(s) Summary
Shared backpressure oracle
middleware/websocket/backpressure_oracle_linux_test.go
Adds bounded, progress-sensitive client writes and close waits, connection timelines, server-error classification, unread-socket stall detection, and wait statistics.
Inbound frame delivery
middleware/websocket/inbound_sequence_linux_test.go
The handler continues reading and counting frames after an echo write fails. It records delivered frames separately from frames not delivered.
Inbound flood and close checks
middleware/websocket/inbound_sequence_linux_test.go
Client flood writes resume after deadline-ended bursts. Tracked close waits and assertions report connection-level delivery, server errors, stalls, and linked-receive metrics.
Pause-cancel handler tracking
middleware/websocket/pause_cancel_linux_test.go
The handler records connection phases and errors through the rig. The test registers the server endpoint and checks linked receives.
Pause-cancel client lifecycle
middleware/websocket/pause_cancel_linux_test.go, middleware/websocket/writeall_closeprobe_linux_test.go
Clients use tracked writes and close waits. The test checks handler errors, receive stalls, unfinished frames, and close outcomes; the build-tagged helper provides writeAll for close-probe builds.

Priority: ⬇️ Low

Estimated code review effort: 4 (Complex) | ~60 minutes

Change: Other · Severity of issue fixed: Low

Merge Risk: ⚪ Minimal · up to 1ee81

This change strengthens the Linux WebSocket backpressure tests and does not modify product code. No concrete regressions remain, so it is ready to merge after normal CI.

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title uses the required test(websocket): summary format, clearly describes the backpressure-oracle changes, and includes issue references.
Description check ✅ Passed The description directly explains the WebSocket test changes, linked issues, verification results, and scope.
Linked Issues check ✅ Passed #623: pause_cancel_linux_test.go now increments clientCloseFail, keeps it separate from closeTimeout, and fails the test when the counter is non-zero. The give-up record includes both connection…
Out of Scope Changes check ✅ Passed The changes are Linux test code and test-only helpers. The shared rig, progress-based waits, socket timelines, receive-stall checks, and io_uring receive-link checks support reliable backpressure orac…

Comment @coderabbitai help to get the list of available commands.

@codecov

codecov Bot commented Sep 27, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

…, with waits that give up on no progress and both ends' timeline (celeris#633, celeris#623)

TestBackpressurePauseDoesNotCancelInflightSend finished its last frame and its
Close frame with fixed 15 s and 10 s deadlines, waited a fixed 10 s for the
server's close, never read while it waited, and counted a connection that
never finished its writes (clientCloseFail) without judging it. celeris#633
measured what those deadlines judged on GitHub's runners: a server that was
behind but moving (457 of 459 close-timeouts were still receiving bytes), a
client that stalled itself because its full receive queue makes Linux 6.17
discard the server's ACKs, and server errors that were the echo of the
client's own give-up (1,165 of 1,165 after that client's close).

The client now writes in 1 s slices, empties its receive queue without
blocking after every slice that times out, and gives up only after 90 s with
no progress (no byte written, no shrink of its send queue) or at 240 s; the
close wait is re-armed by every byte and gives up after 30 s of silence or at
240 s. SO_RCVBUF stays 32 KiB and the flood never reads, so the echo SEND
still blocks, which is what the test is for.

In the same change, clientCloseFail is asserted (celeris#623), and a server
error is judged only when it came before its own client closed. Every give-up
prints a WSO-GIVEUP record: the connection's milestones and a 1 s timeline of
both ends (client socket; handler phase and echoes; chanReader depth, spill,
pause state and the pauses and resumes it applied; the server socket read
through the engine's descriptor: bytes read by the engine, bytes accepted by
the kernel and ACKed, the handler bytes the engine has not handed to the
kernel), and the shape it stalled in. The client joins its handler through
?cid=<its port> and Conn.Query.
…p, and judge the echo failure with its connection (celeris#611)

TestBackpressureInboundSequenceIntegrity's handler returned when its echo
write failed, leaving frames the engine had delivered unread, and the test
then reported them lost (99.8-99.9% of the "missing" frames were still queued
at handler exit); the echo failure itself was logged and never judged.

The handler now stops echoing and keeps reading to the end of the stream, so
the frame count measures delivery; frames read after the echo failed are
counted separately, and a mismatch names the connection and both populations.
An echo write or read error before its own client closed fails the test with
the error and the connection's address; after the client's close it is that
client's teardown and is printed, not judged.

Its client gets the same wait/read pattern as the #482 oracle's (celeris#633):
a flood write that times out is the backpressure the test creates and ends
the burst, instead of taking the connection out of the frame-count verdict;
the last frame and the Close frame are written on progress with drains
between slices; the close wait is re-armed per byte. A client that cannot
finish, and a close that arrives as an RST, now fail the test with the
connection's timeline, as clientCloseFail and clientRST were counted and not
judged.

writeAll, which only the celeris_closeprobe-tagged close-drain test still
uses, moves to a file under that tag, so the default build does not carry
an unused helper.
…e every server error on a connection that closed cleanly (celeris#607 class, celeris#623)

The progress-based client of the previous commits empties its receive queue
between write slices. That is what reopens a client window a blocked server
SEND waits on, so it also gets through a connection whose engine chained its
RECV behind that SEND (celeris#607). Measured in CI's shape on GitHub's
runners with #607 re-introduced (flushSendLink's Detached early-return
deleted), 5 processes per arch: the #482 oracle before this PR counted such
connections as clientCloseFail (19 to 55 of 96, in 6 of 10 processes) and
never judged them; this PR's previous head finished every one of them (0 of
10 processes failed) while the engine held a chained recv for up to 6.5 s.

So the rig now watches the server side of every connection every 250 ms for
its whole life. A connection whose server socket holds unread bytes, whose
chanReader is neither paused nor holding anything, and whose handler has not
exited is one the engine should be reading. If the engine reads nothing from
it for 5 s the test fails with that stretch's state at both ends (WSO-STALL),
whatever the client does afterwards. Bytes read, pauses and resumes only
grow, so equal readings at a stretch's two ends prove it; a watcher run more
than 1 s after its tick was due means the process was starved, and every open
stretch starts over. The healthy maximum measured in CI's shape was 2.3 s
(epoll) and 1.7 s (io_uring). After its flood the #482 client now holds its
backpressure idle for 8 s, so a stretch that begins late in the flood still
reaches the limit. On io_uring both oracles also assert RecvLinkedArms == 0:
every connection on these servers is a detached WebSocket, and a chain on one
is #607's own mechanism.

Server errors: an error is excused only on a connection whose client gave up
or was reset (which already fails the test) and only after that client
closed. An error on a connection whose client read the server's close is
judged whenever the handler noted it: that client's queue was empty and its
close sent a FIN, and an engine error is noted only when the handler next
returns from a read or write.

Also:
  - every wait also ends at the test binary's deadline less 60 s, and a
    subtest that starts with less than 30 s of that left fails at once, so a
    wedge under CI's -timeout=300s ends as each subtest's verdict and
    WSO-GIVEUP record, not a goroutine dump, while a slow run that would
    have finished in time still does;
  - the inbound verdict line counts bursts a flood write deadline ended
    (floodDeadlines, floodDeadlineConns), and a connection whose close wait
    gave up or was reset leaves the frame count, as a failed write already
    did: the frames still in its client's send queue were never the
    engine's to deliver, and the give-up fails the test with its record;
  - the timeline's engine-read and pending-write bytes count from the
    connection's first byte net of the upgrade, whose sizes the client reads
    from its own socket, not from a baseline sampled after the 101 was
    queued; a drain that finds the connection reset or closed ends the wait
    as "rst" or "eof";
  - PauseCancel's protoErr and otherWriteErr keep main's meaning (every
    error); the judged ones are protoErrJudged and otherWriteErrJudged; the
    waits line adds recvStallMax, watchResets and watchMaxLag, and a
    "longest unread stretch" line shows how close a run came to the limit.
@FumingPower3925
FumingPower3925 force-pushed the test/celeris-633-ws-backpressure-oracle branch from 1a6a60f to 1ee81fc Compare September 27, 2026 23:37
@FumingPower3925 FumingPower3925 changed the title test(websocket): the backpressure oracles wait on progress, read while they wait, and judge every give-up with both ends' timeline (celeris#633, celeris#623, celeris#611) test(websocket): the backpressure oracles wait on progress and read while they wait, judge every give-up with both ends' timeline, and fail a connection the engine stops reading (celeris#633, celeris#623, celeris#611, celeris#607 class) Sep 27, 2026
@FumingPower3925
FumingPower3925 marked this pull request as ready for review September 28, 2026 01:52
@FumingPower3925
FumingPower3925 merged commit 6ec8224 into main Sep 28, 2026
17 checks passed
@FumingPower3925
FumingPower3925 deleted the test/celeris-633-ws-backpressure-oracle branch September 28, 2026 02:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working middleware Middleware implementation testing Testing infrastructure and helpers

Projects

None yet

1 participant