You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Split out of #607 so it is not closed along with it. #607's fix touches engine/iouring only and does not address this.
What is observed
The capped WebSocket backpressure suite failed once on the epoll subtest with closeTimeout=1, while io_uring was clean in the same run, and passed on re-run. That is 1 failure in 73 observations under the 8 MiB memlock, one-worker, -race shape.
An epoll close-handshake failure was also seen once in CI, which is what first made #607 look engine-independent.
Measured during the #607 investigation, and this is the discriminator:
io_uring stalls show the server's kernel receive queue holding 43 KB to 845 KB unread, with the middleware reader empty and the handler blocked. The engine had not read the bytes.
epoll stalls of the same shape show server rx=0 in every case. The engine had read everything.
Same symptom at the oracle, opposite state at the socket. #607's cause is a recv chained behind a blocked SEND with IOSQE_IO_LINK, which epoll does not do at all, since it has no submission queue and no linked operations.
In 151 runs during that investigation, epoll produced zero close-handshake failures across 16,747 handshakes. So the rate here is at least two orders of magnitude below io_uring's pre-fix rate, which is why it was invisible until the io_uring noise was removed.
Why it still matters
The release bar requires validations 100% clean with the same behaviour on arm64 and x86, and this is a live one-in-seventy-three failure on a required check. It will surface as an unexplained red on unrelated pull requests, which is exactly how #607, #608, #609, #610, #613, #620 and #622 each first appeared.
It is also the kind of rate where a single run discriminates nothing. Any investigation needs a run count in the low hundreds before it can claim either a cause or a cure.
What would move it
Reproduce with a run count sized for a 1.4% per-run rate, so a cure has something to be measured against. At least 200 runs per arm for a meaningful comparison.
Check whether the epoll async dispatch path can leave a connection with its reader drained and its handler parked with no wakeup pending, which is the epoll-shaped analogue of an armed operation that never starts.
Split out of #607 so it is not closed along with it. #607's fix touches
engine/iouringonly and does not address this.What is observed
The capped WebSocket backpressure suite failed once on the epoll subtest with
closeTimeout=1, while io_uring was clean in the same run, and passed on re-run. That is 1 failure in 73 observations under the 8 MiB memlock, one-worker,-raceshape.An epoll close-handshake failure was also seen once in CI, which is what first made #607 look engine-independent.
Why it is a different defect from #607
Measured during the #607 investigation, and this is the discriminator:
server rx=0in every case. The engine had read everything.Same symptom at the oracle, opposite state at the socket. #607's cause is a recv chained behind a blocked SEND with
IOSQE_IO_LINK, which epoll does not do at all, since it has no submission queue and no linked operations.In 151 runs during that investigation, epoll produced zero close-handshake failures across 16,747 handshakes. So the rate here is at least two orders of magnitude below io_uring's pre-fix rate, which is why it was invisible until the io_uring noise was removed.
Why it still matters
The release bar requires validations 100% clean with the same behaviour on arm64 and x86, and this is a live one-in-seventy-three failure on a required check. It will surface as an unexplained red on unrelated pull requests, which is exactly how #607, #608, #609, #610, #613, #620 and #622 each first appeared.
It is also the kind of rate where a single run discriminates nothing. Any investigation needs a run count in the low hundreds before it can claim either a cause or a cure.
What would move it
/proc/net/tcpfor both sockets at the instant of failure, joined by address.server rx=0with a starved handler is the state to explain.Blocked on nothing. Related: #607, #623, #631.