Summary
middleware/websocket's engine test helper drops the server's start error. startNativeServerWithHandle (engine_linux_test.go:63 at main 9f4d89b) runs s.StartWithListenerAndContext in a goroutine and sends its error to done. It then calls waitForReady (engine_test.go:17), which polls s.Addr() and a dial for 30 s and fails with server not ready within timeout. It never reads done. So when a native engine fails to start, the test burns 30 s and reports a generic timeout instead of the real error. The real error could be io_uring_setup ENOMEM at the CI shape's 8 MiB memlock, or a refused bind.
Seen
In the #633 reproducer matrix (PR #671's lane, round cover633-07-head), io_uring/multishot_recv logged the RLIMIT_MEMLOCK worker cap, never logged that it was listening, and failed after 30.01 s with server not ready within timeout. The log cannot tell whether Start failed or hung. The same file is on main and on #671, so #671 did not cause it.
Why it matters for v1.6.0
The CI shape is a one-worker io_uring ring at ulimit -l 8192, and the WebSocket suites run there. A start failure (a real regression, or an environment ENOMEM that should be a skip) is misreported as a readiness timeout. It is then easy to misfile as a flake of the #633/#652 stall family, which these same suites are used to measure. An oracle that discards an error blocks every diagnosis.
Fix direction
Make waitForReady (or the helper) select on done while polling, and fail with Start's error as soon as it returns. If the start error is the known environment case (io_uring unavailable, or ENOMEM below the documented memlock floor), skip with that reason, but only where the suite's CI step forbids silent skips, so it cannot become a silent hole (#684). Apply the same pattern to any other helper that starts a server in a goroutine and then only polls for readiness; grep for waitForReady( and Start.*Context( in _test.go.
Related: #633, #652, #684, #671.
Summary
middleware/websocket's engine test helper drops the server's start error.startNativeServerWithHandle(engine_linux_test.go:63 at main9f4d89b) runss.StartWithListenerAndContextin a goroutine and sends its error todone. It then callswaitForReady(engine_test.go:17), which pollss.Addr()and a dial for 30 s and fails withserver not ready within timeout. It never readsdone. So when a native engine fails to start, the test burns 30 s and reports a generic timeout instead of the real error. The real error could beio_uring_setupENOMEM at the CI shape's 8 MiB memlock, or a refused bind.Seen
In the #633 reproducer matrix (PR #671's lane, round
cover633-07-head), io_uring/multishot_recv logged the RLIMIT_MEMLOCK worker cap, never logged that it was listening, and failed after 30.01 s withserver not ready within timeout. The log cannot tell whether Start failed or hung. The same file is on main and on #671, so #671 did not cause it.Why it matters for v1.6.0
The CI shape is a one-worker io_uring ring at
ulimit -l 8192, and the WebSocket suites run there. A start failure (a real regression, or an environment ENOMEM that should be a skip) is misreported as a readiness timeout. It is then easy to misfile as a flake of the #633/#652 stall family, which these same suites are used to measure. An oracle that discards an error blocks every diagnosis.Fix direction
Make
waitForReady(or the helper)selectondonewhile polling, and fail with Start's error as soon as it returns. If the start error is the known environment case (io_uring unavailable, or ENOMEM below the documented memlock floor), skip with that reason, but only where the suite's CI step forbids silent skips, so it cannot become a silent hole (#684). Apply the same pattern to any other helper that starts a server in a goroutine and then only polls for readiness; grep forwaitForReady(andStart.*Context(in_test.go.Related: #633, #652, #684, #671.