Found in the review of #746 (celeris#703). Pre-existing on main: measured on 698bed6.
Defect
On shutdown, epoll closes every connection without flushing the writes still pending on it. Loop.run calls shutdown as soon as ctx.Err() != nil (engine/epoll/loop.go), and shutdown closes each fd in phase 3 however much of writeBuf / bodyBuf is still unsent. A response that fits in the socket buffers is unaffected. The tail of a larger one, sent to a client that reads more slowly than the loop shuts down, is lost, and the client sees EOF in the middle of the body. adaptive has the same defect while it runs epoll.
io_uring drains its sends before closing (Worker.hasPendingSends, #595), bounded by shutdownSendDrainNanos = 250 ms (engine/iouring/worker.go). std drains through http.Server.Shutdown.
With #746, Shutdown waits for Listen to return before it runs the hooks. On epoll, "Listen has returned" therefore means "the handlers have returned", not "the responses were sent". #746's godoc and docs#77 say so and point here.
Measurement
A 3 MiB response, below the 4 MiB pending cap. The client is raw, with a 64 KiB SO_RCVBUF, and starts reading 1 s after the handler is released. The handler is released 200 ms into a direct Shutdown with a 30 s budget. Control: the same with no shutdown. REPRO-FLUSH lines, main 698bed6 (bytes include the ~121-byte header):
| engine |
control (no shutdown) |
shutdown in flight |
| std |
3145849 (full) |
3145868 (full) |
| epoll |
3145849 (full) |
2634240, then EOF |
| io_uring |
3145849 (full) |
3145849 (full) |
| adaptive |
3145849 (full) |
2634240, then EOF |
Shutdown returned nil in every case.
Script: lane evidence lanes-20260927/LIFECYCLE/round2/repro/run-repro.sh <ref> (TestReproSendDrain). Log: round2/logs/repro-698bed6-unlimited.log.
Fix direction
Port #595's send drain to epoll's shutdown. Before phase 3 closes the fds, keep flushing conns with pending bytes (EPOLLOUT or a short poll loop), bounded like io_uring's 250 ms or by a deadline, then close. Test: TestReproSendDrain's shape as a failing-first engine test.
Found in the review of #746 (celeris#703). Pre-existing on main: measured on 698bed6.
Defect
On shutdown, epoll closes every connection without flushing the writes still pending on it.
Loop.runcallsshutdownas soon asctx.Err() != nil(engine/epoll/loop.go), andshutdowncloses each fd in phase 3 however much ofwriteBuf/bodyBufis still unsent. A response that fits in the socket buffers is unaffected. The tail of a larger one, sent to a client that reads more slowly than the loop shuts down, is lost, and the client sees EOF in the middle of the body. adaptive has the same defect while it runs epoll.io_uring drains its sends before closing (
Worker.hasPendingSends, #595), bounded byshutdownSendDrainNanos= 250 ms (engine/iouring/worker.go). std drains throughhttp.Server.Shutdown.With #746, Shutdown waits for Listen to return before it runs the hooks. On epoll, "Listen has returned" therefore means "the handlers have returned", not "the responses were sent". #746's godoc and docs#77 say so and point here.
Measurement
A 3 MiB response, below the 4 MiB pending cap. The client is raw, with a 64 KiB
SO_RCVBUF, and starts reading 1 s after the handler is released. The handler is released 200 ms into a directShutdownwith a 30 s budget. Control: the same with no shutdown.REPRO-FLUSHlines, main 698bed6 (bytes include the ~121-byte header):Shutdown returned nil in every case.
Script: lane evidence
lanes-20260927/LIFECYCLE/round2/repro/run-repro.sh <ref>(TestReproSendDrain). Log:round2/logs/repro-698bed6-unlimited.log.Fix direction
Port #595's send drain to epoll's
shutdown. Before phase 3 closes the fds, keep flushing conns with pending bytes (EPOLLOUT or a short poll loop), bounded like io_uring's 250 ms or by a deadline, then close. Test:TestReproSendDrain's shape as a failing-first engine test.