Skip to content

epoll/io_uring: an HTTP/1.1 response of 4 MiB or more sends only its headers; the body is dropped and a keep-alive client waits until its own timeout #761

Description

@FumingPower3925

Found while reviewing #746; not related to shutdown. Pre-existing on main: measured on 698bed6.

Defect

On epoll and io_uring, an HTTP/1.1 response whose body is 4 MiB or more is sent as its headers only. The body is dropped silently, the connection stays open, and a keep-alive client waits for the Content-Length bytes until its own timeout.

epoll: makeWriteBodyFn (engine/epoll/loop.go) returns without writing or erroring when cs.pendingBytes + len(cs.bodyBuf) + len(body) > cs.writeCap(), and makeWriteFn does the same. writeCap() is maxPendingBytes = 4 MiB for an H1/H2 connection (engine/epoll/conn.go). The headers are already staged, so they go out, the body never does, and the pending > cs.writeCap() close in drainRead never fires because nothing was added. io_uring has the same 4 MiB cap, maxSendQueueBytes (engine/iouring/conn.go), and the same result. Its code path has not been traced here. std delivers the response.

The cap is meant to stop a stalled peer from filling memory. Here it drops a response that was never pending, and it does so silently.

Measurement

One request per connection, the handler writing c.Blob at once, no shutdown, a raw keep-alive client reading until the body is complete, EOF or 10 s. REPRO-BIG lines, main 698bed6:

engine 1 MiB 4 MiB - 4 KiB 4 MiB
std complete complete complete
epoll complete complete 200 OK, 121 header bytes, 0 body bytes, still open at 10 s
io_uring complete complete 200 OK, 121 header bytes, 0 body bytes, still open at 10 s

The review of #746 got the same result with net/http's client (http_body=0, context deadline exceeded after 10 s) and with a Connection: close raw client (140 bytes, then EOF).

Script: lane evidence lanes-20260927/LIFECYCLE/round2/repro/run-repro.sh <ref> (TestReproBigResponse). Log: round2/logs/repro-698bed6-unlimited.log.

Fix direction

Stage a single response larger than the cap, since it is not a stalled peer's backlog, and let the back-pressure check apply to what stays unsent. Or, if a hard per-response limit is wanted, close the connection (or answer 500 before the headers) instead of dropping the body. Either way, never leave a connection open with a declared body that will not come. Test: 4 MiB and 64 MiB bodies on every engine, keep-alive and Connection: close.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/engineEngine interface or implementationbugSomething isn't workingengine/epollEpoll engine specificsengine/iouringio_uring engine specificsprotocol/h1HTTP/1.1 protocol

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions