Skip to content

io_uring: an async conn stays on the dirty list while its ring SEND waits for a slow client, so the worker busy-polls (~370 ms CPU per second, measured) #811

Description

@FumingPower3925

Summary

On io_uring, a promoted async conn whose queued bytes the worker sends from its dirty list stays on that list for as long as the ring SEND is in flight. While w.dirtyHead != nil, baseTimeout returns 0, so every loop iteration waits with a zero timeout. When a slow or stalled client holds the SEND, the worker busy-polls for as long as the client holds it.

Found by lane EP-3 while fixing #751 and #750, and measured. It is not caused by either fix: main and the #751 head measure the same.

Measured

Probe TestProbeAsyncSendSpin2 (evidence/lanes-20260927/EP-3/spin/zz_probe_async_send_spin2_linux_test.go, added by -overlay, not committed; script spin/probe2.sh). Docker linux/arm64, 4 CPUs, 8 MiB memlock (one io_uring worker), no -race, -count=3. A client with a 4 KiB SO_RCVBUF asks for 3 MiB and does not read. The measurement is the process's CPU time (getrusage(RUSAGE_SELF)) over the next 1 s:

arm main dfd044f #751 head b340cbe
route=sync: one request, inline handler 0.3, 0.4, 0.6 ms 0.2, 0.6, 0.7 ms
route=async: one request, async handler 0.2, 0.3, 0.4 ms 0.2, 0.5, 0.6 ms
route=async+pipelined: the same plus a second request answered while the SEND waits 368.1, 374.0, 374.4 ms 371.1, 374.3, 375.9 ms

The idle second before each measurement read 5.0 to 10.2 ms in every run. In every run the client then received all 3 MiB and the second response once it read.

Mechanism (shown in the fixture)

Probe TestProbeDirtySpinMechanism (spin/zz_probe_dirty_spin_mechanism_linux_test.go, -overlay on the #751 head for its fixture helpers; spin/probe3.sh): a promoted async conn of an fdlFixture, whose first response's direct write is short, after the worker's per-iteration work (drainDetachQueue, flushDirty):

PROBEDIRTY single    steady sending=true writeBuf=0   dirty=true dirtyHead=true baseTimeout=0s
PROBEDIRTY pipelined steady sending=true writeBuf=103 dirty=true dirtyHead=true baseTimeout=0s
  • The goroutine's short direct write enqueues the conn. drainDetachQueue puts it on the dirty list, and flushDirty moves the rest to sendBuf and submits the SEND.
  • flushDirty removes a conn only when sendBuf and writeBuf are empty (canRemove), and it skips a conn with cs.sending set. So the conn stays listed until completeSend removes it, which is when the SEND completes.
  • baseTimeout returns 0 while w.dirtyHead != nil, and with nothing to submit the loop takes mode 3b, SubmitAndWaitTimeout(0), which returns at once.

Not explained yet: the single-request async arm measured no CPU on the real engine, although the fixture shows the same listed state for it. One candidate is that the TCP send buffer took the whole 3 MiB direct write there, so no ring SEND was made; that is not checked.

Why it matters

Direction

flushDirty could drop a conn whose SEND is in flight and which owes no recv arm, because completeSend already sends whatever writeBuf holds once the SEND completes, and re-lists the conn only if the SQ ring is full. Alternatively baseTimeout could ignore listed conns that only wait on a SEND. Either needs its own failing-first test: the probe above as a test, with a CPU or loop-iteration budget.

Related: #751, #750, #527 (dirty-list hygiene), #712 (the worker's wait and deferred task work).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/engineEngine interface or implementationbugSomething isn't workingengine/iouringio_uring engine specificsperformancePerformance optimizationplatform/linuxLinux-specific (io_uring, epoll)

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions