Skip to content

io_uring resets one or two connections of 96 slightly more often than epoll (12/16 vs 6/16, p=0.037) — needs a cluster reproduction before it is believed #652

Description

@FumingPower3925

Split out of celeris#533 so it is judged on its own evidence rather than inheriting that issue's. #533 is closed: the stall it filed is resolved by celeris#607, 20 of 20 runs affected before the fix against 0 of 36 after.

This is what survives, and it is two orders of magnitude smaller.

The measurement

Clean head-to-head, 16 runs per engine, order alternated, equal workers (workers=4 throughout, zero capped by RLIMIT_MEMLOCK), TestBackpressureInboundSequenceIntegrity, 96 connections, linux/arm64, --cpus 4:

engine runs affected per-run magnitude
io_uring default 12 / 16 same
epoll 6 / 16 same

Fisher p = 0.037. The magnitude per affected run does not differ between the engines; only how often a run is affected.

What "affected" means here

Two shapes, and both occur on epoll as well:

  1. One or two connections of 96 reset. serverRST equals the number of affected connections exactly, and framesSent equals (96 − k) × 64000 for k of 1 or 2. Epoll's worst run equals io_uring's worst run.
  2. A transient dip with framesIn == framesSent, serverRST = 0 and closedOK = 96, so nothing was actually lost.

Why this is not #533 wearing different clothes

So whatever this is, it is not a stalled receive, and it is not unique to io_uring.

What would settle it

Reproduce it on the cluster before spending more on it. This was measured in a 4 GB Docker VM on an Apple-silicon Mac, and a difference of 12 against 6 at p = 0.037 is exactly the size that host noise can manufacture. The nightly and the race tier run the same refapps on real hardware with the error classes now recorded per cell, so the question is nearly free to ask there.

If it does reproduce, the next measurement is which connection resets and why: serverRST equals the affected count exactly, so a per-connection join by address would name it, the way the address join did for #607.

If it does not reproduce on the cluster, close this as a property of the development VM rather than of the engine.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingengine/iouringio_uring engine specifics

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions