Skip to content

Perf checkpoint dccb839 vs 9f4d89b: arm64 RPS is MARGINAL on epoll-auto+upg-async (-1.76%) and adaptive-auto+upg-async (-1.43%) against a 1.33% floor; one replication, then bisect if it holds #789

Description

@FumingPower3925

Summary

The perf checkpoint of main dccb839 against 9f4d89b ends MARGINAL. On arm64, two engines lose RPS by slightly more than the in-run noise floor:

Engine (arm64, RPS) Median d Engine floor Margin past the floor
epoll-auto+upg-async −1.76% 1.33% 0.43 pp
adaptive-auto+upg-async −1.43% 1.33% 0.10 pp

Both margins are inside the pre-registered 0.5 pp near-floor band. The rule for that case is to re-run the same checkpoint once before acting. This issue tracks that replication, and the bisect if the flags hold. It is not yet a confirmed regression.

Everything else is NO CHANGE:

  • amd64: all 9 engines, on both RPS and CPU/req.
  • arm64: CPU/req on all 9 engines, and RPS on the other 7.
  • No single cell passed its class floor on either arch or metric.
  • All 648 cells were ok (324 per arch; no suspect, no dnf).

What was measured

  • Run: goceleris/probatorium Benchmark Tier run 36340530922, branch perf/checkpoint-dccb839-vs-9f4d89b at aa4a5ea9, profile fast, both arches in parallel (msa2-server amd64, msr1 arm64). Bench window: 2026-09-27 19:24Z to 23:57Z.
  • Design: four columns per engine in the order A1 B1 B2 A2. A = dccb839 (v1.5.12-0.20260927173448-dccb83925a50), B = 9f4d89b (v1.5.12-0.20260920100717-9f4d89b171db). Nine engines × nine scenarios = 324 cells per arch.
  • d = (mean A − mean B) / mean B, sign-normalised so that negative means main is worse, for both RPS and CPU/req.
  • Engine floor = max(1.0%, 2 × the largest per-engine same-binary median |A1 vs A2| or |B1 vs B2|). A verdict is the engine's median d over its 9 cells against that floor. The rule and the reader were pre-registered and hash-checked before the read.
  • Integrity:
    • Every column's version label matches its arm, and every column bound its own binary (bind guard 36/36 per host).
    • One server process per column, exit rc=0, no restart.
    • No WARN, ERROR or panic line in any server log.
    • All 16 adaptive columns promote epoll → io_uring exactly once, at +253 to +255 s, the same in every arm. So the adaptive flag is not a placement difference.

The two flagged engines, cell by cell (arm64 RPS)

"A<B" means every A arm is below every B arm.

Scenario adaptive-auto+upg-async epoll-auto+upg-async
churn-close −1.03% −0.05%
driver-pg-read −0.12% A<B −0.25% A<B
driver-redis-pipeline −1.82% (known arm64 async stall) −11.13% (known arm64 async stall)
get-json-1k −4.39% A<B −3.02% A<B
get-simple −1.43% −5.86% A<B
get-simple-512c −0.77% A<B +0.38%
post-4k −7.12% (A/A +25%) −3.51% A<B
sse-fanout-128 −0.99% A<B −1.76% A<B
ws-echo −1.83% −1.74% A<B
median −1.43% −1.76%

Evidence either way

For a real effect:

  • The sign is consistent: 9/9 cells negative for adaptive-auto, 8/9 for epoll-auto.
  • With the floor recomputed without the engine that set it (1.29%), both flags still hold, at +0.14 and +0.47 pp.
  • epoll-auto+upg-async also flags against either B column alone. adaptive-auto+upg-async flags only against B2.
  • A sub-floor main-worse tilt shows on both arches, not only arm64. Counting cells where every A is worse than every B, against the reverse:
    • amd64 CPU/req: 30 against 6 (p≈7e-5). In-run control std-h1: 0 against 3.
    • arm64 RPS: 28 against 11 (p≈0.009).
  • epoll-auto+upg-async CPU/req is also negative on amd64 (−0.72%, 7/9 cells strict), which backs the arm64 RPS flag from the other arch.
  • iouring-h1-sync is negative on all four engine medians: RPS −0.33% amd64, −0.09% arm64; CPU/req −1.08% amd64, −0.23% arm64. amd64 RPS is 9/9 cells negative. The previous checkpoint flagged this engine and its bisect closed it as noise.

For noise:

  • The arm64 RPS floor of this run (1.33%) is unusually low. The expected value is a median of 1.85% (p90 3.84%), and at that floor neither engine would flag.
  • The movement is arm64-wide: 54/81 arm64 RPS cells are negative (amd64 42/81), and the median of the engine medians is −0.29%. The two flagged engines are the tail of that distribution.
  • The h1-async siblings of both flagged engines are flat: adaptive-h1-async +0.94%, epoll-h1-async +0.08%.

A limit of the design. d = (A1 − B1 − B2 + A2)/2 is exactly the quadratic slot contrast (+1, −1, −1, +1). A humped drift that favours the middle slots, where B sits, reads as "main is worse". The previous checkpoint (A1 B A2) showed the same A-worse tilt on non-streaming rows, std-h1 included. A replication in the same order therefore reproduces any such positional effect and cannot tell it apart from a real shift.

Plan

  1. Replicate once with checkpoint 698bed6 vs 9f4d89b: same cells, same four-column design, same unchanged comparison script.

  2. Before any replication result exists, its pre-registration must:

    • define "holds": the same engine gets an arm64 RPS REGRESSION again under the same rule, and say what happens if only one of the two engines repeats;
    • pre-register a pooled strict-ordering secondary across both arches, with std-h1 as the control;
    • add a reversed-order arm (B1 A1 A2 B2), or state that a same-order replication cannot separate a middle-slot effect from a real shift.
  3. If it holds, bisect inside 9f4d89b..dccb839:

    fix(engine): keep serving on a paused listener for 1.5 s with TCP_DEFER_ACCEPT cleared, so a switch or PauseAccept no longer resets clients that had not yet sent a request (celeris#662, #675) #674 is the change both flagged engines share. Before calling it a regression, weigh the by-design costs:

  4. If it does not hold, close this as run noise. The skeptic review notes a risk here: re-basing every checkpoint's B to the latest main while sub-floor main-worse tilts keep recurring means a cumulative cost would never flag. The pooled secondary in step 2 is there to catch that.

Watch items for the replication (below every floor, not findings)

Known apparatus notes

  • The comparison script's "status mismatches: 0" check is vacuous on this artifact, because the merged results carry no cell_statuses key. The statuses were tabulated from the raw per-host files instead: 324/324 ok per arch.
  • arm64 async driver-redis-pipeline is the known stall (same-binary dispersion up to 78%) and cannot be interpreted. It is absorbed by its class floor.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingperformancePerformance optimization

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions