You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Perf checkpoint dccb839 vs 9f4d89b: arm64 RPS is MARGINAL on epoll-auto+upg-async (-1.76%) and adaptive-auto+upg-async (-1.43%) against a 1.33% floor; one replication, then bisect if it holds #789
The perf checkpoint of main dccb839 against 9f4d89b ends MARGINAL. On arm64, two engines lose RPS by slightly more than the in-run noise floor:
Engine (arm64, RPS)
Median d
Engine floor
Margin past the floor
epoll-auto+upg-async
−1.76%
1.33%
0.43 pp
adaptive-auto+upg-async
−1.43%
1.33%
0.10 pp
Both margins are inside the pre-registered 0.5 pp near-floor band. The rule for that case is to re-run the same checkpoint once before acting. This issue tracks that replication, and the bisect if the flags hold. It is not yet a confirmed regression.
Everything else is NO CHANGE:
amd64: all 9 engines, on both RPS and CPU/req.
arm64: CPU/req on all 9 engines, and RPS on the other 7.
No single cell passed its class floor on either arch or metric.
All 648 cells were ok (324 per arch; no suspect, no dnf).
What was measured
Run: goceleris/probatorium Benchmark Tier run 36340530922, branch perf/checkpoint-dccb839-vs-9f4d89b at aa4a5ea9, profile fast, both arches in parallel (msa2-server amd64, msr1 arm64). Bench window: 2026-09-27 19:24Z to 23:57Z.
Design: four columns per engine in the order A1 B1 B2 A2. A = dccb839 (v1.5.12-0.20260927173448-dccb83925a50), B = 9f4d89b (v1.5.12-0.20260920100717-9f4d89b171db). Nine engines × nine scenarios = 324 cells per arch.
d = (mean A − mean B) / mean B, sign-normalised so that negative means main is worse, for both RPS and CPU/req.
Engine floor = max(1.0%, 2 × the largest per-engine same-binary median |A1 vs A2| or |B1 vs B2|). A verdict is the engine's median d over its 9 cells against that floor. The rule and the reader were pre-registered and hash-checked before the read.
Integrity:
Every column's version label matches its arm, and every column bound its own binary (bind guard 36/36 per host).
One server process per column, exit rc=0, no restart.
No WARN, ERROR or panic line in any server log.
All 16 adaptive columns promote epoll → io_uring exactly once, at +253 to +255 s, the same in every arm. So the adaptive flag is not a placement difference.
The two flagged engines, cell by cell (arm64 RPS)
"A<B" means every A arm is below every B arm.
Scenario
adaptive-auto+upg-async
epoll-auto+upg-async
churn-close
−1.03%
−0.05%
driver-pg-read
−0.12% A<B
−0.25% A<B
driver-redis-pipeline
−1.82% (known arm64 async stall)
−11.13% (known arm64 async stall)
get-json-1k
−4.39% A<B
−3.02% A<B
get-simple
−1.43%
−5.86% A<B
get-simple-512c
−0.77% A<B
+0.38%
post-4k
−7.12% (A/A +25%)
−3.51% A<B
sse-fanout-128
−0.99% A<B
−1.76% A<B
ws-echo
−1.83%
−1.74% A<B
median
−1.43%
−1.76%
Evidence either way
For a real effect:
The sign is consistent: 9/9 cells negative for adaptive-auto, 8/9 for epoll-auto.
With the floor recomputed without the engine that set it (1.29%), both flags still hold, at +0.14 and +0.47 pp.
epoll-auto+upg-async also flags against either B column alone. adaptive-auto+upg-async flags only against B2.
A sub-floor main-worse tilt shows on both arches, not only arm64. Counting cells where every A is worse than every B, against the reverse:
amd64 CPU/req: 30 against 6 (p≈7e-5). In-run control std-h1: 0 against 3.
arm64 RPS: 28 against 11 (p≈0.009).
epoll-auto+upg-async CPU/req is also negative on amd64 (−0.72%, 7/9 cells strict), which backs the arm64 RPS flag from the other arch.
iouring-h1-sync is negative on all four engine medians: RPS −0.33% amd64, −0.09% arm64; CPU/req −1.08% amd64, −0.23% arm64. amd64 RPS is 9/9 cells negative. The previous checkpoint flagged this engine and its bisect closed it as noise.
For noise:
The arm64 RPS floor of this run (1.33%) is unusually low. The expected value is a median of 1.85% (p90 3.84%), and at that floor neither engine would flag.
The movement is arm64-wide: 54/81 arm64 RPS cells are negative (amd64 42/81), and the median of the engine medians is −0.29%. The two flagged engines are the tail of that distribution.
The h1-async siblings of both flagged engines are flat: adaptive-h1-async +0.94%, epoll-h1-async +0.08%.
A limit of the design. d = (A1 − B1 − B2 + A2)/2 is exactly the quadratic slot contrast (+1, −1, −1, +1). A humped drift that favours the middle slots, where B sits, reads as "main is worse". The previous checkpoint (A1 B A2) showed the same A-worse tilt on non-streaming rows, std-h1 included. A replication in the same order therefore reproduces any such positional effect and cannot tell it apart from a real shift.
Plan
Replicate once with checkpoint 698bed6 vs 9f4d89b: same cells, same four-column design, same unchanged comparison script.
If it does not hold, close this as run noise. The skeptic review notes a risk here: re-basing every checkpoint's B to the latest main while sub-floor main-worse tilts keep recurring means a cumulative cost would never flag. The pooled secondary in step 2 is there to catch that.
Watch items for the replication (below every floor, not findings)
iouring-h1-sync ws-echo: the read against B1 alone flags it on both arches: amd64 RPS −2.22%, arm64 CPU/req −2.74%.
arm64 sse-fanout-128: median −0.60% against a 3.79% floor. 5 cells have every A below every B.
arm64 driver-pg-read: |d| ≤ 0.52%, but 8/9 cells are negative and 5/9 have every A below every B, all on epoll and adaptive engines.
Known apparatus notes
The comparison script's "status mismatches: 0" check is vacuous on this artifact, because the merged results carry no cell_statuses key. The statuses were tabulated from the raw per-host files instead: 324/324 ok per arch.
arm64 async driver-redis-pipeline is the known stall (same-binary dispersion up to 78%) and cannot be interpreted. It is absorbed by its class floor.
Summary
The perf checkpoint of main
dccb839against9f4d89bends MARGINAL. On arm64, two engines lose RPS by slightly more than the in-run noise floor:Both margins are inside the pre-registered 0.5 pp near-floor band. The rule for that case is to re-run the same checkpoint once before acting. This issue tracks that replication, and the bisect if the flags hold. It is not yet a confirmed regression.
Everything else is NO CHANGE:
ok(324 per arch; no suspect, no dnf).What was measured
perf/checkpoint-dccb839-vs-9f4d89bataa4a5ea9, profilefast, both arches in parallel (msa2-server amd64, msr1 arm64). Bench window: 2026-09-27 19:24Z to 23:57Z.dccb839(v1.5.12-0.20260927173448-dccb83925a50), B =9f4d89b(v1.5.12-0.20260920100717-9f4d89b171db). Nine engines × nine scenarios = 324 cells per arch.The two flagged engines, cell by cell (arm64 RPS)
"A<B" means every A arm is below every B arm.
Evidence either way
For a real effect:
For noise:
A limit of the design. d = (A1 − B1 − B2 + A2)/2 is exactly the quadratic slot contrast (+1, −1, −1, +1). A humped drift that favours the middle slots, where B sits, reads as "main is worse". The previous checkpoint (A1 B A2) showed the same A-worse tilt on non-streaming rows, std-h1 included. A replication in the same order therefore reproduces any such positional effect and cannot tell it apart from a real shift.
Plan
Replicate once with checkpoint
698bed6vs9f4d89b: same cells, same four-column design, same unchanged comparison script.dccb839on the bench server's import graph.dccb839..698bed6is fix(session): keep a copy of a loaded session's ID, so a write-behind save goes to its own session (celeris#731) #734 and fix: give net/http, Prometheus, OTel and slog copies of the request strings they keep, and give Adapt's request its headers (#732, #720) #736, and neither touches code the bench server runs.Before any replication result exists, its pre-registration must:
If it holds, bisect inside
9f4d89b..dccb839:7c123da(fix(epoll): never park the loop on a running async handler, and leave the live set to the loop on an async hijack (celeris#669, celeris#668) #698) and49d2726(fix(engine): keep serving on a paused listener for 1.5 s with TCP_DEFER_ACCEPT cleared, so a switch or PauseAccept no longer resets clients that had not yet sent a request (celeris#662, #675) #674);0046b6f(fix(adaptive): treat io_uring capped to none as not viable; correct stale io_uring env docs (celeris#679) #694) and49d2726(fix(engine): keep serving on a paused listener for 1.5 s with TCP_DEFER_ACCEPT cleared, so a switch or PauseAccept no longer resets clients that had not yet sent a request (celeris#662, #675) #674).fix(engine): keep serving on a paused listener for 1.5 s with TCP_DEFER_ACCEPT cleared, so a switch or PauseAccept no longer resets clients that had not yet sent a request (celeris#662, #675) #674 is the change both flagged engines share. Before calling it a regression, weigh the by-design costs:
If it does not hold, close this as run noise. The skeptic review notes a risk here: re-basing every checkpoint's B to the latest main while sub-floor main-worse tilts keep recurring means a cumulative cost would never flag. The pooled secondary in step 2 is there to catch that.
Watch items for the replication (below every floor, not findings)
Known apparatus notes
cell_statuseskey. The statuses were tabulated from the raw per-host files instead: 324/324okper arch.