You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
TestAdaptiveSettledRouteRetime592/epoll/settled fails as NOT_FIXED when its reference window was never pinned (the probe's worker held no /kv conn, ~2^-8 per run); it should VOID and re-arm #790
TestAdaptiveSettledRouteRetime592/epoll/settled compares the /ping median after promotion with the median from a reference window taken while /kv was still settled. The comparison assumes that the reference window was pinned. The rig never checks this. When the probe's worker holds no /kv connection, the reference window is fast, the ratio comes out at 1.0, and the arm reports NOT_FIXED even though every celeris#592 state assertion passed. The rig should report that run as VOID and set it up again. It should not fail the product.
Test: adaptive_settled_retime_linux_test.go (package celeris, unchanged between the failing head a2a7a2b and main 0cf0c52).
The failure
PR #748, CI run 36360645780, attempt 1, job Unit (root + middleware sub-modules)108737029035, 2026-09-28T00:03Z. Attempt 2 of the same sha (108738539507) passed.
RESULT592 engine=epoll mode=settled run=1 verdict=NOT_FIXED settled_before=true promoted_before=false adaptive=true settled_at_flip=true
settled_after=false promoted_after=true promoted_in_bound=true promote_ms=5422 async_promoted_conns=8 workers=2 kv_conns=8 ...
pre_sample_ms=5425 window_ms=10004 samples=966 stalled=0 ... ping_med_ms=0.149 ... speedup=1.0
kv_samples=264 kv_med_ms=300.987 kv_max_ms=2407.314 ctl_speedup=2013.6
ref_samples=285 ref_window_ms=3013 ref_ping_med_ms=0.151 ref_ping_max_ms=1.479 ref_stalled=0 ref_queued_frac=0.000
celeris#592 fixed-behaviour assertion failed on the settled rig: epoll: /ping median is only 1.0x faster after promotion (want >= 5x):
149.475µs over 966 post-promotion samples against 151.402µs over 285 samples taken while the route was still settled
The fix held. The route promoted 5.4 s after the flip, inside the 8 s bound. All 8 /kv connections were async-promoted, and the route was no longer settled at the end of the window. Only the latency ratio failed.
All 8 /kv connections were on one worker. Before opening the reference window, the rig waits for slowCalls >= 8 (:1120). In this run the wait took pre_sample_ms - ref_window_ms = 5425 - 3013 = 2412 ms, or 8 x D. If both workers hold a /kv connection, the two workers each finish one Set every 300 ms, so the 8th slow call ends at about 1.2 s. The count reaches 8 only at 2.4 s when a single worker runs all of them one after another. kv_max_ms=2407 agrees: one request waited behind the other seven. The /ping connection was on the other worker, which was idle.
The celeris#747 lane saw the same shape locally in CI's configuration. Its reference window had 284 samples with a median of 0.107 ms. #747 attributed that failure from its log, and #778 records the point.
The defect
The rig assumes a condition it never checks. stall589KVConnsX = 4 (:255) opens 4 x workers /kv connections "so every worker holds a /kv conn (p≈1-2·2^-8)". The kernel's SO_REUSEPORT hash places each connection. The epoll engine attaches no steering program, so no worker is guaranteed a /kv connection. assertSettledRetimed592 (:534) applies the ratio whenever ref_samples >= 2. That floor handles a reference window with too few samples, where promotion came too quickly to measure. It does not handle a reference window that was never pinned, and such a window produces more samples, not fewer. The rig already computes ref_stalled and ref_queued_frac, but only prints them.
CI history, last 14 days (every attempt of every job that runs the root package, jobs started 2026-09-14T00:17Z to 2026-09-28T00:20Z; the Release workflow had no runs):
workflow / job
job attempts
jobs with --- FAIL: TestAdaptiveSettledRouteRetime592
The ratio assertion reached main in #625 at 2026-09-14T21:46Z. Since then the root package reported a result in about 349 of these jobs, and 1 of them failed in this shape. At 2^-8 the expected count is about 1.4, and at 2^-7 about 2.7. The data fits both rates and cannot tell them apart. Passing runs are not verbose, so their RESULT592 lines, which would show the miss directly, are not in CI logs.
Fix direction
Check the premise before judging the ratio. Count the reference window as pinned only if its /ping samples waited for a blocked Set. One example is refPingMed >= queuedBar589(delay) (D/2 = 150 ms at the default). Equivalently, a pinned window has ref_queued_frac near 1, because every sample in it is over the bar. The failing run had 0.151 ms and 0.000. Keep the existing ref_samples < stall592MinRefSamples case as it is.
When the premise fails, report VOID instead of NOT_FIXED, and set the rig up again a bounded number of times. Emit RESULT592 ... verdict=VOID attempt=k with the reason, tear the run down, and run runStall589 again. The new engine and new connections get new reuseport hashes. Only the final attempt's verdict counts. With 3 attempts, the chance that all three miss is about 2^-24 at the 2^-8 rate. A settled arm costs about 16 s, so a retry adds little to the root package's 300 s budget.
Optional, cheaper re-arm: check placement before the flip. Before kv.slow.Store(true) (:1065), confirm that the /ping connection's worker holds at least one /kv connection, and redial the /kv connections if it does not. This needs a test-only view of which loop owns each fd, or per-loop connection counts. It avoids warming and settling the route again. It is also the only way to cover the controls: negctrl_async and negctrl_learning rely on the same unchecked placement, and when the probe's worker holds no /kv connection they pass vacuously in the same fraction of runs. In a control, /ping is fast whatever the dispatch policy, so latency cannot expose the miss.
Correct the comment at :255. The event that matters is the probe's worker holding no /kv connection, and the rig should check it rather than rely on it.
How to prove the fix:CELERIS_589_KVCONNS=1 (already a diagnostic override) makes the premise fail on about half of the runs. On the fixed tree that setting should produce VOID and retry lines and never NOT_FIXED. With enough runs it should also exercise the apparatus-failure path, since all 3 attempts miss about 1/8 of the time at that setting. With the celeris#592 re-opener disabled, the arm must still fail, on promoted_in_bound.
TestAdaptiveSettledRouteRetime592/epoll/settledcompares the/pingmedian after promotion with the median from a reference window taken while/kvwas still settled. The comparison assumes that the reference window was pinned. The rig never checks this. When the probe's worker holds no/kvconnection, the reference window is fast, the ratio comes out at 1.0, and the arm reportsNOT_FIXEDeven though every celeris#592 state assertion passed. The rig should report that run as VOID and set it up again. It should not fail the product.Test:
adaptive_settled_retime_linux_test.go(packageceleris, unchanged between the failing heada2a7a2band main0cf0c52).The failure
PR #748, CI run 36360645780, attempt 1, job
Unit (root + middleware sub-modules)108737029035, 2026-09-28T00:03Z. Attempt 2 of the same sha (108738539507) passed./kvconnections were async-promoted, and the route was no longer settled at the end of the window. Only the latency ratio failed./kvconnections were on one worker. Before opening the reference window, the rig waits forslowCalls >= 8(:1120). In this run the wait tookpre_sample_ms - ref_window_ms= 5425 - 3013 = 2412 ms, or 8 x D. If both workers hold a/kvconnection, the two workers each finish one Set every 300 ms, so the 8th slow call ends at about 1.2 s. The count reaches 8 only at 2.4 s when a single worker runs all of them one after another.kv_max_ms=2407agrees: one request waited behind the other seven. The/pingconnection was on the other worker, which was idle.The celeris#747 lane saw the same shape locally in CI's configuration. Its reference window had 284 samples with a median of 0.107 ms. #747 attributed that failure from its log, and #778 records the point.
The defect
The rig assumes a condition it never checks.
stall589KVConnsX = 4(:255) opens 4 x workers/kvconnections "so every worker holds a /kv conn (p≈1-2·2^-8)". The kernel's SO_REUSEPORT hash places each connection. The epoll engine attaches no steering program, so no worker is guaranteed a/kvconnection.assertSettledRetimed592(:534) applies the ratio wheneverref_samples >= 2. That floor handles a reference window with too few samples, where promotion came too quickly to measure. It does not handle a reference window that was never pinned, and such a window produces more samples, not fewer. The rig already computesref_stalledandref_queued_frac, but only prints them.How often
/kvconnection is 2·2^-8 = 2^-7. The premise fails only when the empty worker is the one serving/ping. With 8 independent, uniform placements over 2 workers, that is 2^-8 (about 1 in 256) per settled epoll run. So 2^-7 is an upper bound on the miss rate. The io_uring arms skip on hosted runners (CI: two test groups have never run — the epoll sendfile e2e tests (Workers: 1 fails validation, the helper skips) and the #592 settled-route io_uring subtests (skip at 8 MiB) #709), so each Unit or Coverage job runs this arm once.--- FAIL: TestAdaptiveSettledRouteRetime592Unit (root + middleware sub-modules)Coverage (root + middleware sub-modules)The other four CI failures have different causes, and both are fixed:
server did not become ready (... Workers:1 ...)because of the memlock cap, whichskipIfMemlockCaps589fixed.epoll/negctrl_learningfailed its precondition withsettled=false promoted=true, together with a race report. This was celeris#620, fixed by test(router): wrap the clock stub once so a failing test cannot race the live engine #621.The ratio assertion reached main in #625 at 2026-09-14T21:46Z. Since then the root package reported a result in about 349 of these jobs, and 1 of them failed in this shape. At 2^-8 the expected count is about 1.4, and at 2^-7 about 2.7. The data fits both rates and cannot tell them apart. Passing runs are not verbose, so their RESULT592 lines, which would show the miss directly, are not in CI logs.
Fix direction
/pingsamples waited for a blocked Set. One example isrefPingMed >= queuedBar589(delay)(D/2 = 150 ms at the default). Equivalently, a pinned window hasref_queued_fracnear 1, because every sample in it is over the bar. The failing run had 0.151 ms and 0.000. Keep the existingref_samples < stall592MinRefSamplescase as it is.RESULT592 ... verdict=VOID attempt=kwith the reason, tear the run down, and runrunStall589again. The new engine and new connections get new reuseport hashes. Only the final attempt's verdict counts. With 3 attempts, the chance that all three miss is about 2^-24 at the 2^-8 rate. A settled arm costs about 16 s, so a retry adds little to the root package's 300 s budget.NOT_FIXEDso that nobody reads it as a adaptive dispatch never re-times a settled route: a store-backed handler that turns slow runs inline on the engine worker forever (#493 item 4, measured) #592 regression.kv.slow.Store(true)(:1065), confirm that the/pingconnection's worker holds at least one/kvconnection, and redial the/kvconnections if it does not. This needs a test-only view of which loop owns each fd, or per-loop connection counts. It avoids warming and settling the route again. It is also the only way to cover the controls:negctrl_asyncandnegctrl_learningrely on the same unchecked placement, and when the probe's worker holds no/kvconnection they pass vacuously in the same fraction of runs. In a control,/pingis fast whatever the dispatch policy, so latency cannot expose the miss.:255. The event that matters is the probe's worker holding no/kvconnection, and the rig should check it rather than rely on it.How to prove the fix:
CELERIS_589_KVCONNS=1(already a diagnostic override) makes the premise fail on about half of the runs. On the fixed tree that setting should produce VOID and retry lines and neverNOT_FIXED. With enough runs it should also exercise the apparatus-failure path, since all 3 attempts miss about 1/8 of the time at that setting. With the celeris#592 re-opener disabled, the arm must still fail, onpromoted_in_bound.Related
-voutput will make VOIDs countable.