Skip to content

TestAdaptiveSettledRouteRetime592/epoll/settled fails as NOT_FIXED when its reference window was never pinned (the probe's worker held no /kv conn, ~2^-8 per run); it should VOID and re-arm #790

Description

@FumingPower3925

TestAdaptiveSettledRouteRetime592/epoll/settled compares the /ping median after promotion with the median from a reference window taken while /kv was still settled. The comparison assumes that the reference window was pinned. The rig never checks this. When the probe's worker holds no /kv connection, the reference window is fast, the ratio comes out at 1.0, and the arm reports NOT_FIXED even though every celeris#592 state assertion passed. The rig should report that run as VOID and set it up again. It should not fail the product.

Test: adaptive_settled_retime_linux_test.go (package celeris, unchanged between the failing head a2a7a2b and main 0cf0c52).

The failure

PR #748, CI run 36360645780, attempt 1, job Unit (root + middleware sub-modules) 108737029035, 2026-09-28T00:03Z. Attempt 2 of the same sha (108738539507) passed.

RESULT592 engine=epoll mode=settled run=1 verdict=NOT_FIXED settled_before=true promoted_before=false adaptive=true settled_at_flip=true
settled_after=false promoted_after=true promoted_in_bound=true promote_ms=5422 async_promoted_conns=8 workers=2 kv_conns=8 ...
pre_sample_ms=5425 window_ms=10004 samples=966 stalled=0 ... ping_med_ms=0.149 ... speedup=1.0
kv_samples=264 kv_med_ms=300.987 kv_max_ms=2407.314 ctl_speedup=2013.6
ref_samples=285 ref_window_ms=3013 ref_ping_med_ms=0.151 ref_ping_max_ms=1.479 ref_stalled=0 ref_queued_frac=0.000
celeris#592 fixed-behaviour assertion failed on the settled rig: epoll: /ping median is only 1.0x faster after promotion (want >= 5x):
149.475µs over 966 post-promotion samples against 151.402µs over 285 samples taken while the route was still settled
  • The fix held. The route promoted 5.4 s after the flip, inside the 8 s bound. All 8 /kv connections were async-promoted, and the route was no longer settled at the end of the window. Only the latency ratio failed.
  • The reference window was not pinned. It contained 285 samples in 3.0 s, with a median of 0.151 ms. None was over the 5 ms stall bar, and none was over the 150 ms queued bar (D/2). The rig's own comments give 2-13 samples for a pinned window, with medians of 0.29-2.10 s. fix(server): a Start that never serves releases the caller's listener, the CPU monitor and the settle re-opener (#737) #747's re-runs gave 2-7 samples, with medians of 0.59-2.10 s.
  • All 8 /kv connections were on one worker. Before opening the reference window, the rig waits for slowCalls >= 8 (:1120). In this run the wait took pre_sample_ms - ref_window_ms = 5425 - 3013 = 2412 ms, or 8 x D. If both workers hold a /kv connection, the two workers each finish one Set every 300 ms, so the 8th slow call ends at about 1.2 s. The count reaches 8 only at 2.4 s when a single worker runs all of them one after another. kv_max_ms=2407 agrees: one request waited behind the other seven. The /ping connection was on the other worker, which was idle.

The celeris#747 lane saw the same shape locally in CI's configuration. Its reference window had 284 samples with a median of 0.107 ms. #747 attributed that failure from its log, and #778 records the point.

The defect

The rig assumes a condition it never checks. stall589KVConnsX = 4 (:255) opens 4 x workers /kv connections "so every worker holds a /kv conn (p≈1-2·2^-8)". The kernel's SO_REUSEPORT hash places each connection. The epoll engine attaches no steering program, so no worker is guaranteed a /kv connection. assertSettledRetimed592 (:534) applies the ratio whenever ref_samples >= 2. That floor handles a reference window with too few samples, where promotion came too quickly to measure. It does not handle a reference window that was never pinned, and such a window produces more samples, not fewer. The rig already computes ref_stalled and ref_queued_frac, but only prints them.

How often

workflow / job job attempts jobs with --- FAIL: TestAdaptiveSettledRouteRetime592 this shape
CI / Unit (root + middleware sub-modules) 315 (282 success, 32 failure, 1 cancelled) 5 1 (job 108737029035, 2026-09-28T00:01Z, PR #748)
Coverage / Coverage (root + middleware sub-modules) 97 (84 success, 9 failure, 4 cancelled) 0 0

The other four CI failures have different causes, and both are fixed:

The ratio assertion reached main in #625 at 2026-09-14T21:46Z. Since then the root package reported a result in about 349 of these jobs, and 1 of them failed in this shape. At 2^-8 the expected count is about 1.4, and at 2^-7 about 2.7. The data fits both rates and cannot tell them apart. Passing runs are not verbose, so their RESULT592 lines, which would show the miss directly, are not in CI logs.

Fix direction

  1. Check the premise before judging the ratio. Count the reference window as pinned only if its /ping samples waited for a blocked Set. One example is refPingMed >= queuedBar589(delay) (D/2 = 150 ms at the default). Equivalently, a pinned window has ref_queued_frac near 1, because every sample in it is over the bar. The failing run had 0.151 ms and 0.000. Keep the existing ref_samples < stall592MinRefSamples case as it is.
  2. When the premise fails, report VOID instead of NOT_FIXED, and set the rig up again a bounded number of times. Emit RESULT592 ... verdict=VOID attempt=k with the reason, tear the run down, and run runStall589 again. The new engine and new connections get new reuseport hashes. Only the final attempt's verdict counts. With 3 attempts, the chance that all three miss is about 2^-24 at the 2^-8 rate. A settled arm costs about 16 s, so a retry adds little to the root package's 300 s budget.
  3. If every attempt is VOID, fail loudly as an apparatus failure, with a message like "the rig could not pin a reference window in N attempts: the probe's worker held no /kv connection". Keep it distinct from NOT_FIXED so that nobody reads it as a adaptive dispatch never re-times a settled route: a store-backed handler that turns slow runs inline on the engine worker forever (#493 item 4, measured) #592 regression.
  4. Optional, cheaper re-arm: check placement before the flip. Before kv.slow.Store(true) (:1065), confirm that the /ping connection's worker holds at least one /kv connection, and redial the /kv connections if it does not. This needs a test-only view of which loop owns each fd, or per-loop connection counts. It avoids warming and settling the route again. It is also the only way to cover the controls: negctrl_async and negctrl_learning rely on the same unchecked placement, and when the probe's worker holds no /kv connection they pass vacuously in the same fraction of runs. In a control, /ping is fast whatever the dispatch policy, so latency cannot expose the miss.
  5. Correct the comment at :255. The event that matters is the probe's worker holding no /kv connection, and the rig should check it rather than rely on it.

How to prove the fix: CELERIS_589_KVCONNS=1 (already a diagnostic override) makes the premise fail on about half of the runs. On the fixed tree that setting should produce VOID and retry lines and never NOT_FIXED. With enough runs it should also exercise the apparatus-failure path, since all 3 attempts miss about 1/8 of the time at that setting. With the celeris#592 re-opener disabled, the arm must still fail, on promoted_in_bound.

Related

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/ciCI/CD pipelinebugSomething isn't workingtestingTesting infrastructure and helpers

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions