Skip to content

adaptive tests: TestIdleConnsFollowSwitchRevert{Sync,Async} panic on a nil e.secondary when the io_uring standby cannot start (ENOMEM); the engine correctly stays on epoll #804

Description

@FumingPower3925

Summary

When the lazy io_uring standby cannot start (here io_uring_setup returned ENOMEM), TestIdleConnsFollowSwitchRevertSync and TestIdleConnsFollowSwitchRevertAsync panic on a nil e.secondary. They should skip, or fail cleanly under CELERIS_REQUIRE_UPSWITCH=1. The panic takes down the whole ./adaptive test binary, so every test after it in the run is lost.

Classification: a TEST bug, not an engine bug. The adaptive engine does the right thing. It aborts the switch, stays on epoll, leaves the io_uring slot nil, and keeps serving every connection. Its Metrics() and Shutdown() handle the nil slot. The dereference is in the test.

The two Promote cells do not panic in the same condition. They fail with a false celeris#657 PLACEMENT verdict instead ("nothing examines a connection that sends nothing"), because they judge a switch that never happened.

Where it was seen

This was seen in the celeris#685 lane (draft PR #793), round 2, on 2026-09-28 at 04:52Z. The run was the whole ./adaptive package with -race in a Docker container with memlock 8 MiB, on linuxkit 7.0.12 arm64 and go1.27.1. The tree was e48adee, which is main dfd044f plus the #685 fix. adaptive/idle_follow_switch_test.go is byte-identical to main's.

The ring setup failed with ENOMEM. The lane attributes this to root containers sharing uid 0's locked-memory accounting while another container with unlimited memlock was running. The same cell passed in 20 of 20 targeted re-runs (10 on each tree) and in both whole-suite m8 re-runs, so whether the ENOMEM happens depends on what else uid 0 has locked at that moment.

=== RUN   TestIdleConnsFollowSwitchRevertSync
INFO epoll engine listening addr=127.0.0.1:37035 loops=4
INFO adaptive engine listening addr=127.0.0.1:37035 active=epoll
WARN io_uring workers capped by RLIMIT_MEMLOCK requested=4 capped_to=1 memlock_cur_bytes=8388608 ...
WARN lazy standby Listen returned error standby=io_uring error="worker 0 ring setup: io_uring_setup: cannot allocate memory (likely RLIMIT_MEMLOCK; current=8388608 bytes ...)"
WARN aborting switch: lazy standby build failed; staying on current active standby=io_uring error="io_uring standby failed to start: ..." consecutive_failures=1 retry_in=30s
--- FAIL: TestIdleConnsFollowSwitchRevertSync (1.22s)
panic: runtime error: invalid memory address or nil pointer dereference [recovered, repanicked]
[signal SIGSEGV: segmentation violation code=0x1 addr=0x28 pc=0x4ee098]

goroutine 1110 [running]:
testing.tRunner.func1.2({0x9d3d80, 0xa9c0a0})
	/usr/local/go/src/testing/testing.go:2123 +0x2b8
testing.tRunner.func1()
	/usr/local/go/src/testing/testing.go:2126 +0x438
panic({0x9d3d80?, 0xa9c0a0?})
	/usr/local/go/src/runtime/panic.go:859 +0x120
github.com/goceleris/celeris/adaptive.idleFollowSwitch(0xc000293b08, {0xa2bad0, 0xadab40}, 0x0, 0x1)
	/src/adaptive/idle_follow_switch_test.go:202 +0x328
github.com/goceleris/celeris/adaptive.TestIdleConnsFollowSwitchRevertSync(0xc000293b08)
	/src/adaptive/idle_follow_switch_test.go:162 +0x44
testing.tRunner(0xc000293b08, 0xa30558)
	/usr/local/go/src/testing/testing.go:2193 +0x168
created by testing.(*T).Run in goroutine 1
	/usr/local/go/src/testing/testing.go:2258 +0x7c0
go_test_rc=2

The package result was PASS 49, FAIL 1, SKIP 1 out of 51 tests started. The other 44 tests in the package never ran.

Mechanism (line numbers at main dfd044f)

Test side, where the panic is:

  1. adaptive/idle_follow_switch_test.go:186: in the revert cells, the test calls e.ForceSwitch() to get onto io_uring before any client dials. It never checks that the switch happened.
  2. adaptive/engine.go:763-767 (performSwitch): e.secondary == nil, so the engine calls buildAndStartStandby(engine.IOUring). The worker's ring setup fails at engine/iouring/worker.go:932-935 ("worker 0 ring setup: ..."), and buildAndStartStandby returns "io_uring standby failed to start" (adaptive/engine.go:677). abortStandbyBuild records the backoff and logs "staying on current active" (adaptive/engine.go:700-708), then performSwitch returns. The slot is filled only on success (adaptive/engine.go:772-775), so e.secondary stays nil and epoll stays active.
  3. adaptive/idle_follow_switch_test.go:192 sets srcIOU := revert, which is true.
  4. adaptive/idle_follow_switch_test.go:202: promoted := subEngine(e, srcIOU).Metrics().AsyncPromotedConns. subEngine(e, true) returns e.secondary (:115-120), and that is a nil engine.Engine interface. Calling .Metrics() on it faults.
    • addr=0x28 is itab.fun[2]. engine.Engine's methods in sorted order are Addr, Listen, Metrics, Shutdown, Type, so this is exactly the call to Metrics on a nil interface.
    • subActive (:122-128) guards this nil case, and the file's own comment at :112-114 says the standby "is nil" before the first promotion. The call at :202 bypasses that guard.

Promote cells, which get a misleading failure instead: srcIOU is false, so :202 reads epoll and does not panic. The ForceSwitch at :210 aborts the same way. The test then judges placement against an io_uring engine that does not exist, and reports PLACEMENT: 64 of 64 idle conns were still on the engine switched away from and PLACEMENT-RESUME. Anyone triaging that output would read it as a #657 regression.

Engine side, which is correct:

  • performSwitch returns before it touches the nil slot (adaptive/engine.go:763-767).
  • Shutdown guards the nil slot (adaptive/engine.go:978, 995-996), and so does Metrics (adaptive/engine.go:1006, 1013-1014).
  • The controller backs off before recommending the switch again: retry_in went 30s, then 1m0s, then 2m0s (celeris#656's recordStandbyBuildFailure).

Deterministic reproduction

The repro starts from main dfd044f plus a scratch-only overlay, which is not for merge. engine/iouring/zz_repro_enomem_inject.go swaps newWorkerRing (engine/iouring/worker.go:5849, the var the io_uring init-failure regression CI job's failRingSetup656 uses) for one that returns io_uring_setup: ENOMEM. That is the same failure point as the natural run. adaptive/zz_repro_enomem_test.go calls the unchanged idleFollowSwitch with the injection active.

The container ran with --cpus 4 --ulimit memlock=-1:-1 (unlimited) and -race, so the injection is the only thing that can fail:

run result
control: unmodified TestIdleConnsFollowSwitchRevertSync, no injection PASS (converged in 25 ms, err=0)
RevertSync + ENOMEM, 3 runs 3/3 panic at idle_follow_switch_test.go:202, addr=0x28, 1.22 s, identical to the natural run
RevertAsync + ENOMEM, 3 runs 3/3 panic at idle_follow_switch_test.go:202, addr=0x28
PromoteSync / PromoteAsync + ENOMEM FAIL with the false PLACEMENT / PLACEMENT-RESUME verdict (64 of 64 conns on epoll, passes=0)
engine check: 64 keep-alive clients, 3 forced switches each, sync and async, + ENOMEM PASS: after every switch active=epoll, secondary_nil=true, requests kept flowing (about 25k per 300 ms), 0 client errors out of about 227k requests, clean shutdown

The evidence (the overlay, the scripts, the repro log and a copy of the natural log) is in evidence/celeris-adaptive-standby-nil-panic/.

Same pattern elsewhere in ./adaptive (by reading, not reproduced)

These tests also call ForceSwitch() toward io_uring and then dereference e.secondary without checking that the switch happened. Some may hit a PREMISE Fatalf before they reach the dereference, and then they fail with a misleading verdict instead of panicking:

  • reverse_transplant_test.go:131 → :142-143
  • switch_adopt_order_test.go:44 → :46
  • flap_conns_per_ring_test.go:100 → :103
  • flap_converge_poll_test.go:66 → :87
  • slow_async_converge_test.go:89 → :119

switch_idle_conns_linux_test.go:297-300 (buildStandby662) already does it right. After its ForceSwitch it checks e.ActiveEngine().Type() == engine.IOUring, and otherwise calls skipOrFailUpswitch662.

Proposed fix (test-only)

  • After every ForceSwitch that is meant to bring up the io_uring standby, assert e.ActiveEngine().Type() == engine.IOUring. If it is not, skip, or fail under CELERIS_REQUIRE_UPSWITCH=1, the way buildStandby662 and s0SkipUnlessRequired do. This applies to the revert setup switch at idle_follow_switch_test.go:186 and to the promote switch at :210, before any PLACEMENT verdict.
  • Read sub-engine metrics through the nil-guarded accessor (subActive, or a subMetrics twin) and never through a bare e.secondary.Metrics().
  • Apply the same fix to the five sites listed above.
  • Add a regression cell that runs idleFollowSwitch with ring setup forced to fail. It must end in SKIP, or in FAIL under CELERIS_REQUIRE_UPSWITCH=1, and never in a panic. It must also not report a PLACEMENT verdict.

Related: #684 covers the case where io_uring is unavailable at probe time (the skip that ignores CELERIS_REQUIRE_UPSWITCH). This issue is the other case: the probe says io_uring is available, but the standby fails at switch time. #656 (closed) is the engine-side abort and backoff that this issue confirms still holds.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/testTesting infrastructurebugSomething isn't workingplatform/linuxLinux-specific (io_uring, epoll)testingTesting infrastructure and helpers

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions