Consolidated on 2026-10-02 from the per-PR 'Follow-ups from #N' bundles, by maintainer decision. From now on, review nits are fixed inside the PR before it merges, so no new bundles are filed. Each line names the bundle and item it came from; the bundle's thread keeps the full context. Everything here stays in v1.6.0.
Rewrite the ci.yml ramp-quarantine comment (now ci.yml:670-678 on main 24cdb7b ) that still says main without the epoll: PauseAccept silently drops connections already waiting in the accept queue — their requests get EOF (8/8, deterministic), and a promotion on a GitHub runner loses ~1/3 of 2048 connections #662 fix fails with read resets; CI: lift the adaptive job's quarantine of TestRampH1Sync and TestRampH1Async once #674 is on main (>= 6 green GitHub-hosted runs) #708 says main passed 12/12, so match CI: lift the adaptive job's quarantine of TestRampH1Sync and TestRampH1Async once #674 is on main (>= 6 green GitHub-hosted runs) #708 . (from Follow-ups from #674: stale ci.yml ramp-quarantine comment, stall write-up accuracy, B1 instrument note #724 1)
Fix the ci.yml comment on the 'celeris#691 io_uring driver tests' step: it accounts for 17 (7+2+2+3+3) of the 19 names in new=; the first group is nine and TestDriverUnregisterWaitForOnCloseThenClose and TestDriverShutdownReleasesDescriptors are uncounted. (from Follow-ups from #696: provider.go close-before-unregister contract text, the #691 CI step's test-group count, and its job not being required #726 2)
The skip-forbidding io_uring: UnregisterConn then Close leaks the driver socket — the fd-keyed ASYNC_CANCEL misses once the caller has closed the fd, so onClose never fires and the peer never sees EOF #691 (and io_uring: a worker whose ring setup or first submit fails leaves its SO_REUSEPORT listen socket open, and adaptive retries the build every tick (found by reading) #656 ) steps live in the non-required 'io_uring init-failure regression' job; either say so where the step is described or make that job a required check (org-admin ruleset change). (from Follow-ups from #696: provider.go close-before-unregister contract text, the #691 CI step's test-group count, and its job not being required #726 3)
Add TestADeferredTransplantFinishesOnlyAfterItsGoroutineExits to the Unit job's named -v step so a silent skip would show (ci.yml has no match for it). (from Follow-ups from #698: shutdown leaves a deferred transplant's descriptor open (pre-existing), EPOLLOUT comment premise, CI visibility of the new test, stale comments #727 3)
Add a CI step (or an entry in a raised-memlock step) that runs -run '^TestAdaptiveSettledRouteRetime592$' . with sudo prlimit --memlock=unlimited, -v, CELERIS_REQUIRE_IOURING_WORKERS=1 and a 6 PASS / 0 SKIP tally; the root Unit step at 8 MiB silently skips its io_uring subtests. (Also tracked as CI: two test groups have never run — the epoll sendfile e2e tests (Workers: 1 fails validation, the helper skips) and the #592 settled-route io_uring subtests (skip at 8 MiB) #709 item 2.) (from Follow-ups from #702: a raised-memlock CI step for the celeris#592 rig, with the io_uring stall assertion #739 1)
Make the c714ProbeIOUring probe retry (detach_capture_alias_linux_test.go:323 and the ws/sse/otel copies) t.Logf when it fires so CI can show it; narrow fix: clone the request values a detached stream keeps (websocket Conn.Query, SSE Last-Event-ID, Context.Detach, requestid and otel context values) (#714, #717, #718) #723 's 'No ENOMEM retry fired' sentence to the counted runs (it fired 3-15 times in ~12 superseded local mutant runs). (from Follow-ups from #723: otel traceparent clone cost, SSE OnConnect and session c.Set views, stale #714/#718 scope, ENOMEM claim wording, deliberate-red record #740 5)
Maintainer decision on the CodSpeed Free-plan budget (600 macro-runner min/month shared by celeris and loadgen; measured burn lasts ~3.6-4.1 days): gate pull_request runs on the 'performance' label in both repos, and/or move main to a daily schedule; also read real usage in CodSpeed's org settings. (from Follow-ups from #748: decide the CodSpeed budget lever now, the fork-PR exemption for org members, H2 benchmarks, bench-ab.sh result checks #754 1)
.github/scripts/bench-ab.sh goes straight to benchstat even if the regex matched no benchmark in one ref; collect benchmark names from A.txt, B.txt, A2.txt and exit non-zero naming the missing side if any set is empty or A and B differ; check with a head-only and a no-match regex. (from Follow-ups from #748: decide the CodSpeed budget lever now, the fork-PR exemption for org members, H2 benchmarks, bench-ab.sh result checks #754 4)
In start_context_shutdown_test.go, TestStartContextWatcherDoesNotRepeatADirectShutdown's check races the Shutdown goroutine under the negative control (300/300 FAIL without -race, 296/300 with); fix its doc so it does not present itself as the deterministic detector, which is TestShutdownRunsHooksOnlyAfterListenReturns. (from Follow-ups from #746: the adapted #692 assertion is not forced, the ctx-bound test's unbounded-wait failure mode, CI runs the drain-order test in one shape #777 1)
TestShutdownWaitForListenIsBoundedByCtx (shutdown_waits_listen_test.go) catches an unbounded wait only through the package -timeout panic (no --- FAIL line, 300 s, aborts the root package), and takes start after context.WithTimeout; run Shutdown on a goroutine with its own 10 s watchdog and take start before WithTimeout. (from Follow-ups from #746: the adapted #692 assertion is not forced, the ctx-bound test's unbounded-wait failure mode, CI runs the drain-order test in one shape #777 2)
CI runs TestShutdownHooksRunAfterTheDrain only in the Unit root step (no -v, 8 MiB memlock, one io_uring worker, no skip path on ENOMEM); add a named -v step so the multi-worker drain shape is covered, as Follow-ups from #698: shutdown leaves a deferred transplant's descriptor open (pre-existing), EPOLLOUT comment premise, CI visibility of the new test, stale comments #727 item 3 asks for another test. (from Follow-ups from #746: the adapted #692 assertion is not forced, the ctx-bound test's unbounded-wait failure mode, CI runs the drain-order test in one shape #777 3)
start_failure_release_linux_test.go:24 says other tests' descriptors 'cannot move' the /proc/stat count, but every celeris server in the binary opens and closes a /proc/stat monitor; reword to say the tests are safe only because they are non-parallel top-level tests. (from Follow-ups from #747: /proc/stat test comments and direction, a dial assertion that cannot fail, the uncommitted fd-delta probe #778 1)
In start_failure_release_linux_test.go change after != before to after > before (:64, :129), the leak direction, so a monitor closed elsewhere in the window cannot fail the check. (from Follow-ups from #747: /proc/stat test comments and direction, a dial assertion that cannot fail, the uncommitted fd-delta probe #778 2)
TestFailedStartClosesTheSuppliedListener (start_failure_release_test.go) closes the listener itself before dialling, so its dial assertion cannot fail; dial first, or drop the claim from the test's doc. (from Follow-ups from #747: /proc/stat test comments and direction, a dial assertion that cannot fail, the uncommitted fd-delta probe #778 3)
The fd-delta probe behind fix(server): a Start that never serves releases the caller's listener, the CPU monitor and the settle re-opener (#737) #747 's scenario table lives only in the maintainer's evidence root; commit it (as a test, or in probatorium) or mark the table maintainer-local. (from Follow-ups from #747: /proc/stat test comments and direction, a dial assertion that cannot fail, the uncommitted fd-delta probe #778 4)
TestHijackKeepsRequestViews has no reuse witness: count per round that B was served in the defect's conditions (epoll: B's connState is A's recycled one; io_uring: from the pool) and require a minimum, so a later pool change cannot make the test vacuous. (from Follow-ups from #773: an AsyncHandlers arm after #774, io_uring multishot strings read before Hijack, the copy on std, the engine-side cost, doc and witness nits #785 6)
Run the test-only commit aa2331d (TestRegisterConnDeliversBytesQueuedBeforeIt) once on amd64 (cluster celeris-stress target=cluster arches=x86 -race, or a GitHub runner) expecting FAIL; if it passes, force the window harder on amd64. (from Follow-ups from #776: a test for RegisterConn's rollback, an amd64 failing-first run, the 200 ms bound, the Coverage linger flake on pre-#744 branches #787 2)
Add the six tests of engine/iouring/park_close_fin_test.go (TestParkedWorkerSendsFINFor{AHeaderTimeout,AReadTimeout}Close, TestRunningWorkerSendsFINForAHeaderTimeoutClose, TestParkedAsyncWorkerSendsFINForAHeaderTimeoutClose, TestParkedWorkerSendsFINToMany{ReadTimeout,HeaderTimeout}Closes) to a named CI interlock in .github/workflows/ci.yml reporting want N, ran N, passed N, SKIP lines 0. Still absent on main 24cdb7b ; test(iouring): regression arms for #712 (fixed by #793) #767 's merged body also lists it as open. (from Follow-ups from #767: a named CI interlock for the #712 tests, and the bare-metal failing-first #795 1)
Add the io_uring: a parked worker keeps a stale cachedNow; if the park outlasted ReadTimeout, conns it adopts or accepts on waking get a past timestamp and the first checkTimeouts closes them (adaptive: 6/6 promotes 31 s after a demote, 0/8 at 18 s; A/B not run) #713 tests to named CI interlocks (want N, ran N, passed N, SKIP lines 0): epoll twins TestAcceptAfterALongParkIsNotTimedOutEpoll and TestAdoptAfterALongParkIsNotTimedOutEpoll (which can t.Skip silently) and the ten io_uring tests in engine/iouring/park_stale_clock_test.go; or turn the twins' skips into failures under a CELERIS_REQUIRE_* variable. Absent from ci.yml on main 24cdb7b . (from Follow-ups from #768: named CI interlocks for the #713 tests and epoll twins, and why the async standby's drain grew from ~1 s to ~11 s #797 1)
Triage the flake TestHeldRecvIsReArmedWhenTheHandOffDoesNotHappen/target_refuses on main (sync mode, -race -count=10: 1 FAIL, all 64 conns read_timeout); file a separate issue if it reproduces. No tracker exists in the four repos. (from Follow-ups from #799: assert the h2c arms' outcome, write down the exit contract, a rerunHandOff witness, an unattributed merged-tree failure, two flakes on main #816 5a)
Triage the flake TestFlapConvergesPollSync (FLAPPOLL PLACEMENT: flap 3 left 8 of 64 conns, want <=2; once in an adaptive run, then 10/10 pass on main and merged tree); file a separate issue if it reproduces. No tracker exists. (from Follow-ups from #799: assert the h2c arms' outcome, write down the exit contract, a rerunHandOff witness, an unattributed merged-tree failure, two flakes on main #816 5b)
Replace sleep-as-synchronisation in fix(epoll, iouring): send a large response whole, keep pipelined responses in order, never send a body the handler has given back (celeris#761, celeris#802, celeris#817) #805 's tests with server-side conditions: TestBackloggedPeerIsClosed's 300 ms sleep (large_response_linux_test.go:418 on main; can fail falsely), and the 2 ms / 100 ms sleeps in TestSplitBodyResponsesKeepTheConnection and TestConcurrentResponsesOwnTheirBodies (CodeRabbit on fix(epoll, iouring): send a large response whole, keep pipelined responses in order, never send a body the handler has given back (celeris#761, celeris#802, celeris#817) #805 ). (from Follow-ups from #805: the memory bound after the per-request cap, epoll's closing-conn reap by read time, no deterministic H2 refusal test, io_uring's dormant WRITEV path, body nits #818 7)
Delete skipIouringAsyncPipelined751 and its two call sites (large_response_linux_test.go:334,400,495-501 on main) now that fix(iouring): an async handler's direct write waits for a ring SEND of the conn's earlier bytes (celeris#751) #800 is merged (65d1755 ); the trial merge ran the cases 30/30 and 15/15 with the skip disabled. (from Follow-ups from #805: the memory bound after the per-request cap, epoll's closing-conn reap by read time, no deterministic H2 refusal test, io_uring's dormant WRITEV path, body nits #818 9)
TestWriteBufBackpressureClosesSlowConsumer stays gated on GOTEST_BACKPRESSURE=1 (engine/epoll/backpressure_test.go:67), so CI skips it without a record; give it a CELERIS_REQUIRE_* switch or a CI tally, or document it as a manual test. (from Follow-ups from #805: the memory bound after the per-request cap, epoll's closing-conn reap by read time, no deterministic H2 refusal test, io_uring's dormant WRITEV path, body nits #818 10)
TestShutdownH2PoolWaitIsBounded (shutdown_h2_pool_linux_test.go, bound = 3*time.Second, unchanged on main) passes on main and allows budget+3s; add a lower bound (handler still held when Start returns, >= budget) and tighten the upper bound so the test pins the wait. (from Follow-ups from #808: std h2c conns get no GOAWAY and outlive the drain, a pool-wait bound test that passes on main, cost wording #820 2)
TestShutdownWaitsForFlowControlledH2Response/*/sync (shutdown_h2_pool_linux_test.go:590 on main) synchronises on time.Sleep(300ms); on a slow runner GOAWAY can name last-stream-id 0 and the test fails for the wrong reason. Signal on stream 1's first DATA frame instead. (from Follow-ups from #808: std h2c conns get no GOAWAY and outlive the drain, a pool-wait bound test that passes on main, cost wording #820 5)
engine/std/listen_cancel_budget_test.go (~line 133 on main): shutCtx is created before start is taken, so 'el < budget' has a sub-microsecond theoretical margin; create the context after start. (from Follow-ups from #803: overlapping Shutdown calls share the shortest budget, the Engine.Shutdown doc sentence, a test margin nit #821 3)
Add the variant-rebuild case (rebuild only the .br/.gz, keep the original's mtime, resume with If-Range) to middleware/static/range_435_test.go and show it fails on e410772 . (from Follow-ups from #834: If-Range on pre-compressed .br/.gz variants (rest of #435), the IfRange date-safety comment, an unsatisfiable If-Range test, the stat/open window #846 1c)
Add an If-Range mismatch on an unsatisfiable range (must be 200, RFC 9110 13.2.2, not 416) to middleware/static/range_435_test.go for both os and fs.FS paths and to range_engines_435_linux_test.go; mutant M1 (Parse before If-Range) is caught only by one root case today. (from Follow-ups from #834: If-Range on pre-compressed .br/.gz variants (rest of #435), the IfRange date-safety comment, an unsatisfiable If-Range test, the stat/open window #846 3)
When h2: an async handler's streamed response can lose its last frames in the write queue (lost wakeup in h2ShardedQueue.DrainTo): the stream never ends for a quiet client #837 lands, put /stream and /stream-head back into the async-route raw-HEAD check in auto_head_options_engines_421_linux_test.go:215-218. (from Follow-ups from #841: HEAD now reaches #837 and #835, HEAD on Detach-streaming GET routes, the per-request isHEAD read on live streams, FullPath sentinels, OPTIONS opt-out #847 1b)
Add a HEAD case to sse: an OnConnect rejection never reaches the client: no response on epoll/io_uring/adaptive (the client hangs), an empty 200 on std #835 's regression test. (from Follow-ups from #841: HEAD now reaches #837 and #835, HEAD on Detach-streaming GET routes, the per-request isHEAD read on live streams, FullPath sentinels, OPTIONS opt-out #847 1c)
Add a Lint-job step that fails when the modules of the tools ((cd .github/tools && go list -f '{{.Module.Path}}' tool | sort -u)) differ from the dependency-name entries of the /.github/tools entry in .github/dependabot.yml (:47-51); a missing entry or a package path instead of the module path silently stops Dependabot (verified with Dependabot CLI v1.93.0). Probe fix-r2/scripts/allow-guard-probe.sh already exists. Same gap in probatorium (deps: bump the all-go-deps group across 2 directories with 6 updates #464 ) and loadgen (perf: x86 multi-socket optimization + AVX2 SIMD (v1.1.0-beta.6) #104 ), tracked there. (from Follow-ups from #839: an unchecked Dependabot allow list for .github/tools, the Scorecard side of its security trade-off, inline go install pins, a stale evidence label #853 1a)
Mention that check in the dependabot.yml comment (:44-45). (from Follow-ups from #839: an unchecked Dependabot allow list for .github/tools, the Scorecard side of its security trade-off, inline go install pins, a stale evidence label #853 1b)
Move the four inline go install pins (ci.yml:38 mage@v1.17.2, ci.yml:80 actionlint@v1.7.12, release.yml:61 and release-checklist.yml:52 mage@v1.17.2) into .github/tools as tool directives with allow entries, run via go tool -modfile, so tools are pinned one way and Dependabot bumps them. (from Follow-ups from #839: an unchecked Dependabot allow list for .github/tools, the Scorecard side of its security trade-off, inline go install pins, a stale evidence label #853 3)
Added 2026-10-02 18:40Z from #860 comment 5958063514 , the second review of #843 at f718f05 (posted 4 min before the consolidation closed #860 ):
Pin the write half of UnregisterConn's promise: in TestTeardownMarksTheConnClosedBeforeItLeavesTheWorker784/UnregisterConn, assert unregErr is still empty while the test holds c.mu. Mutant X10 (rmutate_r2.py X10) must fail it (from Follow-ups from #843: the cost of forget's EPOLL_CTL_DEL under the map lock, the test hook on the read path, the repro's thin P arm, a worker starved by one conn's inflow #860 item 5).
Six of the 13 caller-side reads in WriteAndPoll* are unpinned after a teardown (Phase B/C of WriteAndPoll and WriteAndPollBusy; WriteAndPollMulti :1010/:1046). Commit the review's zz_review2_multi_seam_784 probe, and add a hook before poll(2) for Phase B/C, or state that they are pinned by inspection only (from Follow-ups from #843: the cost of forget's EPOLL_CTL_DEL under the map lock, the test hook on the read path, the repro's thin P arm, a worker starved by one conn's inflow #860 item 6).
Add tests for X1 (forget's DEL after releasing w.mu) and X2 (shutdown marking conns closed after closing epfd). Both mutants survive the whole *784 suite. Add X1 before item 1's locking change (from Follow-ups from #843: the cost of forget's EPOLL_CTL_DEL under the map lock, the test hook on the read path, the repro's thin P arm, a worker starved by one conn's inflow #860 item 7).
read_after_unregister_784_other_test.go:80-87 handles only pair[0] == a; reuse c784TakeNumber's switch with Dup2 (from Follow-ups from #843: the cost of forget's EPOLL_CTL_DEL under the map lock, the test hook on the read path, the repro's thin P arm, a worker starved by one conn's inflow #860 item 8).
test/drivercmp/redis tests do not build off Linux: integrated_async_test.go has no build constraint but uses celerisRedisEnv, which integrated_test.go declares under //go:build linux, so go vet ./... in that module fails on darwin (same on main 24cdb7b ; CI runs on Linux, so CI is unaffected). Give integrated_async_test.go the same //go:build linux. (from fix(api): remove every deprecated public API and basicauth.HashPassword before the public v1.6.0 (celeris#830, celeris#826) #894 )
TestNoDeprecatedPublicAPI830 (deprecated_api_830_test.go, fix(api): remove every deprecated public API and basicauth.HashPassword before the public v1.6.0 (celeris#830, celeris#826) #894 at a519143 ): (a) its notice half bans every Deprecated: paragraph outside internal/ for good (doc comment :45-60), and its message (:126) says to remove the identifier "before v1.6.0" or move it under internal/. After the v1.6.0 tag both of those are breaking changes; the v1.x way is to deprecate now and remove in v2. GOVERNANCE.md and CONTRIBUTING.md state no deprecation policy, so this test is the only place one is stated. Keep the removed-names half as a permanent invariant, and turn the notice half into an explicit allowlist (empty at v1.6.0, each entry naming its v2 removal issue) whose message says so. (b) Its marker deprecationMarker = "Deprecated" + ": " (:20, used at :244) needs a trailing space, so a notice written as // Deprecated: at the end of a line passes the guard although the issue's done condition git grep 'Deprecated:' finds it (correctness reviewer's arm S5); also match Deprecated: at end of line. Test-only: fix both in fix(api): remove every deprecated public API and basicauth.HashPassword before the public v1.6.0 (celeris#830, celeris#826) #894 before it merges. (from fix(api): remove every deprecated public API and basicauth.HashPassword before the public v1.6.0 (celeris#830, celeris#826) #894 review r1: API/perf/security MINOR 1, correctness NIT 1) DONE in fix(api): remove every deprecated public API and basicauth.HashPassword before the public v1.6.0 (celeris#830, celeris#826) #894 (2c81073 , from 8ee8257 ): (a) the deprecationsAllowedInV1 allowlist, empty at v1.6.0, each entry with its notice count and v2 removal issue; (b) the marker is Deprecated: with or without text after it.
Put the three celeris#905 tests in a by-name CI interlock. Moved into fix(epoll, iouring): a stopped engine no longer hands its loop threads back to the scheduler pinned to one CPU (celeris#905) #907 (round 1, once fix(epoll, iouring): an adopted conn's carried requests get the async dispatch rule under AsyncHandlers (celeris#543) #898 had merged): the Unit job's step celeris#905 stopped engines leave no thread pinned to one CPU (skipping forbidden) runs all three by name (epoll and io_uring at the runner's memlock, adaptive with memlock raised and CELERIS_REQUIRE_UPSWITCH=1) and requires 3 PASS, no SKIP line and 3 RESULT lines with off_mask_after_stop=0 and locked_off_mask=0, and since round 2 pinned_not_main above 0 and alive_after_stop=0. Tick this when fix(epoll, iouring): a stopped engine no longer hands its loop threads back to the scheduler pinned to one CPU (celeris#905) #907 merges. (from fix(epoll, iouring): a stopped engine no longer hands its loop threads back to the scheduler pinned to one CPU (celeris#905) #907 )
The celeris#905 io_uring test checks the restore half of the fix(epoll, iouring): a stopped engine no longer hands its loop threads back to the scheduler pinned to one CPU (celeris#905) #907 fix (the main thread gets its mask back) only in runs where the one io_uring worker ran on the main thread: 8 of 20 first cycles locally, none in CI run 37087402233. More cycles do not raise that: with up to 12 cycles, 5 of 10 runs never had the worker on the main thread (L905 logs/r1-x1/iouring-x10.log). A deterministic check needs the worker started on the main thread on purpose (a test hook in the engine, say). TestSaveThreadAffinityRestoresThePin covers Save and Restore themselves, and epoll and adaptive have a loop on the main thread in 20 of 20 first cycles. (from fix(epoll, iouring): a stopped engine no longer hands its loop threads back to the scheduler pinned to one CPU (celeris#905) #907 )
Added 2026-10-03 (from #899 , the last pre-rule bundle):
A half-done async promotion in a transplant replay is caught only by -race. Add an assertion that does not need the race detector, so that mutants C/D (cs.asyncPromoted not set; evidence lanes-20261002/L616_543/review-correctness-r2/) fail without -race (from Follow-ups from #898 (celeris#543): a promoting replay's RequestCount differs by engine; a replay parse error closes without the error response; a replayed h2c upgrade closes the conn #899 item 3).
CONTRIBUTING.md copies CI's tool pins by hand: mage v1.17.2 (CONTRIBUTING.md:19 after fix(mage): Tools installs h2spec 2.6.0 on every platform and prints the version it ended with; pin the mage and benchstat hints (celeris#838) #910 ) and golangci-lint v2.13 (:22), while ci.yml:38, release.yml:61 and release-checklist.yml:52 pin mage and ci.yml pins golangci-lint, so the next bump leaves the hints stale with nothing failing. When the item that moves mage into .github/tools (from Follow-ups from #839: an unchecked Dependabot allow list for .github/tools, the Scorecard side of its security trade-off, inline go install pins, a stale evidence label #853 3) lands, point the hint at GOWORK=off go install -modfile=.github/tools/go.mod tool as probatorium's CONTRIBUTING does (probatorium#485), and have the Lint job compare CONTRIBUTING's golangci-lint version with ci.yml's. (from fix(mage): Tools installs h2spec 2.6.0 on every platform and prints the version it ended with; pin the mage and benchstat hints (celeris#838) #910 review round 1, nit)
Added 2026-10-03 (#907 's round-2 reviews, lane L905):
Codecov counts test-only helper packages as library code: fix(epoll, iouring): a stopped engine no longer hands its loop threads back to the scheduler pinned to one CPU (celeris#905) #907 's patch coverage read 3.52%, with 218 of its 219 missing lines in internal/platform/pintest/pintest_linux.go (the Codecov comment on fix(epoll, iouring): a stopped engine no longer hands its loop threads back to the scheduler pinned to one CPU (celeris#905) #907 , issuecomment-5964324517), and middleware/internal/testutil is counted the same way. Add both, and any later test-support package, to codecov.yml's ignore list, which today holds only test/**, /testdata/ and **/*.pb.go. celeristest is public API and stays counted. (from fix(epoll, iouring): a stopped engine no longer hands its loop threads back to the scheduler pinned to one CPU (celeris#905) #907 )
Added 2026-10-03 (#910 's round-2 reviews, lane CI1):
mage tools accepts on its second run the h2spec its first run refused (magefile.go Tools, fix(mage): Tools installs h2spec 2.6.0 on every platform and prints the version it ended with; pin the mage and benchstat hints (celeris#838) #910 at fe509c8 ). The early return (:106-111) checks only --version, which the -ldflags stamp makes say 2.6.0 for any code. So the 2.6.0-stamped v2.2.1+incompatible build that the post-install build-info check refuses (:131-137) stays in GOBIN, and the next mage tools reports it "(already installed)" and exits 0. The same branch accepts such a build anywhere on PATH. Fix: in the early path, read h2specBuiltFrom(path) and reinstall when the build info names github.com/summerwind/h2spec at a version other than h2specModVersion. Keep binaries with no Go build info (the release tarballs) and homebrew's build, which records exactly v1.5.1-0.20200804131034-70ac22940108 (go version -m /opt/homebrew/bin/h2spec). Optionally os.Remove the install before returning a refusal. Failing test: the review's TestReviewerRefusalIsSticky (lane evidence lanes-20261003/CI1/review-correctness-r1/r1x-tests/magetools/reviewer_rerun_test.go), which FAILs at fe509c8 on darwin/arm64, linux/arm64 and linux/amd64 (r1x-logs/838-{host,linux-arm64,linux-amd64}-P-probe-sticky-head.log). (from fix(mage): Tools installs h2spec 2.6.0 on every platform and prints the version it ended with; pin the mage and benchstat hints (celeris#838) #910 review round 2, minor and nit; a behaviour change, so not fixed in the PR)
.coderabbit.yaml:184 (main fa76fe6 ) tells the review bot, as a fact, "Code from a fork pull request never runs on a self-hosted or CodSpeed macro runner". An organization member's fork pull request runs without approval (CI: the fork-PR approval policy exempts org members, so a read-only member's fork PR runs codspeed.yml on the macro runner without approval #864 ), and the if: that skips forks sits in the workflow file, which a fork can edit. Reword it as the rule it is, the way codspeed.yml:108-120 already does: such a job must skip forks with head.repo.full_name == github.repository, and the fork-PR approval policy is what holds a fork that edits the workflow. probatorium#486 corrects the same claim in probatorium's matrix-pr-tier.yml and CONTRIBUTING.md, and loadgen#108 in loadgen's codspeed.yml. (found by the CI: the fork-PR approval policy exempts org members, so a read-only member's fork PR runs codspeed.yml on the macro runner without approval #864 fork-claim audit in fix(mage): Tools installs h2spec 2.6.0 on every platform and prints the version it ended with; pin the mage and benchstat hints (celeris#838) #910 's round-2 fix)
Added 2026-10-03 (#923 's round-1 reviews, lane F3T):
middleware/cache/slow_leader_913_linux_test.go (fix(cache): the coalesced fill stores the response and releases its followers before any client is written, so a slow leader's client no longer holds them (celeris#913) #923 at 87ecd45 , :86-98): the handler only ever returns a cacheable 200, so the test covers only the coalesced fill. fix(cache): the coalesced fill stores the response and releases its followers before any client is written, so a slow leader's client no longer holds them (celeris#913) #923 also takes the leader's write out of the call on every uncacheable path, where followers fall back to executeAndStore (no-store, non-2xx, 206, a body over MaxBodyBytes, a handler that writes its body and returns an error), and reorders the non-coalesced path to store before it writes (cache.go:139-149). No test checks either. Add one uncacheable row (no-store or 206) and one Singleflight: false row to the same harness, with failing-first on base + fix(h2): bound the response bytes a connection holds for the peer's window; over the budget, handlers run on the pool and send as the windows open (celeris#893) #906 . The review's probe TestReview923SlowLeaderPaths (7 paths x 4 engines) already shows the shape: all 28 rows FAIL on base + fix(h2): bound the response bytes a connection holds for the peer's window; over the budget, handlers run on the pool and send as the windows open (celeris#893) #906 and meet the fixed semantics on head + fix(h2): bound the response bytes a connection holds for the peer's window; over the budget, handlers run on the pool and send as the windows open (celeris#893) #906 , except three native errbody rows whose assertion is over-strict (followers on the leader's sync loop arrive after the call ends). Evidence: lanes-20261003/F3T/review923-corr-r1/probes/review923_paths_probe_test.go, logs/q1/rp-{test,fix,906test,906fix}.probe.log. (from fix(cache): the coalesced fill stores the response and releases its followers before any client is written, so a slow leader's client no longer holds them (celeris#913) #923 review r1, correctness MINOR 1)
Added 2026-10-03 (#935 's round-1 reviews, lane D2):
redis and memcached: add a portable recording-loop test for Close, like postgres's checkCalls in fix(postgres): keep a conn's descriptor number until Close has made its last call on it (celeris#859) #935 (driver/postgres/close_after_teardown_859_test.go): a WorkerLoop that records, at each call by number, which socket the number names (by inode). fix(postgres): keep a conn's descriptor number until Close has made its last call on it (celeris#859) #935 's end-to-end guards (driver/redis/close_after_teardown_859_linux_test.go, driver/memcached/close_after_teardown_859_linux_test.go) catch onClose releasing the number. They do not catch Close releasing it before UnregisterConn: Close unregisters (driver/redis/conn.go:946, driver/memcached/conn.go:667 at 56a6c1e ) before it closes the fd (:953, :674). With that order inverted, UnregisterConn would act on whatever conn has taken the number. The correctness review's arm x9 (Close calls file.Close() before UnregisterConn) passes every postgres: after a worker-side teardown, Close writes Terminate to and unregisters the conn that has taken the closed descriptor number #859 test, PASS 70 / FAIL 0 / SKIP 0 for each driver, and the whole redis and memcached suites, PASS 159 / FAIL 0 / SKIP 0. Evidence: lanes-20261003/D2/review-correctness-r1/logs/ctl-x9-{redis,memcached}-release-before-unregister-arm64.log and suites-x9-arm64.log; injections in review-correctness-r1/scripts/make-variants.py. The new test must FAIL under x9 and PASS on main. (from fix(postgres): keep a conn's descriptor number until Close has made its last call on it (celeris#859) #935 review r1, correctness MINOR)
Added 2026-10-03 (#932 's round-1 review, lane D1):
Added 2026-10-03 (#932 's round-2 review, lane D1):
fix(eventloop): a Write or RegisterConn racing Loop.Close no longer uses the worker's eventfd or epoll fd number after shutdown closed it (celeris#862) #932 register tests: TestRegisterConnRacingCloseNeverAddsToAClosedEpoll862 and TestRegisterConnRacingUnregisterConnLeavesNoEpollEntry862 (driver/internal/eventloop/wake_close_862_linux_test.go) run their hook at testHookBeforeAdd, which RegisterConn calls before it takes c.mu (loop_linux.go:328 at dce5a3c ). So they do not pin that the c.closed check and the EPOLL_CTL_ADD are one critical section under c.mu (:346-361). A refactor that checks c.closed under c.mu, unlocks, and then issues the ADD passes every eventloop: a Write racing Loop.Close can write the closed eventfd's number in wake(), which another socket may hold by then (data race) #862 test. The correctness review's mutant CHECKOUTSIDE (that shape) passed all three tests in 20 of 20 -race processes, 60/0/0, with 0 race reports (evidence lanes-20261003/D1/review-correctness-r2/logs/r2a/mut-CHECKOUTSIDE.log). Add a deterministic pin in the style of TestTeardownMarksTheConnClosedBeforeItLeavesTheWorker784. A nil-by-default hook inside the c.mu section, after the check and before the ADD, holds the section while the test calls Loop.Close on another goroutine. The test asserts that Close has not returned while the hook is held, then releases it and checks that the ADD landed on the worker's own epoll instance. CHECKOUTSIDE must fail it. (from fix(eventloop): a Write or RegisterConn racing Loop.Close no longer uses the worker's eventfd or epoll fd number after shutdown closed it (celeris#862) #932 review round 2, correctness NIT)
Added 2026-10-04 (from #938 's review round 1, lane M443):
The 14 //go:build !race files never run in CI: every go test in ci.yml and test-coverage.yml, and mage test, passes -race. They hold the alloc guards of driver/{memcached,postgres,redis}, middleware/{cache,idempotency,overload,ratelimit,timeout} and celeristest (added by refactor!: the engine, protocol and driver codec packages move under internal/ (celeris#443) #938 ), plus the race_off files. That is how refactor!: the engine, protocol and driver codec packages move under internal/ (celeris#443) #938 's commit b shipped green while middleware/timeout's TestNoTimeoutCommonPathSingleAlloc failed (3 allocs/op, want <=1; fixed in refactor!: the engine, protocol and driver codec packages move under internal/ (celeris#443) #938 by e55fa03 ). Add a short non-race step to the Unit job, for example go test -count=1 -run 'TestAllocBudgets|TestAllocBudgetsPreparedExec|TestAllowUnderLimitZeroAlloc|TestNoTimeoutCommonPathSingleAlloc|TestNewContextZeroAlloc' ./driver/... ./middleware/... ./celeristest/, about 10 s in a container. (from refactor!: the engine, protocol and driver codec packages move under internal/ (celeris#443) #938 r1, MINOR)
TestServerMixedAsync_ConcurrentNoCrosstalk (server_async_test.go:134) flakes on darwin with "connection reset by peer" on the std engine: go test -count=10 -run 'TestServerMixedAsync_ConcurrentNoCrosstalk$' . failed 2 of 10 at main a43aa1d and 1 of 10 at refactor!: the engine, protocol and driver codec packages move under internal/ (celeris#443) #938 's head, on darwin/arm64. It passes on Linux, where CI runs it. (from refactor!: the engine, protocol and driver codec packages move under internal/ (celeris#443) #938 r1, NIT)
The api/ golden files (ci: check the exported API against golden files in api/, so every API change shows in review (celeris#443) #914 ) do not show five kinds of breaking change, which api/README.md lists for review by hand: a struct that stops being comparable; the first unexported field added to an all-exported struct; reordered exported fields (both break unkeyed literals); a renamed package clause; methods promoted through an unexported embedded interface of another module. Each fix changes what apidump reports, so each comes with a mage api regeneration; the plan's estimate is 2-3 h for all five. (from ci: check the exported API against golden files in api/, so every API change shows in review (celeris#443) #914 's review MINORs, decision D-F in refactor!: the engine, protocol and driver codec packages move under internal/ (celeris#443) #938 )
Coverage's root step (test-coverage.yml) runs ./internal/engine/iouring in the same go test -race -cover as the rest of ci.yml's unit set, at the runner's 8 MiB memlock, and that package sometimes cannot set up a ring: io_uring_setup: cannot allocate memory (likely RLIMIT_MEMLOCK ...). It failed TestHandoffHasNothingInFlight at refactor!: the engine, protocol and driver codec packages move under internal/ (celeris#443) #938 's ee8ae8b (run 37162905087, 2 of 2 attempts) and TestAsyncH2CUpgradeOnPromotionLeavesH1StateToTheGoroutine at refactor!: the engine, protocol and driver codec packages move under internal/ (celeris#443) #938 's 998d7bf (run 37169188620, attempt 1; the rerun passed). Neither commit touches that package's code. ci.yml's Unit job runs the package in a step of its own and has not hit it. The likely cause, not proven, is the other test binaries of the same go test holding io_uring memory at the same time. Fix: run ./internal/engine/iouring in its own step in test-coverage.yml as ci.yml does, with its own profile and upload, and keep the package-set check in step. (from refactor!: the engine, protocol and driver codec packages move under internal/ (celeris#443) #938 r2, CI)
Consolidated on 2026-10-02 from the per-PR 'Follow-ups from #N' bundles, by maintainer decision. From now on, review nits are fixed inside the PR before it merges, so no new bundles are filed. Each line names the bundle and item it came from; the bundle's thread keeps the full context. Everything here stays in v1.6.0.
Added 2026-10-02 18:40Z from #860 comment 5958063514, the second review of #843 at f718f05 (posted 4 min before the consolidation closed #860):
Pin the write half of UnregisterConn's promise: in TestTeardownMarksTheConnClosedBeforeItLeavesTheWorker784/UnregisterConn, assert unregErr is still empty while the test holds c.mu. Mutant X10 (rmutate_r2.py X10) must fail it (from Follow-ups from #843: the cost of forget's EPOLL_CTL_DEL under the map lock, the test hook on the read path, the repro's thin P arm, a worker starved by one conn's inflow #860 item 5).
Six of the 13 caller-side reads in WriteAndPoll* are unpinned after a teardown (Phase B/C of WriteAndPoll and WriteAndPollBusy; WriteAndPollMulti :1010/:1046). Commit the review's zz_review2_multi_seam_784 probe, and add a hook before poll(2) for Phase B/C, or state that they are pinned by inspection only (from Follow-ups from #843: the cost of forget's EPOLL_CTL_DEL under the map lock, the test hook on the read path, the repro's thin P arm, a worker starved by one conn's inflow #860 item 6).
Add tests for X1 (forget's DEL after releasing w.mu) and X2 (shutdown marking conns closed after closing epfd). Both mutants survive the whole *784 suite. Add X1 before item 1's locking change (from Follow-ups from #843: the cost of forget's EPOLL_CTL_DEL under the map lock, the test hook on the read path, the repro's thin P arm, a worker starved by one conn's inflow #860 item 7).
read_after_unregister_784_other_test.go:80-87 handles only pair[0] == a; reuse c784TakeNumber's switch with Dup2 (from Follow-ups from #843: the cost of forget's EPOLL_CTL_DEL under the map lock, the test hook on the read path, the repro's thin P arm, a worker starved by one conn's inflow #860 item 8).
test/drivercmp/redistests do not build off Linux:integrated_async_test.gohas no build constraint but usescelerisRedisEnv, whichintegrated_test.godeclares under//go:build linux, sogo vet ./...in that module fails on darwin (same on main 24cdb7b; CI runs on Linux, so CI is unaffected). Giveintegrated_async_test.gothe same//go:build linux. (from fix(api): remove every deprecated public API and basicauth.HashPassword before the public v1.6.0 (celeris#830, celeris#826) #894)TestNoDeprecatedPublicAPI830(deprecated_api_830_test.go, fix(api): remove every deprecated public API and basicauth.HashPassword before the public v1.6.0 (celeris#830, celeris#826) #894 at a519143): (a) its notice half bans everyDeprecated:paragraph outside internal/ for good (doc comment :45-60), and its message (:126) says to remove the identifier "before v1.6.0" or move it under internal/. After the v1.6.0 tag both of those are breaking changes; the v1.x way is to deprecate now and remove in v2. GOVERNANCE.md and CONTRIBUTING.md state no deprecation policy, so this test is the only place one is stated. Keep the removed-names half as a permanent invariant, and turn the notice half into an explicit allowlist (empty at v1.6.0, each entry naming its v2 removal issue) whose message says so. (b) Its markerdeprecationMarker = "Deprecated" + ": "(:20, used at :244) needs a trailing space, so a notice written as// Deprecated:at the end of a line passes the guard although the issue's done conditiongit grep 'Deprecated:'finds it (correctness reviewer's arm S5); also matchDeprecated:at end of line. Test-only: fix both in fix(api): remove every deprecated public API and basicauth.HashPassword before the public v1.6.0 (celeris#830, celeris#826) #894 before it merges. (from fix(api): remove every deprecated public API and basicauth.HashPassword before the public v1.6.0 (celeris#830, celeris#826) #894 review r1: API/perf/security MINOR 1, correctness NIT 1) DONE in fix(api): remove every deprecated public API and basicauth.HashPassword before the public v1.6.0 (celeris#830, celeris#826) #894 (2c81073, from 8ee8257): (a) thedeprecationsAllowedInV1allowlist, empty at v1.6.0, each entry with its notice count and v2 removal issue; (b) the marker isDeprecated:with or without text after it.Put the three celeris#905 tests in a by-name CI interlock. Moved into fix(epoll, iouring): a stopped engine no longer hands its loop threads back to the scheduler pinned to one CPU (celeris#905) #907 (round 1, once fix(epoll, iouring): an adopted conn's carried requests get the async dispatch rule under AsyncHandlers (celeris#543) #898 had merged): the Unit job's step
celeris#905 stopped engines leave no thread pinned to one CPU (skipping forbidden)runs all three by name (epoll and io_uring at the runner's memlock, adaptive with memlock raised and CELERIS_REQUIRE_UPSWITCH=1) and requires 3 PASS, no SKIP line and 3 RESULT lines with off_mask_after_stop=0 and locked_off_mask=0, and since round 2 pinned_not_main above 0 and alive_after_stop=0. Tick this when fix(epoll, iouring): a stopped engine no longer hands its loop threads back to the scheduler pinned to one CPU (celeris#905) #907 merges. (from fix(epoll, iouring): a stopped engine no longer hands its loop threads back to the scheduler pinned to one CPU (celeris#905) #907)The celeris#905 io_uring test checks the restore half of the fix(epoll, iouring): a stopped engine no longer hands its loop threads back to the scheduler pinned to one CPU (celeris#905) #907 fix (the main thread gets its mask back) only in runs where the one io_uring worker ran on the main thread: 8 of 20 first cycles locally, none in CI run 37087402233. More cycles do not raise that: with up to 12 cycles, 5 of 10 runs never had the worker on the main thread (L905
logs/r1-x1/iouring-x10.log). A deterministic check needs the worker started on the main thread on purpose (a test hook in the engine, say).TestSaveThreadAffinityRestoresThePincovers Save and Restore themselves, and epoll and adaptive have a loop on the main thread in 20 of 20 first cycles. (from fix(epoll, iouring): a stopped engine no longer hands its loop threads back to the scheduler pinned to one CPU (celeris#905) #907)Added 2026-10-03 (from #899, the last pre-rule bundle):
A half-done async promotion in a transplant replay is caught only by -race. Add an assertion that does not need the race detector, so that mutants C/D (cs.asyncPromoted not set; evidence lanes-20261002/L616_543/review-correctness-r2/) fail without -race (from Follow-ups from #898 (celeris#543): a promoting replay's RequestCount differs by engine; a replay parse error closes without the error response; a replayed h2c upgrade closes the conn #899 item 3).
CONTRIBUTING.md copies CI's tool pins by hand: mage v1.17.2 (CONTRIBUTING.md:19 after fix(mage): Tools installs h2spec 2.6.0 on every platform and prints the version it ended with; pin the mage and benchstat hints (celeris#838) #910) and golangci-lint v2.13 (:22), while ci.yml:38, release.yml:61 and release-checklist.yml:52 pin mage and ci.yml pins golangci-lint, so the next bump leaves the hints stale with nothing failing. When the item that moves mage into .github/tools (from Follow-ups from #839: an unchecked Dependabot allow list for .github/tools, the Scorecard side of its security trade-off, inline go install pins, a stale evidence label #853 3) lands, point the hint at
GOWORK=off go install -modfile=.github/tools/go.mod toolas probatorium's CONTRIBUTING does (probatorium#485), and have the Lint job compare CONTRIBUTING's golangci-lint version with ci.yml's. (from fix(mage): Tools installs h2spec 2.6.0 on every platform and prints the version it ended with; pin the mage and benchstat hints (celeris#838) #910 review round 1, nit)Added 2026-10-03 (#907's round-2 reviews, lane L905):
ignorelist, which today holds only test/**, /testdata/ and **/*.pb.go. celeristest is public API and stays counted. (from fix(epoll, iouring): a stopped engine no longer hands its loop threads back to the scheduler pinned to one CPU (celeris#905) #907)Added 2026-10-03 (#910's round-2 reviews, lane CI1):
mage toolsaccepts on its second run the h2spec its first run refused (magefile.goTools, fix(mage): Tools installs h2spec 2.6.0 on every platform and prints the version it ended with; pin the mage and benchstat hints (celeris#838) #910 at fe509c8). The early return (:106-111) checks only--version, which the-ldflagsstamp makes say 2.6.0 for any code. So the 2.6.0-stamped v2.2.1+incompatible build that the post-install build-info check refuses (:131-137) stays in GOBIN, and the nextmage toolsreports it "(already installed)" and exits 0. The same branch accepts such a build anywhere on PATH. Fix: in the early path, readh2specBuiltFrom(path)and reinstall when the build info namesgithub.com/summerwind/h2specat a version other thanh2specModVersion. Keep binaries with no Go build info (the release tarballs) and homebrew's build, which records exactlyv1.5.1-0.20200804131034-70ac22940108(go version -m /opt/homebrew/bin/h2spec). Optionallyos.Removethe install before returning a refusal. Failing test: the review'sTestReviewerRefusalIsSticky(lane evidencelanes-20261003/CI1/review-correctness-r1/r1x-tests/magetools/reviewer_rerun_test.go), which FAILs at fe509c8 on darwin/arm64, linux/arm64 and linux/amd64 (r1x-logs/838-{host,linux-arm64,linux-amd64}-P-probe-sticky-head.log). (from fix(mage): Tools installs h2spec 2.6.0 on every platform and prints the version it ended with; pin the mage and benchstat hints (celeris#838) #910 review round 2, minor and nit; a behaviour change, so not fixed in the PR).coderabbit.yaml:184(main fa76fe6) tells the review bot, as a fact, "Code from a fork pull request never runs on a self-hosted or CodSpeed macro runner". An organization member's fork pull request runs without approval (CI: the fork-PR approval policy exempts org members, so a read-only member's fork PR runs codspeed.yml on the macro runner without approval #864), and theif:that skips forks sits in the workflow file, which a fork can edit. Reword it as the rule it is, the way codspeed.yml:108-120 already does: such a job must skip forks withhead.repo.full_name == github.repository, and the fork-PR approval policy is what holds a fork that edits the workflow. probatorium#486 corrects the same claim in probatorium's matrix-pr-tier.yml and CONTRIBUTING.md, and loadgen#108 in loadgen's codspeed.yml. (found by the CI: the fork-PR approval policy exempts org members, so a read-only member's fork PR runs codspeed.yml on the macro runner without approval #864 fork-claim audit in fix(mage): Tools installs h2spec 2.6.0 on every platform and prints the version it ended with; pin the mage and benchstat hints (celeris#838) #910's round-2 fix)Added 2026-10-03 (#923's round-1 reviews, lane F3T):
executeAndStore(no-store, non-2xx, 206, a body overMaxBodyBytes, a handler that writes its body and returns an error), and reorders the non-coalesced path to store before it writes (cache.go:139-149). No test checks either. Add one uncacheable row (no-store or 206) and oneSingleflight: falserow to the same harness, with failing-first on base + fix(h2): bound the response bytes a connection holds for the peer's window; over the budget, handlers run on the pool and send as the windows open (celeris#893) #906. The review's probeTestReview923SlowLeaderPaths(7 paths x 4 engines) already shows the shape: all 28 rows FAIL on base + fix(h2): bound the response bytes a connection holds for the peer's window; over the budget, handlers run on the pool and send as the windows open (celeris#893) #906 and meet the fixed semantics on head + fix(h2): bound the response bytes a connection holds for the peer's window; over the budget, handlers run on the pool and send as the windows open (celeris#893) #906, except three native errbody rows whose assertion is over-strict (followers on the leader's sync loop arrive after the call ends). Evidence:lanes-20261003/F3T/review923-corr-r1/probes/review923_paths_probe_test.go,logs/q1/rp-{test,fix,906test,906fix}.probe.log. (from fix(cache): the coalesced fill stores the response and releases its followers before any client is written, so a slow leader's client no longer holds them (celeris#913) #923 review r1, correctness MINOR 1)Added 2026-10-03 (#935's round-1 reviews, lane D2):
checkCallsin fix(postgres): keep a conn's descriptor number until Close has made its last call on it (celeris#859) #935 (driver/postgres/close_after_teardown_859_test.go): a WorkerLoop that records, at each call by number, which socket the number names (by inode). fix(postgres): keep a conn's descriptor number until Close has made its last call on it (celeris#859) #935's end-to-end guards (driver/redis/close_after_teardown_859_linux_test.go, driver/memcached/close_after_teardown_859_linux_test.go) catch onClose releasing the number. They do not catch Close releasing it before UnregisterConn: Close unregisters (driver/redis/conn.go:946, driver/memcached/conn.go:667 at 56a6c1e) before it closes the fd (:953, :674). With that order inverted, UnregisterConn would act on whatever conn has taken the number. The correctness review's arm x9 (Close callsfile.Close()beforeUnregisterConn) passes every postgres: after a worker-side teardown, Close writes Terminate to and unregisters the conn that has taken the closed descriptor number #859 test, PASS 70 / FAIL 0 / SKIP 0 for each driver, and the whole redis and memcached suites, PASS 159 / FAIL 0 / SKIP 0. Evidence:lanes-20261003/D2/review-correctness-r1/logs/ctl-x9-{redis,memcached}-release-before-unregister-arm64.logandsuites-x9-arm64.log; injections inreview-correctness-r1/scripts/make-variants.py. The new test must FAIL under x9 and PASS on main. (from fix(postgres): keep a conn's descriptor number until Close has made its last call on it (celeris#859) #935 review r1, correctness MINOR)Added 2026-10-03 (#932's round-1 review, lane D1):
TestWriteRacingCloseNeverTouchesTheClosedEventfd862(driver/internal/eventloop/wake_close_862_linux_test.go) catches its defect only through the race detector. Its second oracle (two eventfds that take the freed numbers must receive no write) needs a wake to land in the few microseconds between the close and the reuse: on main without-raceit fired in 0 of 10 processes. The likeliest regression of the fixed design is a wake that writesw.wakeFD.FD()itself (FD()is exported and lock-free, andrunalready uses it). The correctness review's mutant WAKEATOMIC (that shape) failed the test in 1 of 10 processes with-race, by the second oracle, and 0 of 10 without (evidencelanes-20261003/D1/review-correctness-r1/logs/run-r1/m-WAKEATOMIC.log,m-WAKEATOMIC-norace.log,base-wake-norace.log). Add a deterministic test: a nil-by-default hook inshutdownbetween the eventfd's close and the return ofw.mu, where the test opens the taker eventfd and then lets a wake that loaded the number before the close go ahead, so the regression fails without-race. fix(eventloop): a Write or RegisterConn racing Loop.Close no longer uses the worker's eventfd or epoll fd number after shutdown closed it (celeris#862) #932 states in the test's comment and its body that coverage rests on-race. (from fix(eventloop): a Write or RegisterConn racing Loop.Close no longer uses the worker's eventfd or epoll fd number after shutdown closed it (celeris#862) #932 review round 1, correctness NIT 1)Added 2026-10-03 (#932's round-2 review, lane D1):
TestRegisterConnRacingCloseNeverAddsToAClosedEpoll862andTestRegisterConnRacingUnregisterConnLeavesNoEpollEntry862(driver/internal/eventloop/wake_close_862_linux_test.go) run their hook attestHookBeforeAdd, whichRegisterConncalls before it takesc.mu(loop_linux.go:328at dce5a3c). So they do not pin that thec.closedcheck and theEPOLL_CTL_ADDare one critical section underc.mu(:346-361). A refactor that checksc.closedunderc.mu, unlocks, and then issues the ADD passes every eventloop: a Write racing Loop.Close can write the closed eventfd's number in wake(), which another socket may hold by then (data race) #862 test. The correctness review's mutant CHECKOUTSIDE (that shape) passed all three tests in 20 of 20-raceprocesses, 60/0/0, with 0 race reports (evidencelanes-20261003/D1/review-correctness-r2/logs/r2a/mut-CHECKOUTSIDE.log). Add a deterministic pin in the style ofTestTeardownMarksTheConnClosedBeforeItLeavesTheWorker784. A nil-by-default hook inside thec.musection, after the check and before the ADD, holds the section while the test callsLoop.Closeon another goroutine. The test asserts thatClosehas not returned while the hook is held, then releases it and checks that the ADD landed on the worker's own epoll instance. CHECKOUTSIDE must fail it. (from fix(eventloop): a Write or RegisterConn racing Loop.Close no longer uses the worker's eventfd or epoll fd number after shutdown closed it (celeris#862) #932 review round 2, correctness NIT)Added 2026-10-04 (from #938's review round 1, lane M443):
//go:build !racefiles never run in CI: everygo testin ci.yml and test-coverage.yml, andmage test, passes-race. They hold the alloc guards of driver/{memcached,postgres,redis}, middleware/{cache,idempotency,overload,ratelimit,timeout} and celeristest (added by refactor!: the engine, protocol and driver codec packages move under internal/ (celeris#443) #938), plus the race_off files. That is how refactor!: the engine, protocol and driver codec packages move under internal/ (celeris#443) #938's commit b shipped green whilemiddleware/timeout'sTestNoTimeoutCommonPathSingleAllocfailed (3 allocs/op, want <=1; fixed in refactor!: the engine, protocol and driver codec packages move under internal/ (celeris#443) #938 by e55fa03). Add a short non-race step to the Unit job, for examplego test -count=1 -run 'TestAllocBudgets|TestAllocBudgetsPreparedExec|TestAllowUnderLimitZeroAlloc|TestNoTimeoutCommonPathSingleAlloc|TestNewContextZeroAlloc' ./driver/... ./middleware/... ./celeristest/, about 10 s in a container. (from refactor!: the engine, protocol and driver codec packages move under internal/ (celeris#443) #938 r1, MINOR)TestServerMixedAsync_ConcurrentNoCrosstalk(server_async_test.go:134) flakes on darwin with "connection reset by peer" on the std engine:go test -count=10 -run 'TestServerMixedAsync_ConcurrentNoCrosstalk$' .failed 2 of 10 at main a43aa1d and 1 of 10 at refactor!: the engine, protocol and driver codec packages move under internal/ (celeris#443) #938's head, on darwin/arm64. It passes on Linux, where CI runs it. (from refactor!: the engine, protocol and driver codec packages move under internal/ (celeris#443) #938 r1, NIT)api/golden files (ci: check the exported API against golden files in api/, so every API change shows in review (celeris#443) #914) do not show five kinds of breaking change, whichapi/README.mdlists for review by hand: a struct that stops being comparable; the first unexported field added to an all-exported struct; reordered exported fields (both break unkeyed literals); a renamed package clause; methods promoted through an unexported embedded interface of another module. Each fix changes what apidump reports, so each comes with amage apiregeneration; the plan's estimate is 2-3 h for all five. (from ci: check the exported API against golden files in api/, so every API change shows in review (celeris#443) #914's review MINORs, decision D-F in refactor!: the engine, protocol and driver codec packages move under internal/ (celeris#443) #938)./internal/engine/iouringin the samego test -race -coveras the rest of ci.yml's unit set, at the runner's 8 MiB memlock, and that package sometimes cannot set up a ring:io_uring_setup: cannot allocate memory (likely RLIMIT_MEMLOCK ...). It failedTestHandoffHasNothingInFlightat refactor!: the engine, protocol and driver codec packages move under internal/ (celeris#443) #938's ee8ae8b (run 37162905087, 2 of 2 attempts) andTestAsyncH2CUpgradeOnPromotionLeavesH1StateToTheGoroutineat refactor!: the engine, protocol and driver codec packages move under internal/ (celeris#443) #938's 998d7bf (run 37169188620, attempt 1; the rerun passed). Neither commit touches that package's code. ci.yml's Unit job runs the package in a step of its own and has not hit it. The likely cause, not proven, is the other test binaries of the samego testholding io_uring memory at the same time. Fix: run./internal/engine/iouringin its own step in test-coverage.yml as ci.yml does, with its own profile and upload, and keep the package-set check in step. (from refactor!: the engine, protocol and driver codec packages move under internal/ (celeris#443) #938 r2, CI)