Consolidated on 2026-10-02 from the per-PR 'Follow-ups from #N' bundles, by maintainer decision. From now on, review nits are fixed inside the PR before it merges, so no new bundles are filed. Each line names the bundle and item it came from; the bundle's thread keeps the full context. Everything here stays in v1.6.0.
Note in merged PR fix(websocket): apply pause/resume under the lock that decides them, and lift a stale pause (celeris#667, celeris#672) #671 's A/B section that its pre-registration stated no predicted outcome, as a deviation from the review's MAJOR 2 request. (from Follow-ups from #671: measure the engine-worker stall under pausedMu, the epoll tail lean, #705-repro vs M5 coverage, prereg predictions #716 4)
Optional: if a ~1% empirical bound on fix(engine): keep serving on a paused listener for 1.5 s with TCP_DEFER_ACCEPT cleared, so a switch or PauseAccept no longer resets clients that had not yet sent a request (celeris#662, #675) #674 's cost is wanted before release, run B1 on bare metal (needs probatorium#434 perf + perf_event_paranoid and core isolation); the cluster A/B resolves only 1.0-4.4%. (from Follow-ups from #674: stale ci.yml ramp-quarantine comment, stall write-up accuracy, B1 instrument note #724 3)
Traceability only: fix(epoll): never park the loop on a running async handler, and leave the live set to the loop on an async hijack (celeris#669, celeris#668) #698 's PR comment claims local vet, cross-build and golangci-lint with no saved log in the lane's round2/ evidence (CI corroborates it); save the log or note it. (from Follow-ups from #698: shutdown leaves a deferred transplant's descriptor open (pre-existing), EPOLLOUT comment premise, CI visibility of the new test, stale comments #727 6)
Record fix: clone the request values a detached stream keeps (websocket Conn.Query, SSE Last-Event-ID, Context.Detach, requestid and otel context values) (#714, #717, #718) #723 's deliberately red CI runs 36320632412 (9a89e0c ) and 36321102421 (50825c1 ), paired with green 36321745156, in the maintainers' deliberate-red ledger. The squash-merge requirement is already met (fix: clone the request values a detached stream keeps (websocket Conn.Query, SSE Last-Event-ID, Context.Detach, requestid and otel context values) (#714, #717, #718) #723 landed as the single squash commit 3fe9620 ). (from Follow-ups from #723: otel traceparent clone cost, SSE OnConnect and session c.Set views, stale #714/#718 scope, ENOMEM claim wording, deliberate-red record #740 6)
No CodSpeed benchmark runs HTTP/2 request code since ci: take the ns-scale micro-benchmarks out of CodSpeed, and fix the #699 follow-ups (celeris#725) #748 removed BenchmarkInternH2HeaderName; if wanted, add a us-scale benchmark (one HPACK-decoded request's headers through processor.go) and add protocol/h2/** back to codspeed.yml's trigger paths. (from Follow-ups from #748: decide the CodSpeed budget lever now, the fork-PR exemption for org members, H2 benchmarks, bench-ab.sh result checks #754 3)
Add a two-P BenchmarkAsyncFeedHeaderDeadline variant where a second goroutine locks/unlocks cs.detachMu each iteration, so the +5.4 ns TryLock cost is measured under cross-core contention. (from Follow-ups from #743: epoll checkTimeouts h1State race, arm-path test coverage, a contended bench for the TryLock feed #762 3)
In the next perf checkpoint, read the epoll driver cells (redis, memcached, postgres) against the checkpoint before fix(epoll): never read a driver conn's descriptor number after UnregisterConn has returned (celeris#710) #772 , including a writer-contended shape; fix(epoll): never read a driver conn's descriptor number after UnregisterConn has returned (celeris#710) #772 's merge-gating cluster timing was INCONCLUSIVE (A/A band up to +-11%). (from Follow-ups from #772: the standalone driver event loop reads a reused descriptor number after UnregisterConn (#710 there), the WorkerLoop contract sentence, a test for the check-and-read critical section, UnregisterConn doc #784 5)
Context.Hijack calls cloneRequestValues() with no engine check (context_response.go:1280), so every std WebSocket upgrade pays a copy of values that are already copies (mock bench 257.6 ns to 1810 ns, 37 allocs); measure a std upgrade before/after and skip the copy on std if it matters. (from Follow-ups from #773: an AsyncHandlers arm after #774, io_uring multishot strings read before Hijack, the copy on std, the engine-side cost, doc and witness nits #785 3)
Measure the engine-side cost of fix: keep a hijacked request's receive buffer out of the pool on epoll and io_uring, and copy the request values at Hijack (celeris#733) #773 (epoll gives up a receive buffer, io_uring a connState per hijack) with a hijack-then-accept benchmark per native engine, main vs fix: keep a hijacked request's receive buffer out of the pool on epoll and io_uring, and copy the request values at Hijack (celeris#733) #773 , or bound it with a number. (from Follow-ups from #773: an AsyncHandlers arm after #774, io_uring multishot strings read before Hijack, the copy on std, the engine-side cost, doc and witness nits #785 4)
Run and record the bare-metal failing-first for io_uring: a sync-mode HTTP/1 connection closed in the iteration that parks a paused worker gets no FIN until the worker wakes (the cancelled recv still holds the file; argued: its DEFER_TASKRUN completion needs a GETEVENTS enter the park never makes) #712 as main before fix(iouring): never release a descriptor number while an op can still resolve it: close paths, hijack, shutdown (celeris#685) #793 (1fdfcd4 ) against main after it (3e7abba ) on the cluster; the queued rows for d9307b5 and 8013189 test a park change that was dropped and no longer apply. (from Follow-ups from #767: a named CI interlock for the #712 tests, and the bare-metal failing-first #795 2)
Find the commit between 9f4d89b and 698bed6 that moved the async io_uring standby's drain from ~1 s to ~11 s after a demote (stby0_s; candidates fix(engine): keep serving on a paused listener for 1.5 s with TCP_DEFER_ACCEPT cleared, so a switch or PauseAccept no longer resets clients that had not yet sent a request (celeris#662, #675) #674 49d2726 , fix(epoll): never park the loop on a running async handler, and leave the live set to the loop on an async hijack (celeris#669, celeris#668) #698 , fix(iouring): run every driver op through the engine's own duplicate of the socket, and count every cancel until its CQE, so closing after UnregisterConn is safe (celeris#691, celeris#707) #696 ), e.g. by bisecting the revert-cell probe, and state whether it is intended. Investigation; no tracker elsewhere. (from Follow-ups from #768: named CI interlocks for the #713 tests and epoll twins, and why the async standby's drain grew from ~1 s to ~11 s #797 2)
Run the queued bare-metal A/B for lane 685 (cluster.tsv, fix(engine): keep serving on a paused listener for 1.5 s with TCP_DEFER_ACCEPT cleared, so a switch or PauseAccept no longer resets clients that had not yet sent a request (celeris#662, #675) #674 ABA template, both arches) of the deferred close's churn cost (fdOwed, closeFDOwed park gate, fast-path shutdown(SHUT_RD)) and record its verdict; laptop rows cannot exclude a <5%/<10% regression. (from Follow-ups from #793: a SEND_ZC notification counted as an owed op, the unpinned park gate and shutdown drain, and two counter/exit nits #798 7)
Include TestHandoffHasNothingInFlight in the cluster row that runs the lane's merged tree (one unattributed read_timeout failure, 1 of 37 on the merged tree vs 0 of 27 on main), and file a separate issue if it recurs on bare metal. (from Follow-ups from #799: assert the h2c arms' outcome, write down the exit contract, a rerunHandOff witness, an unattributed merged-tree failure, two flakes on main #816 4)
Optionally add a -benchmem micro-benchmark of rerunHandOff's exiting branch on an fdlFixture conn for the new atomic loads (off the per-request path; CodeRabbit on fix(iouring): leave a conn whose dispatch goroutine exited to close it, or upgraded it to h2c, to its queued exit (celeris#780) #799 ). (from Follow-ups from #799: assert the h2c arms' outcome, write down the exit contract, a rerunHandOff witness, an unattributed merged-tree failure, two flakes on main #816 7)
Correct fix(epoll, iouring): send a large response whole, keep pipelined responses in order, never send a body the handler has given back (celeris#761, celeris#802, celeris#817) #805 's PR-body claims: the per-cell H2 failing-first numbers were single samples of a timing-dependent result on epoll/adaptive, and 'well under the end-to-end noise floor' was inferred from a micro-benchmark with the benchmark-tier row unrun and the H2 queue loop's per-turn resync unmeasured (record the correction where the claims are cited). (from Follow-ups from #805: the memory bound after the per-request cap, epoll's closing-conn reap by read time, no deterministic H2 refusal test, io_uring's dormant WRITEV path, body nits #818 5)
Add a BenchmarkWriteHooks case that forces a short write (small-sndbuf socket or pipe) so epoll's partial-writev branch, where unstage copies the body remainder into writeBuf, is measured with -benchmem (CodeRabbit on fix(epoll, iouring): send a large response whole, keep pipelined responses in order, never send a body the handler has given back (celeris#761, celeris#802, celeris#817) #805 ). (from Follow-ups from #805: the memory bound after the per-request cap, epoll's closing-conn reap by read time, no deterministic H2 refusal test, io_uring's dormant WRITEV path, body nits #818 6)
fix(epoll, iouring, std): graceful shutdown waits for the HTTP/2 streams on the shared worker pool, and std for its h2c streams (celeris#759) #808 's body cost claim should read 'within +-5 %' rather than 'no measurable cost' (BenchmarkBridgeServeHTTP head bias -4 to -5.5 %, BenchmarkPoolDispatch spread +-11-22 %); also note the GOAWAY wire test catches the m1 mutant on io_uring only intermittently (PR merged, body/record fix only). (from Follow-ups from #808: std h2c conns get no GOAWAY and outlive the drain, a pool-wait bound test that passes on main, cost wording #820 3)
fix(iouring): hold a closed connection's send buffer until its SEND_ZC notification, past the release backstop (celeris#812) #813 's nothing-held cost (one load+compare per send in prepSendSQE, a few branches per closed conn with an op owed) is queued as a bare-metal same-run ABA with A/A twin, cluster row 63 (reader analyze_cluster813.py); still 'queued' in evidence/_queue/cluster.tsv on 2026-10-02. Run it and record the verdict. (from Follow-ups from #813: the SEND_ZC hold's memory-bound comment after #805, a WARN that becomes a Debug line under the #801 lag #845 cost-ABA)
fix(api, static): answer an unsatisfiable range with 416 and honour If-Range (celeris#435) #834 body: relabel the 'TestFileRange435 (root)' column (it also counts TestContextFileRange: 36+1=37), and say in the Cost section that the load gate is checked only before round 1 (window 2 load 2.45-3.22 after rounds 4-10; same point as Follow-ups from #841: HEAD now reaches #837 and #835, HEAD on Detach-streaming GET routes, the per-request isHEAD read on live streams, FullPath sentinels, OPTIONS opt-out #847 item 7). (from Follow-ups from #834: If-Range on pre-compressed .br/.gz variants (rest of #435), the IfRange date-safety comment, an unsatisfiable If-Range test, the stat/open window #846 5)
fix(router, conn, sse): answer HEAD with the GET route and OPTIONS with the Allow list; a streamed HEAD sends no body (celeris#421, celeris#833) #841 body Cost section: say the load gate (<2.0) is checked only before round 1 and give the per-round load (1.76-1.85 after rounds 1-3, 2.45-3.22 after 4-10, bench/run-e410772-521d31d); say the adapters' new bool test in Write/Close (response.go:964,:982,:1031) is unmeasured and why that is acceptable. (from Follow-ups from #841: HEAD now reaches #837 and #835, HEAD on Detach-streaming GET routes, the per-request isHEAD read on live streams, FullPath sentinels, OPTIONS opt-out #847 7)
Evidence logs/tools-vulns.log:6 names f6d6eb3 (a local commit amended into 7ae3467 , on no remote); re-run scripts/tools-vulns.sh at 7ae3467 or note the identical tree (f31d374). (from Follow-ups from #839: an unchecked Dependabot allow list for .github/tools, the Scorecard side of its security trade-off, inline go install pins, a stale evidence label #853 4)
driver/internal/eventloop/loop_linux.go forget (fix(eventloop): never read a driver conn's descriptor number after UnregisterConn has returned (celeris#784) #843 ) runs EPOLL_CTL_DEL under the worker's exclusive w.mu.Lock (fe9264f ran it after Unlock), so each teardown stalls dispatch and lookups for one epoll_ctl; UnregisterConn takes w.mu twice. Add a register/unregister churn benchmark next to read_cost_784_bench_linux_test.go with a reader on another conn of the same worker, compare base/head under the timing lock; if it shows, do the identity check and DEL under RLock and Lock only for the map delete. (from Follow-ups from #843: the cost of forget's EPOLL_CTL_DEL under the map lock, the test hook on the read path, the repro's thin P arm, a worker starved by one conn's inflow #860 1)
Repro run 20261002T142740Z arm 4-P opened 27,214 redis conns vs 735,499 for 1-P (and half the probes), unexplained (port/TIME_WAIT pressure?). Before reusing the repro as a witness, check ss -s and the port range, repeat arm order P F P F with pauses, and make the tally refuse an arm whose conn count is far below the others. (from Follow-ups from #843: the cost of forget's EPOLL_CTL_DEL under the map lock, the test hook on the read path, the repro's thin P arm, a worker starved by one conn's inflow #860 3)
The embedded Swagger UI default page now serves ~2,039,923 bytes uncompressed from the origin; behind middleware/compress (zstd/br/gzip, MinLength 256, no size cap, no result cache) every bundle request recompresses 1.5 MB (brotli level 6 for br-only clients). Measure CPU per bundle request per encoding, then choose precompressed embedded variants (measure binary growth) or a doc note recommending a cache/CDN. Not measured yet; would be a defect if the cost is large on a public page. (from Follow-ups from #851: Scalar's AssetsPath file name, re-compressing the embedded bundle behind compress, Scalar's default third-party origins, a relative default spec URL #861 2)
Added 2026-10-02 18:40Z from #860 comment 5958063514 , the second review of #843 at f718f05 (posted 4 min before the consolidation closed #860 ):
fix(eventloop): never read a driver conn's descriptor number after UnregisterConn has returned (celeris#784) #843 's field repro: in a second session the drift repeats, and the last arm is thin whichever binary runs it (F P F: 1,310,091 / 142,830 / 12,951 conns). Only 1-F vs 2-P is valid evidence (from Follow-ups from #843: the cost of forget's EPOLL_CTL_DEL under the map lock, the test hook on the read path, the repro's thin P arm, a worker starved by one conn's inflow #860 item 3, new data).
fix(eventloop): never read a driver conn's descriptor number after UnregisterConn has returned (celeris#784) #843 's cost on amd64: in the next perf checkpoint, read probatorium's driver_redis refapp cells on both arches against the checkpoint before fix(eventloop): never read a driver conn's descriptor number after UnregisterConn has returned (celeris#784) #843 (from Follow-ups from #843: the cost of forget's EPOLL_CTL_DEL under the map lock, the test hook on the read path, the repro's thin P arm, a worker starved by one conn's inflow #860 item 10).
Added 2026-10-03 (#923 's round-1 reviews, lane F3T):
Added 2026-10-03 (#906 's round-2 verification at 1c1cce4 , lane F3T):
fix(h2): bound the response bytes a connection holds for the peer's window; over the budget, handlers run on the pool and send as the windows open (celeris#893) #906 (protocol/h2/stream/outbound.go:193-201, one time.AfterFunc per wait that blocks): BenchmarkWindowUpdateWithWaiters893 keeps every waiter's stream window at 0, so no waiter ever returns and re-waits, and the per-wait timer is never measured. A probe of the reading-client regime (stream windows open, the connection window the bottleneck), BenchmarkProbeConnWindowRace906 at waiters=1, went from 281.6 ns to 731.9 ns sec/op (+159.86%, p=0.000) and 325.3 ns to 942.6 ns CPU (+189.76%, p=0.000) per connection WINDOW_UPDATE, round 1 eb93228 against 1c1cce4; A/A ~ (p=0.481); at waiters=99 the head is faster (-25.83% sec/op). That is about +0.6 µs of CPU per connection WINDOW_UPDATE with one waiter, off the dispatch path and only on connections over the budget. The attribution to the timer is by code reading only: run the same-binary zero-deadline arm (lanes-20261003/F3T/verify906-perfsec/scripts/72-timing-attrib.sh, 73-timing-attrib-session.sh, scripted, not run) under the TIMING lock and record whether the timer accounts for the difference. Probe verify906-perfsec/probes/zz_probe_conn_window_race_906_test.go.tmpl, logs verify906-perfsec/logs/bench/. (from fix(h2): bound the response bytes a connection holds for the peer's window; over the budget, handlers run on the pool and send as the windows open (celeris#893) #906 verify r2, perf/security MINOR 2)
Consolidated on 2026-10-02 from the per-PR 'Follow-ups from #N' bundles, by maintainer decision. From now on, review nits are fixed inside the PR before it merges, so no new bundles are filed. Each line names the bundle and item it came from; the bundle's thread keeps the full context. Everything here stays in v1.6.0.
Added 2026-10-02 18:40Z from #860 comment 5958063514, the second review of #843 at f718f05 (posted 4 min before the consolidation closed #860):
Added 2026-10-03 (#923's round-1 reviews, lane F3T):
BenchmarkCacheHit128 B/3,MissSingleflight296 B/7,MissNoSingleflight280 B/6 on both trees;lanes-20261003/F3T/review923-perfsec/logs/31-alloc-counts.log). The timing run was abandoned because the host never went quiet within 20 min (logs/30-bench.out). Statically the MISS paths do the same operations in a new order with no new allocation,sf.Call's result shrinks from 32 B to 24 B, and the HIT path is untouched. Run it in the next quiet window or perf checkpoint. (from fix(cache): the coalesced fill stores the response and releases its followers before any client is written, so a slow leader's client no longer holds them (celeris#913) #923 review r1, perf/security NIT 3)Added 2026-10-03 (#906's round-2 verification at 1c1cce4, lane F3T):
protocol/h2/stream/outbound.go:193-201, onetime.AfterFuncper wait that blocks):BenchmarkWindowUpdateWithWaiters893keeps every waiter's stream window at 0, so no waiter ever returns and re-waits, and the per-wait timer is never measured. A probe of the reading-client regime (stream windows open, the connection window the bottleneck),BenchmarkProbeConnWindowRace906at waiters=1, went from 281.6 ns to 731.9 ns sec/op (+159.86%, p=0.000) and 325.3 ns to 942.6 ns CPU (+189.76%, p=0.000) per connection WINDOW_UPDATE, round 1eb93228against1c1cce4; A/A ~ (p=0.481); at waiters=99 the head is faster (-25.83% sec/op). That is about +0.6 µs of CPU per connection WINDOW_UPDATE with one waiter, off the dispatch path and only on connections over the budget. The attribution to the timer is by code reading only: run the same-binary zero-deadline arm (lanes-20261003/F3T/verify906-perfsec/scripts/72-timing-attrib.sh,73-timing-attrib-session.sh, scripted, not run) under the TIMING lock and record whether the timer accounts for the difference. Probeverify906-perfsec/probes/zz_probe_conn_window_race_906_test.go.tmpl, logsverify906-perfsec/logs/bench/. (from fix(h2): bound the response bytes a connection holds for the peer's window; over the budget, handlers run on the pool and send as the windows open (celeris#893) #906 verify r2, perf/security MINOR 2)