Skip to content

Follow-ups from #674: stale ci.yml ramp-quarantine comment, stall write-up accuracy, B1 instrument note #724

Description

@FumingPower3925

Follow-ups from PR #674 (#662, #675), deferred under the maintainer's two-round review cap (2026-09-27). #674 merges at 7aeb2ec after three re-review rounds that APPROVED, and a cluster A/B (probatorium run 36308789830) with the verdict NO REGRESSION.

  1. The ci.yml ramp-quarantine comment is stale. At ci.yml:456-460 (7aeb2ec) it says "on main without the epoll: PauseAccept silently drops connections already waiting in the accept queue — their requests get EOF (8/8, deterministic), and a promotion on a GitHub runner loses ~1/3 of 2048 connections #662 fix they fail with the accept hand-over's read resets". CI: lift the adaptive job's quarantine of TestRampH1Sync and TestRampH1Async once #674 is on main (>= 6 green GitHub-hosted runs) #708, as corrected in round 4.3, says the opposite: in the r6 gate main passed 12/12 ramp runs, and "main fails this too" is not a known fact. Rewrite the comment to match CI: lift the adaptive job's quarantine of TestRampH1Sync and TestRampH1Async once #674 is on main (>= 6 green GitHub-hosted runs) #708, and keep it accurate when CI: lift the adaptive job's quarantine of TestRampH1Sync and TestRampH1Async once #674 is on main (>= 6 green GitHub-hosted runs) #708 lifts the quarantine.
  2. The stall write-up. The PR body's stall section (section 11) says another lane's container ran beside every one of the 80 stall containers. In fact 27 of the 80 had none, and 40 ran beside this lane's own containers. Several figures in that section (the Fisher p = 0.0023 against D2, the host-load ranges, the 1,501 = 1,501 transplant ledger) are hand-computed and not regenerated by stall/analyze.py. The TL;DR should also carry the weak-control caveat (B 0/30 vs C 1/10, p = 0.25). Fix the text, or name a script for the numbers.
  3. The B1 instrument. B1 (the laptop S1b timing) ended not_run, and the maintainer replaced it with the cluster A/B. That A/B resolves a few percent (engine floors 1.0-4.4%), not B1's ~1%. The static and syscall witnesses (E2, E3b) bound the cost at ~0.02%/op. If a tighter empirical bound is wanted before the release, run B1 on bare metal (Fallback A: needs probatorium#434's perf and perf_event_paranoid, plus core isolation on the hosts).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions