You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
#725 item 3 asked to check the organization's real CodSpeed minutes and decide whether pull request runs need the performance-label gate. #748 fixes the rest of #725 but cannot make that decision, so it is tracked here, with the burn measured so far. Two more things from #748's review are listed below.
1. CodSpeed budget: the decision is needed now, not after a month
The Free plan gives the organization 600 macro-runner minutes a month, shared by celeris and loadgen. These are the job times GitHub recorded for every CodSpeed job from the first one (2026-09-26 15:57Z) to 2026-09-27 18:49Z, 26.9 hours:
repo
jobs
pull_request
push
dispatch
minutes, raw
minutes, rounded up per job
celeris
21
12
6
3
94
101
loadgen
22
17
4
1
69
88
org
43
29
10
4
162
189
At that rate the 600 minutes last 3.6 days (rounded) to 4.1 days (raw). CodSpeed's own meter is not public, so which of the two it counts is unknown.
Pull request runs are 124 of the 189 minutes, pushes to main 47, dispatches 18.
Pushes to main alone, at this merge rate, would use the 600 minutes in about 14 days.
Gate pull request runs on the performance label: add labeled to the pull_request types and contains(github.event.pull_request.labels.*.name, 'performance') to the job's if:, in both repositories. On the data above, that removes about two thirds of the minutes.
Also move main from every push to a daily schedule: run. CodSpeed compares a pull request with the latest main run anyway.
Do nothing and let the minutes run out. What CodSpeed does then (a queued job, a failed job, a paused check) is not documented where we can see it.
Someone with access to the organization's CodSpeed settings should also read the real usage there.
2. The fork-PR approval policy exempts organization members
The approval policy is all_external_contributors on celeris, loadgen, probatorium and docs. GitHub's rule for that setting requires approval only from users who are "not a member or owner of this repository and not a member of the organization". The organization's base repository permission is read, so an organization member without write access to a repository can open a fork pull request there, and its workflows run without approval.
celeris: both organization members have write or admin, so the exemption is not open today.
loadgen: one organization member has read. loadgen's codspeed.yml runs on the same macro-runner group, with the same head.repo.full_nameif:, which a fork pull request can delete. So that member's fork pull request can run code on the macro runner without approval.
Fix, a settings decision: give every organization member write access to each repository in the macro-runner group, or make a read-only collaborator an outside collaborator instead of an organization member.
Check it with gh api orgs/goceleris/members and gh api repos/goceleris/<repo>/collaborators/<login>/permission.
3. HTTP/2 request handling is not benchmarked on CodSpeed
#748 removed BenchmarkInternH2HeaderName (49-64 ns on a byte-identical binary), and it was the only benchmark in the CodSpeed set that ran HTTP/2 request code. If H2 header handling should be watched, add a µs-scale benchmark, for example one HPACK-decoded request's headers through processor.go, and add protocol/h2/** back to codspeed.yml's trigger paths.
The measurements are in the evidence for #748: scripts/job-minutes.sh, scripts/burn.py, scripts/budget_replay.py.
4. bench-ab.sh prints a comparison even when a ref measured nothing
From CodeRabbit on #748 (thread, minor). .github/scripts/bench-ab.sh runs each arm with -test.bench "$re" and then goes straight to benchstat. If re matches no benchmark in one ref (a rename, or a benchmark that exists only at the head), that binary exits 0 with no result lines, and benchstat prints a table with nothing to compare.
Fix: after the rounds, collect the benchmark names from A.txt, B.txt and A2.txt (the ^Benchmark\S+ field of each result line, without the -N GOMAXPROCS suffix). Exit non-zero, naming the missing side, if any set is empty or if A's and B's sets differ. A2 is a copy of A's binary, so its set must equal A's.
Check: run it with a re that only the head has, and with one that matches nothing; both must stop before benchstat.
5. codspeed.yml: a difference inside the A/A floor is inconclusive, not "not a change"
From CodeRabbit on #748 (thread, minor). The header's "How to read a flag" step 2 says "A difference inside the floor is not a change." A real 5% slowdown measured against an 8% A/A floor would be dismissed by that rule. The comparison cannot tell it from noise; it does not show there is no change.
Fix: change the sentence to say that a difference inside the floor is inconclusive, and that the measurement must be repeated or extended (more rounds, a longer -benchtime) until the floor is below the flagged difference, or the flag stays open. Say the same in bench-ab.sh's usage text, which also calls A/A the floor.
#725 item 3 asked to check the organization's real CodSpeed minutes and decide whether pull request runs need the
performance-label gate. #748 fixes the rest of #725 but cannot make that decision, so it is tracked here, with the burn measured so far. Two more things from #748's review are listed below.1. CodSpeed budget: the decision is needed now, not after a month
The Free plan gives the organization 600 macro-runner minutes a month, shared by celeris and loadgen. These are the job times GitHub recorded for every CodSpeed job from the first one (2026-09-26 15:57Z) to 2026-09-27 18:49Z, 26.9 hours:
The options, for the maintainer:
performancelabel: addlabeledto thepull_requesttypes andcontains(github.event.pull_request.labels.*.name, 'performance')to the job'sif:, in both repositories. On the data above, that removes about two thirds of the minutes.schedule:run. CodSpeed compares a pull request with the latest main run anyway.Someone with access to the organization's CodSpeed settings should also read the real usage there.
2. The fork-PR approval policy exempts organization members
The approval policy is
all_external_contributorson celeris, loadgen, probatorium and docs. GitHub's rule for that setting requires approval only from users who are "not a member or owner of this repository and not a member of the organization". The organization's base repository permission isread, so an organization member without write access to a repository can open a fork pull request there, and its workflows run without approval.read. loadgen'scodspeed.ymlruns on the same macro-runner group, with the samehead.repo.full_nameif:, which a fork pull request can delete. So that member's fork pull request can run code on the macro runner without approval.gh api orgs/goceleris/membersandgh api repos/goceleris/<repo>/collaborators/<login>/permission.3. HTTP/2 request handling is not benchmarked on CodSpeed
#748 removed
BenchmarkInternH2HeaderName(49-64 ns on a byte-identical binary), and it was the only benchmark in the CodSpeed set that ran HTTP/2 request code. If H2 header handling should be watched, add a µs-scale benchmark, for example one HPACK-decoded request's headers throughprocessor.go, and addprotocol/h2/**back tocodspeed.yml's trigger paths.The measurements are in the evidence for #748:
scripts/job-minutes.sh,scripts/burn.py,scripts/budget_replay.py.4.
bench-ab.shprints a comparison even when a ref measured nothingFrom CodeRabbit on #748 (thread, minor).
.github/scripts/bench-ab.shruns each arm with-test.bench "$re"and then goes straight tobenchstat. Ifrematches no benchmark in one ref (a rename, or a benchmark that exists only at the head), that binary exits 0 with no result lines, and benchstat prints a table with nothing to compare.A.txt,B.txtandA2.txt(the^Benchmark\S+field of each result line, without the-NGOMAXPROCS suffix). Exit non-zero, naming the missing side, if any set is empty or if A's and B's sets differ. A2 is a copy of A's binary, so its set must equal A's.rethat only the head has, and with one that matches nothing; both must stop before benchstat.5.
codspeed.yml: a difference inside the A/A floor is inconclusive, not "not a change"From CodeRabbit on #748 (thread, minor). The header's "How to read a flag" step 2 says "A difference inside the floor is not a change." A real 5% slowdown measured against an 8% A/A floor would be dismissed by that rule. The comparison cannot tell it from noise; it does not show there is no change.
-benchtime) until the floor is below the flagged difference, or the flag stays open. Say the same inbench-ab.sh's usage text, which also calls A/A the floor.