Skip to content

A perf cell is timed by tm, a cell under 2 s grades the median of three runs, --pin pins the median of five graded runs, and the run times are regraved at their own commits - #1421

Merged
Lorenzobattistela merged 3 commits into
bendlang:mainfrom
nicolas-abril:perf-short-median
Oct 8, 2026

Conversation

@nicolas-abril

@nicolas-abril nicolas-abril commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

The perf gate graded a cell on one timed run: a perl clock before and after /usr/bin/time -l ./cell. On the short GPU cells that single run oscillated by more than the gate's 15% slack, so they failed on some runs and passed on others with the same code: mandelbrot's PAR-GPU (pinned at 0.068 s) read 0.075–0.082 s over 13 runs of main and failed 5 of them.

Where the oscillation came from

Measured on the minis (macOS 26.5.2):

  • The clock. The interval included one perl exiting and another starting: 9.5 ms outside main, against 3.3 ms with one posix_spawn and wait4.
  • The first run after the cell's builds is slow. After bend -o and cc, the run after the warm run pays 10–40 ms more in start-up and exit (mandelbrot read 59.7–73.6 ms over 24 cells; later runs held at about 60). The gate timed exactly that run.
  • Some programs vary run to run. terrain's GPU execution takes about 297 ms, and one run in four takes 213–264 ms, on every mini, with the same output (start-up, pipeline and exit stay flat).

The change

  • tm, a 20-line C program each cell writes and builds with cc, runs the cell with posix_spawn and wait4 and writes its exit status, wall time and maximum RSS. It replaces the perl clocks and /usr/bin/time -l; the RSS comes from the same rusage /usr/bin/time read.
  • A cell whose warm run took under 2 s (SHORT) times three runs and grades their median, on the same mini. That drops the slow first run and terrain's fast outliers. A failing run stops the cell, so its note shows that run's output. Longer cells keep one run, so the cells that set the gate's length do not grow: the four gates run in 22–23 s.
  • --pin runs the gate PIN_RUNS (5) times, 2 s apart, each run dispatched and measured exactly as a graded run, prints every run's times, and pins each cell's median. It used to put three copies of each cell into one run, which share that run's moment: mandelbrot's PAR-GPU pinned at 0.060 s that way, then read 0.060–0.070 s at the commit it was pinned at.

Spread of a cell across 3 gate runs (old timing vs this PR, both graded against the old pins):

cell old this PR
terrain PAR-GPU 0.231–0.314 s (36%) 0.304–0.308 s (1.3%)
tree-matmul PAR-GPU 0.740–0.921 s (25%) 0.893–0.911 s (2.0%)
raytrace PAR-GPU 0.264–0.293 s (11%) 0.265–0.275 s (3.8%)
nbody PAR-GPU 0.069–0.076 s (10%) 0.062–0.064 s (3.2%)
gameoflife PAR-GPU 0.100–0.110 s (10%) 0.088–0.092 s (4.5%)
widest of any cell 36% 10.5% (editdist PAR-GPU)

The pins

The aim is the same baseline under the new method, so only what this PR measures differently is regraved: the SEQ-CPU, PAR-CPU and PAR-GPU times. Each is the median of 5 gate runs, 2 s apart, at the commit its old pin was taken at:

  • 35dbdb75 (the 2026-09-09 pins, stamped 26de367b): every time cell of the 16 original benches, except
  • 386b7fc1 (stamped b20509fd): bitonic, kmeans and nbody SEQ-CPU and PAR-CPU, which it repinned;
  • d04cf05e: histogram's three times.

Each commit ran its own gate (its build line, -fno-slp-vectorize included at 35dbdb7, its flags and bench names) with only this PR's cell timing swapped in, and today's _lib.ts to reach the cluster. The old commits don't have the new --pin loop, so their 5 runs were ordinary gate runs taken one after another and the medians taken from those, which is what the new --pin does. The COMPILER column, the spaces and the checker pins are untouched: this PR doesn't change how they're measured. bitonic, matmul and radix keep their tree- names.

Most times land within 1% of their old pins. The ones that moved:

cell old pin new pin
gameoflife PAR-GPU 0.100 s 0.085 s
nbody PAR-GPU 0.065 s 0.061 s
bfs PAR-GPU 0.956 s 0.894 s
terrain PAR-GPU 0.291 s 0.313 s
histogram SEQ-CPU / PAR-CPU 1.717 / 0.367 s 1.466 / 0.312 s
tree-bitonic SEQ-CPU / PAR-CPU 5.847 / 1.246 s 5.490 / 1.185 s
kmeans SEQ-CPU 3.969 s 3.772 s

mandelbrot's PAR-GPU pin is 0.067 s (it was 0.068). The 5 runs at each commit show no drift from the first run to the last (mandelbrot: 0.065, 0.067, 0.069, 0.067, 0.068).

Checks

  • main + this PR against the new pins: perf 124 / 124 on 5 of 5 runs; the highest cells sit around 1.11x (tree-radix PAR-GPU, tree-bitonic PAR-CPU, and the generics_3200 checker, whose pin is unchanged).
  • Gates on the final commit: repo 54 / 54, ping 47 / 47, test 1593 / 1593, perf 124 / 124.
  • RSS from wait4 matches /usr/bin/time -l (e.g. 13.5 MB on the GPU cells).

Not in this PR

The slot lock lives in /tmp on the machine that starts the gate, and every job runs as the same user in $HOME/bend-perf/<bench>-<mode>, so two people's gates can take the same minis and delete each other's binaries mid-run (one such collision failed a bfs SEQ-CPU cell with exit 127 during this work). A cell sharing a mini with another gate's also runs slower. That needs a lock on the minis and a per-run directory.

🤖 Generated with Claude Code

…er 1 s grades the median of three runs; the apple_m4 pins are regraved at 386b7fc (histogram at d04cf05) under this method

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@Lorenzobattistela

Copy link
Copy Markdown
Collaborator

Thanks @nicolas-abril! The timing change is good: tm plus the median of three takes most of the flicker out of the short GPU cells. Over repeated runs of main + this PR, terrain PAR-GPU spread went from 35% to 2%, raytrace from 13% to 6%, and tree-matmul from 12% to 7%. The gate also takes the same time as before (perf about 19 s, _run.ts about 21 s).

The problem is four of the new PAR-GPU pins. I put this PR's timing into 386b7fc's own perf.ts and ran it 5 times. Even at the commit they were pinned at, these cells read above their pins:

gameoflife 0.085-0.093 s (pin 0.079)
mandelbrot 0.066-0.069 s (pin 0.060)
merkle 0.080-0.083 s (pin 0.075)
nbody 0.062-0.063 s (pin 0.058)

Every other cell landed within 4% of its pin. Current main reads the same as 386b7fc on these four (gameoflife 0.087-0.095, mandelbrot 0.066-0.070), so the mandelbrot and gameoflife misses aren't regressions. They come from the pins. With these pins, perf passed 124/124 on only 2 of 12 runs of main + PR, while main's current gate passed on 6 of 7. Cells between 1 and 2 s that are timed once (tree-bitonic PAR-CPU) can still flicker over 1.15x too.

Smaller things:

  • If a later tm run fails before writing ran.txt, the previous run's line gets printed again.
  • tm doesn't check fopen.
  • When run 1 fails and run 3 passes, the note shows run 3's output.
  • The checker pin stamp says "timed by tm", but the checker is still timed by the perl clock.

Gates on main + this PR: repo 54/54, ping 48/48, test 1593/1593, perf 123/124.

@Lorenzobattistela Lorenzobattistela self-assigned this Oct 8, 2026
nicolas-abril and others added 2 commits October 8, 2026 20:44
…ing result as exit 127 and stops at the first failing run; SHORT is 2 s; the checker pin stamp says the checker keeps the perl clock

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…hed and measured, and pins each cell's median; only the run times are regraved, each at the commit it was pinned at, and COMPILER, the spaces and the checker keep their pins

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@nicolas-abril nicolas-abril changed the title A perf cell is timed by tm, a cell under 1 s grades the median of three runs, and the pins are regraved at their own commits A perf cell is timed by tm, a cell under 2 s grades the median of three runs, --pin pins the median of five graded runs, and the run times are regraved at their own commits Oct 8, 2026
@nicolas-abril

Copy link
Copy Markdown
Collaborator Author

@Lorenzobattistela you were right about the four PAR-GPU pins. --pin sent three copies of each cell into one run, and that run read low: mandelbrot pinned at 0.060 s, while 386b7fc itself then read 0.060–0.070 s.

--pin now runs the gate five times, 2 s apart, each run dispatched and measured like a graded one, and pins each cell's median. Only the run times are regraved, each at the commit its old pin came from (35dbdb7; 386b7fc for bitonic, kmeans and nbody SEQ and PAR; d04cf05 for histogram), as the median of 5 gate runs there. COMPILER, the spaces and the checker keep their pins, since this PR doesn't change how they're measured. mandelbrot's pin is now 0.067 s (it was 0.068), and main + this PR passed perf 124/124 on 5 of 5 runs, with the highest cells around 1.11x (tree-radix PAR-GPU, tree-bitonic PAR-CPU).

The smaller points are fixed in 2baec10: tm checks fopen; each run clears ran.txt first and a missing result reads as exit 127; the loop stops at the first failing run, so the note shows that run's output; the checker stamp says it keeps the perl clock. SHORT is 2 s, so cells between 1 and 2 s (tree-bitonic PAR-CPU) get the median of three too.

One more thing we found: the slot lock lives in /tmp on the machine that starts the gate, and every job runs as the same user in $HOME/bend-perf/<bench>-<mode>, so two people's gates can share minis and delete each other's binaries. Two of our gates collided that way today (exit 127 on bfs SEQ-CPU), and a cell sharing a mini with another gate's runs slower, which may explain part of the run-to-run drift. That's for a separate PR.

@Lorenzobattistela
Lorenzobattistela merged commit 60fa05d into bendlang:main Oct 8, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants