Repository navigation
A perf cell is timed by tm, a cell under 2 s grades the median of three runs, --pin pins the median of five graded runs, and the run times are regraved at their own commits - #1421
Conversation
|
Thanks @nicolas-abril! The timing change is good: tm plus the median of three takes most of the flicker out of the short GPU cells. Over repeated runs of main + this PR, terrain PAR-GPU spread went from 35% to 2%, raytrace from 13% to 6%, and tree-matmul from 12% to 7%. The gate also takes the same time as before (perf about 19 s, _run.ts about 21 s). The problem is four of the new PAR-GPU pins. I put this PR's timing into 386b7fc's own perf.ts and ran it 5 times. Even at the commit they were pinned at, these cells read above their pins: gameoflife 0.085-0.093 s (pin 0.079) Every other cell landed within 4% of its pin. Current main reads the same as 386b7fc on these four (gameoflife 0.087-0.095, mandelbrot 0.066-0.070), so the mandelbrot and gameoflife misses aren't regressions. They come from the pins. With these pins, perf passed 124/124 on only 2 of 12 runs of main + PR, while main's current gate passed on 6 of 7. Cells between 1 and 2 s that are timed once (tree-bitonic PAR-CPU) can still flicker over 1.15x too. Smaller things:
Gates on main + this PR: repo 54/54, ping 48/48, test 1593/1593, perf 123/124. |
…ing result as exit 127 and stops at the first failing run; SHORT is 2 s; the checker pin stamp says the checker keeps the perl clock Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…hed and measured, and pins each cell's median; only the run times are regraved, each at the commit it was pinned at, and COMPILER, the spaces and the checker keep their pins Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
|
@Lorenzobattistela you were right about the four PAR-GPU pins.
The smaller points are fixed in 2baec10: tm checks One more thing we found: the slot lock lives in /tmp on the machine that starts the gate, and every job runs as the same user in |
The perf gate graded a cell on one timed run: a perl clock before and after
/usr/bin/time -l ./cell. On the short GPU cells that single run oscillated by more than the gate's 15% slack, so they failed on some runs and passed on others with the same code: mandelbrot's PAR-GPU (pinned at 0.068 s) read 0.075–0.082 s over 13 runs of main and failed 5 of them.Where the oscillation came from
Measured on the minis (macOS 26.5.2):
main, against 3.3 ms with oneposix_spawnandwait4.bend -oandcc, the run after the warm run pays 10–40 ms more in start-up and exit (mandelbrot read 59.7–73.6 ms over 24 cells; later runs held at about 60). The gate timed exactly that run.The change
cc, runs the cell withposix_spawnandwait4and writes its exit status, wall time and maximum RSS. It replaces the perl clocks and/usr/bin/time -l; the RSS comes from the samerusage/usr/bin/timeread.SHORT) times three runs and grades their median, on the same mini. That drops the slow first run and terrain's fast outliers. A failing run stops the cell, so its note shows that run's output. Longer cells keep one run, so the cells that set the gate's length do not grow: the four gates run in 22–23 s.--pinruns the gatePIN_RUNS(5) times, 2 s apart, each run dispatched and measured exactly as a graded run, prints every run's times, and pins each cell's median. It used to put three copies of each cell into one run, which share that run's moment: mandelbrot's PAR-GPU pinned at 0.060 s that way, then read 0.060–0.070 s at the commit it was pinned at.Spread of a cell across 3 gate runs (old timing vs this PR, both graded against the old pins):
The pins
The aim is the same baseline under the new method, so only what this PR measures differently is regraved: the SEQ-CPU, PAR-CPU and PAR-GPU times. Each is the median of 5 gate runs, 2 s apart, at the commit its old pin was taken at:
35dbdb75(the 2026-09-09 pins, stamped 26de367b): every time cell of the 16 original benches, except386b7fc1(stamped b20509fd): bitonic, kmeans and nbody SEQ-CPU and PAR-CPU, which it repinned;d04cf05e: histogram's three times.Each commit ran its own gate (its build line,
-fno-slp-vectorizeincluded at 35dbdb7, its flags and bench names) with only this PR's cell timing swapped in, and today's_lib.tsto reach the cluster. The old commits don't have the new--pinloop, so their 5 runs were ordinary gate runs taken one after another and the medians taken from those, which is what the new--pindoes. The COMPILER column, the spaces and the checker pins are untouched: this PR doesn't change how they're measured. bitonic, matmul and radix keep their tree- names.Most times land within 1% of their old pins. The ones that moved:
mandelbrot's PAR-GPU pin is 0.067 s (it was 0.068). The 5 runs at each commit show no drift from the first run to the last (mandelbrot: 0.065, 0.067, 0.069, 0.067, 0.068).
Checks
wait4matches/usr/bin/time -l(e.g. 13.5 MB on the GPU cells).Not in this PR
The slot lock lives in
/tmpon the machine that starts the gate, and every job runs as the same user in$HOME/bend-perf/<bench>-<mode>, so two people's gates can take the same minis and delete each other's binaries mid-run (one such collision failed a bfs SEQ-CPU cell with exit 127 during this work). A cell sharing a mini with another gate's also runs slower. That needs a lock on the minis and a per-run directory.🤖 Generated with Claude Code