Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 3 additions & 2 deletions .claude/agents/perf.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,8 +31,9 @@ it is never unmeasured again. You report measured figures only.
* **The scaling ceiling.** Refine speeds up at most about 2.0-2.2× from 1 to
20 threads, flat from about 7-8. The 2026-09-26 sweep
(`docs/benchmarks/2026-09-26/scaling/`, medians per thread count) gives 2.2×
on AC and 2.1× on battery; 2026-09-27 gives 2.02-2.03× on battery (`docs/benchmarks/2026-09-27/serial-profile/README.md`). Report
each increment's ceiling against the matching power-state baseline.
on AC and 2.1× on battery; 2026-09-27 gives 2.02-2.03× on battery
(`docs/benchmarks/2026-09-27/serial-profile/README.md`). Report each
increment's ceiling against the matching power-state baseline.

## 2. How a run is made
* **Release build, rebuilt.** `bench.py run` builds Release into
Expand Down
8 changes: 6 additions & 2 deletions .github/workflows/main.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -70,7 +70,8 @@ jobs:

# Increment 14, R7: refinement is the first code path that starts threads,
# and asan+ubsan cannot see a data race. TSan cannot share a build with ASan,
# so this is its own job, building only the four refinement suites. T6 runs
# so this is its own job, building only the suites listed below (refinement and
# the parallel helpers they use). T6 runs
# the loop at 1, 2, 7 and all threads; a clean run is the independent check
# of the design's "no shared writes" claim. The binaries are run directly
# because ctest names the tests by Catch case, not by target.
Expand Down Expand Up @@ -102,6 +103,7 @@ jobs:
test_mesh_row_spans prop_refinement_scan_equivalence
test_mesh_quality prop_refinement_quality
test_refinement_constraint_feet
test_refinement_chunks_dynamic test_refinement_active test_mesh_lawson_stack

- name: Test
env:
Expand All @@ -112,7 +114,9 @@ jobs:
test_refinement_scan_offnode test_refinement_refine \
test_mesh_row_spans prop_refinement_scan_equivalence \
test_mesh_quality prop_refinement_quality \
test_refinement_constraint_feet; do
test_refinement_constraint_feet \
test_refinement_chunks_dynamic test_refinement_active \
test_mesh_lawson_stack; do
build-tsan/tests/cpp/$suite
done

Expand Down
2 changes: 1 addition & 1 deletion ROADMAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,7 +42,7 @@ increment that most needs a picture to check against
| 20 | Quality start: before DEM refinement, Steiner points at the DEM node nearest each bad triangle's circumcentre until the start mesh has a 25° minimum angle (`--start-min-angle`, 0 is off); geometry only, input segments not split, serial and deterministic. Removes increment 16's boundary fans | landed with 20b as interim (C1-C3 open, to 20c; C4 -> 20b) | `docs/increments/20-start-quality.md` |
| 20b | Minimum insertion distance from constraints: when refinement's worst node lies within ε = clamp(tol / slope, cell/100, cell/2) of a constraint segment, insert the foot on the segment (off-node, bilinear z) instead; a footed node still above tolerance is inserted after all, so the tolerance guarantee is unchanged. Removes the 0.0117° needle at 1 m (increment 20's C4) | shipped with 20 (#97) | `docs/increments/20b-min-insertion-distance.md` |
| — | `tools/bench.py`: the 1 m benchmark and the thread-scaling sweep from one checked-in command, with power state, quality and commit recorded per run; the one-off scripts in `docs/benchmarks/2026-09-26/` are its specification. Rule 2's acceptance run needs it | shipped with branch `tools-bench`'s PR | `docs/benchmarks/bench-py.md` |
| — | The serial phase: profile refine's serial insert-and-flip phase, then parallelise what the profile blames. Scaling tops out at about 2.0-2.2×. Profiled 2026-09-27: serial part about a third of 1-thread refine, mostly Lawson legalisation; the scan stops near 5× from load imbalance | profiled; designed as increment 21 (Ola's rulings 2026-09-27: L1 determinism, one path, at most 2 % more triangles): 21a and 21b quick wins next, then 21c measurements, then 21d | `docs/increments/21-parallel-refine.md`, `docs/benchmarks/2026-09-27/serial-profile/README.md` |
| — | The serial phase: profile refine's serial insert-and-flip phase, then parallelise what the profile blames. Scaling tops out at about 2.0-2.2×. Profiled 2026-09-27: serial part about a third of 1-thread refine, mostly Lawson legalisation; the scan stops near 5× from load imbalance | profiled; designed as increment 21 (Ola's rulings 2026-09-27: L1 determinism, one path, at most 2 % more triangles). **21a shipped with branch `increment21a-quick-wins`'s PR**: dynamic scan blocks, active merge, reused flip stack; mesh bit-identical; on AC refine -10 to -15 % at 8 threads, ceiling 2.1x -> 2.3x (`docs/benchmarks/2026-09-27/21a-acceptance.md`). 21b next, then 21c measurements, then 21d | `docs/increments/21-parallel-refine.md`, `docs/benchmarks/2026-09-27/serial-profile/README.md` |
| 16b | Interior polygons and polylines as constraints ("terrain polygons": lakes, land cover, roads, rivers): `--features PATH`, a GeoJSON `FeatureCollection`, each feature naming a vocabulary property; closed or open `Breakline`s, crossings noded, off-node vertices with bilinear z. **Its working example is real data** (Ola, 2026-09-27): CORINE Land Cover 2018 over the benchmark tile `7908_3_10m_z33.tif`. The source is Ola's local copy, `rasputin_data/corine_sql/.../U2018_CLC2018_V2020_20u1.gpkg` (8.2 GB, EPSG:3035, a sibling of this repository, which also holds the 254-tile DTM10 archive for gap 6). It reads without GDAL: sqlite3 over its R-tree, the GeoPackage blob header stripped, `shapely.wkb`, then `pyproj` to 25833, so CRS stops in Python as before. Probed 2026-09-27 in 0.2 s: 60 polygons in 8 classes (heath, bare rock, sparse vegetation, bogs, intertidal flats, water, sea, urban), 11 068 vertices clipped to the tile, median segment 54 m against 10 m cells. The EEA's public ArcGIS service (`image.discomap.eea.europa.eu`, `Corine/CLC2018_WM`) returns the same 11 068 clipped vertices and is the route for anyone without the file. What it forces on 16b's design: neighbouring polygons share their boundaries, so each shared edge arrives twice; the polygons run past the domain and must be clipped; the extract is committed as a fixture with the Copernicus attribution. Placed before 20c because 20c may split constraint segments and should be designed and measured on inputs that have interior ones | designed in 16's R6, no increment file yet; after the serial phase, before 20c | `docs/increments/16-domain-polygon.md` (R6) |
| 20c | Soft quality criterion: a penalty that each Steiner node or constraint split must pay for in angle gained, instead of 20's hard 25°; applied at the start and during DEM refinement; may split constraint segments when that improves the mesh. Ola's rulings on 20's C1-C3 | to design after 16b (`@architect` measures cost against 20 first) | `docs/increments/20-start-quality.md` (Ola's rulings) |
| — | Auto-catchment: the watershed upstream of a coordinate, computed from the DEM and handed to `--domain`, so a catchment no longer has to be supplied as a file (Ola, 2026-09-27: "not far into the future"). The textbook route is depression handling (Priority-Flood, Barnes, Lehman and Mulla 2014), D8 flow directions (O'Callaghan and Mark 1984) and accumulation, the pour point snapped to the strongest flow nearby, the upstream cells traced and their outline turned into a polygon; the literature check is `@architect`'s. Open for its design: whether it runs in the C++ core (a 10 m tile is 25 M cells); how a stair-stepped cell outline becomes a domain polygon, which meets input coarsening; and that a real catchment crosses tile edges, so it needs gap 6 (a DEM in several tiles) first. Legacy has nothing on it (`grep -rliE "watershed|flow.?acc|flow.?dir|pour.?point|catchment" legacy` returns no files) | to design; after gap 6, which it needs; placed after 20c, can move ahead of it on Ola's word | none yet |
Expand Down
145 changes: 145 additions & 0 deletions docs/benchmarks/2026-09-27/21a-acceptance.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,145 @@
# Increment 21a acceptance run (@perf, 2026-09-27)

**Verdict: ACCEPTED.** On AC, against the previous merge (`227522c`) measured
back to back: refine at 8 threads is 11 % faster on the tile and 15 % faster on
the quarter circle. The mesh sha256 is identical on both domains. No measure
regressed.

## Method

- **Why a re-measured base.** The only stored `bench.py` run
(`serial-profile-bench/`) is on battery, so there is no AC baseline. As
`docs/increments/README.md` ("Acceptance") says to, master `227522c` (the
previous merge, #101) was checked out in a temporary `git worktree` and
measured with `--tree`, back to back with 21a's HEAD `cade752`. The worktree
was removed afterwards.
- **Two pairs.** Each tree was measured twice, in the order base, 21a, base,
21a, between 12:54 and 13:05 CEST. The second pair (`-r2`) is there to
measure how much the median of 5 moves between runs (see "Spread" below).
`bench.py`'s verdicts: `21a` against `21a-base-227522c` (`--baseline`),
`21a-r2` against `21a-base-227522c-r2` (`--baseline`), and
`21a-base-227522c-r2` against `21a-base-227522c` (found by the baseline
search). All three are ACCEPTED.
- **Command** (the main tree's `tools/bench.py`, blob `77765b1`, for all four
runs; each run did its own Release build into `<tree>/build-bench`):

```
.venv/bin/python tools/bench.py run --label <label> [--tree <worktree of 227522c>] \
[--baseline docs/benchmarks/2026-09-27/<base label>] --mesh-dir <scratch> \
--domain docs/benchmarks/2026-09-26/quarter.geojson --domain tile \
--dem tests/fixtures/dem_archive/7908_3_10m_z33.tif --tolerance 1 \
--repeats 5 --threads 1,2,...,20
```
- **Machine and power.** Apple M1 Max (8 P + 2 E), 32 GiB, macOS 27.0, on
**AC** for all four runs (96-98 %, charging; `pmset -g batt` before and after
each run is in its `run.json`). `caffeinate -ims` was held for the whole
sequence.
- **The `(dirty)` on `21a` and `21a-r2` is not a source change.** `bench.py`
sets `dirty` from `git status --porcelain`, which also lists untracked files.
During the 21a runs the only untracked files were the base runs' evidence
directories, written minutes earlier. `git status --porcelain
--untracked-files=no` was empty. The `_core` sha256 is the same in `21a` and
`21a-r2` (`654a2442…`), and so is the base's in both of its runs
(`c1d29205…`). Any back-to-back pair where the new tree is the repo will show
this until the first run's evidence is committed.
- **Scan time** is not in `bench.py`'s child record. It was measured afterwards
on AC by a scratch driver that uses the child's technique: it imports each
tree's `build-bench/pkg`, forces `threads` into `cli.refine`, and reads
`RefineOutcome.scan_seconds`. That was 5 repeats at 1, 8 and 16 threads, with
the two trees interleaved. The driver is not checked in. Its figures are the
"scan" rows below.
- **Meshes** stayed in the session scratchpad and are not kept. To regenerate
one, rerun `bench.py` as above. The quality run is one `--ascii` child per
domain at the CLI's default thread count.

## Results (refine seconds, median of 5; ms)

Pair 1 (`21a-base-227522c` then `21a`); pair 2 in brackets.

| domain | threads | base 227522c | 21a | change |
|---|---:|---:|---:|---:|
| tile | 1 | 520.8 [518.0] | 496.9 [500.8] | -4.6 % [-3.3 %] |
| tile | 8 | 249.5 [247.8] | 222.4 [222.1] | -10.9 % [-10.4 %] |
| tile | 16 | 245.8 [245.3] | 220.1 [220.1] | -10.5 % [-10.3 %] |
| tile | default (10) | 254.5 [254.5] | 219.4 [219.0] | -13.8 % [-13.9 %] |
| tile | best | 245.8 @16 [245.3 @16] | 218.4 @10 [218.4 @12] | -11.1 % [-11.0 %] |
| quarter | 1 | 464.0 [486.5] | 444.8 [445.0] | -4.1 % [-8.5 %] |
| quarter | 8 | 230.4 [233.9] | 196.5 [197.0] | -14.7 % [-15.8 %] |
| quarter | 16 | 224.4 [227.0] | 194.4 [195.4] | -13.3 % [-13.9 %] |
| quarter | default (10) | 233.5 [232.8] | 196.0 [194.3] | -16.1 % [-16.5 %] |
| quarter | best | 224.4 @16 [227.0 @16] | 194.2 @13 [194.0 @15] | -13.5 % [-14.5 %] |

Every swept thread count, 1 to 20, is faster in pair 1: by 4.6-14.7 % on the
tile and 4.1-16.3 % on the quarter circle.

**Scan at 8 threads** (scratch driver, median of 5, min-max in brackets):
tile 65.3 (64.8-68.1) -> 57.4 (57.0-59.2) ms; quarter 59.9 (59.3-60.5) ->
46.6 (46.5-47.4) ms. At 1 thread the scan is unchanged: tile 342.0 -> 342.8,
quarter 303.7 -> 303.1 ms.

**Ceiling** (speed-up over 1 thread; best, and at 20 threads):

| domain | base 227522c | 21a | 2026-09-26 (AC) |
|---|---|---|---|
| tile | 2.12x @16 (2.09x @20) [2.11x, 2.07x] | 2.28x @10 (2.26x @20) [2.29x, 2.28x] | 2.2x |
| quarter | 2.07x @16 (2.04x @20) [2.14x, 2.11x] | 2.29x @13 (2.27x @20) [2.29x, 2.28x] | 2.2x |

In 21a the curve is flat from 9 threads on: every median from 9 to 20 threads is within 1.3 % of the best. The base is
not monotone: on the tile it rises from 249.5 ms at 8 threads to 257.7 at 9,
then falls slowly to its best at 16. 21a's ceiling is above the 2026-09-26 AC
figure of 2.2x. The base, re-measured today with `bench.py`, is below it
(2.07-2.14x). The 2026-09-26 figure came from a different method (3 repeats,
timings parsed from stdout), so it is context here, not a baseline.

**Quality: identical.** Both trees, both pairs:

| domain | worst angle | max degree | within tolerance | Delaunay violations | mesh sha256 |
|---|---:|---:|---|---:|---|
| quarter | 0.3955 deg | 18 | yes | 0 of 641,791 | `1e531976f33b38dd883f9cf1945f6303dbac94c92f30b52b8bd7b46db7eab9ee` |
| tile | 0.6296 deg | 74 | yes | 0 of 692,056 | `11741a81adfa17b34a1ea56875ec9e791414d17625eb1c58305e7c0cf10e005c` |

The mesh sha256 of 21a equals the base's on both domains, as the design
requires for a bit-identical increment.

## Spread of the median of 5 (Ola's ruling on the 5 % threshold)

For each tree, the two runs' medians were compared over all 42 cells (2
domains × thread counts 0-20), as |median(run 1) / median(run 2) - 1|:

| tree | median abs. difference | 90th percentile | max |
|---|---:|---:|---:|
| base 227522c | 0.66 % | 1.65 % | 4.62 % (quarter, 1 thread) |
| 21a | 0.24 % | 0.72 % | 0.85 % |

Within a cell, (max - min) / median of the 5 samples had a median of 1.1-1.9 %
per run, and a maximum of 3.8-5.2 % for 21a and 8.5-13.9 % for the base.

The 4.62 % was not one outlier absorbed by the median. In `21a-base-227522c-r2`
all five quarter 1-thread samples were high (0.482-0.550 s, against
0.460-0.467 s in the first run; 4.8 % measured the other way round, against
run 1). The rest of that run's quarter cells were 0-3.4 % slower than run 1,
mostly 1-2 %. The cause has not been measured. So on AC the median of 5 usually moves
by under 2 % between back-to-back runs of the same build, but it has been seen
to move by 4.6 %, just under the 5 % threshold. A single-cell regression of
about 5 % is therefore not conclusive from one run: rerun that cell before
calling it. 21a's gains, 10-16 % at 8 threads and above, are several times that
spread, and they reproduced in both pairs.

## Against the estimates

| source | measure | estimate | measured (pair 1 [pair 2]) |
|---|---|---|---|
| design, "What the quick wins add up to" | 8-thread refine, all quick wins (21a + 21b) | 0.229 -> ~0.18 s | 21a alone: tile 0.249 -> 0.222 s, quarter 0.230 -> 0.197 s |
| design, QW1 + QW3 + stack | 21a alone at 8 threads | ~-10 % (QW1 -6 %, QW3 ~14 ms, stack 0.7 %) | tile -10.9 % [-10.4 %], quarter -14.7 % [-15.8 %] |
| design, QW1 | scan at 8 threads | ~62 -> 47 ms | tile 65.3 -> 57.4 ms, quarter 59.9 -> 46.6 ms |
| design, QW3 + stack | 1-thread refine | ~-17 ms (QW3 14 ms, stack ~3 ms) | tile -24 [-17] ms, quarter -19 [-42] ms |
| @developer quick sweep (vs red `7b54a4e`) | tile, 8 threads: scan / refine | 65.0 -> 57.7 ms / 243 -> 219 ms | 65.3 -> 57.4 ms / 249.5 -> 222.4 ms |

The design's 0.229 s baseline is the battery profile's figure, so it is
compared here by ratio, not by absolute value. 21a alone meets or beats its
share of the estimate. On the tile, the scan gain (-7.9 ms at 8 threads) is
about 20 ms less than the refine gain (-27 ms). That difference is thought to
be the QW3 merge and the caller-owned stack, as it matches the 1-thread gain,
where the scan did not change. No profile has confirmed that split. The
quarter's 1-thread pair-2 figure (-42 ms) is inflated by the base's high cell
described above.
Loading
Loading