Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
17 changes: 10 additions & 7 deletions .claude/agents/perf.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,13 +23,16 @@ it is never unmeasured again. You report measured figures only.
`docs/benchmarks/bench-py.md`. It is code under `tools/`, so a change to it
follows the TDD loop: a failing test in `tests/python/test_bench.py` first,
from `@tester`.
* **The serial-phase profile.** About half of single-thread refine time does not
parallelise, and the serial insert and flip phase is the suspected cause. It
has not been profiled. Profile it before anyone designs a fix for it.
* **The scaling ceiling.** Refine speeds up at most about 2.2× from 1 to 20
threads, flat from about 7, on AC as on battery
(`docs/benchmarks/2026-09-26/README.md`). Report each increment's ceiling
against it.
* **The serial-phase profile.** Profiled on 2026-09-27
(`docs/benchmarks/2026-09-27/serial-profile/README.md`): the serial part is
about a third of single-thread refine, mostly Lawson legalisation, and the
scan itself stops speeding up near 5× from load imbalance. Re-profile before
a design relies on those figures after refine changes.
* **The scaling ceiling.** Refine speeds up at most about 2.0-2.2× from 1 to
20 threads, flat from about 7-8. The 2026-09-26 sweep
(`docs/benchmarks/2026-09-26/scaling/`, medians per thread count) gives 2.2×
on AC and 2.1× on battery; 2026-09-27 gives 2.02-2.03× on battery (`docs/benchmarks/2026-09-27/serial-profile/README.md`). Report
each increment's ceiling against the matching power-state baseline.

## 2. How a run is made
* **Release build, rebuilt.** `bench.py run` builds Release into
Expand Down
2 changes: 1 addition & 1 deletion ROADMAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,7 +42,7 @@ increment that most needs a picture to check against
| 20 | Quality start: before DEM refinement, Steiner points at the DEM node nearest each bad triangle's circumcentre until the start mesh has a 25° minimum angle (`--start-min-angle`, 0 is off); geometry only, input segments not split, serial and deterministic. Removes increment 16's boundary fans | landed with 20b as interim (C1-C3 open, to 20c; C4 -> 20b) | `docs/increments/20-start-quality.md` |
| 20b | Minimum insertion distance from constraints: when refinement's worst node lies within ε = clamp(tol / slope, cell/100, cell/2) of a constraint segment, insert the foot on the segment (off-node, bilinear z) instead; a footed node still above tolerance is inserted after all, so the tolerance guarantee is unchanged. Removes the 0.0117° needle at 1 m (increment 20's C4) | shipped with 20 (#97) | `docs/increments/20b-min-insertion-distance.md` |
| — | `tools/bench.py`: the 1 m benchmark and the thread-scaling sweep from one checked-in command, with power state, quality and commit recorded per run; the one-off scripts in `docs/benchmarks/2026-09-26/` are its specification. Rule 2's acceptance run needs it | shipped with branch `tools-bench`'s PR | `docs/benchmarks/bench-py.md` |
| — | The serial phase: profile refine's serial insert-and-flip phase, then parallelise what the profile blames. Scaling tops out at about 2.2× from 1 to 20 threads | next (Ola, 2026-09-27: order bench.py, serial phase, 20c); `tools/bench.py`'s runs are its baseline | `docs/benchmarks/2026-09-26/README.md` |
| — | The serial phase: profile refine's serial insert-and-flip phase, then parallelise what the profile blames. Scaling tops out at about 2.0-2.2×. Profiled 2026-09-27: serial part about a third of 1-thread refine, mostly Lawson legalisation; the scan stops near 5× from load imbalance | profiled; designed as increment 21 (Ola's rulings 2026-09-27: L1 determinism, one path, at most 2 % more triangles): 21a and 21b quick wins next, then 21c measurements, then 21d | `docs/increments/21-parallel-refine.md`, `docs/benchmarks/2026-09-27/serial-profile/README.md` |
| 16b | Interior polygons and polylines as constraints ("terrain polygons": lakes, land cover, roads, rivers): `--features PATH`, a GeoJSON `FeatureCollection`, each feature naming a vocabulary property; closed or open `Breakline`s, crossings noded, off-node vertices with bilinear z. **Its working example is real data** (Ola, 2026-09-27): CORINE Land Cover 2018 over the benchmark tile `7908_3_10m_z33.tif`. The source is Ola's local copy, `rasputin_data/corine_sql/.../U2018_CLC2018_V2020_20u1.gpkg` (8.2 GB, EPSG:3035, a sibling of this repository, which also holds the 254-tile DTM10 archive for gap 6). It reads without GDAL: sqlite3 over its R-tree, the GeoPackage blob header stripped, `shapely.wkb`, then `pyproj` to 25833, so CRS stops in Python as before. Probed 2026-09-27 in 0.2 s: 60 polygons in 8 classes (heath, bare rock, sparse vegetation, bogs, intertidal flats, water, sea, urban), 11 068 vertices clipped to the tile, median segment 54 m against 10 m cells. The EEA's public ArcGIS service (`image.discomap.eea.europa.eu`, `Corine/CLC2018_WM`) returns the same 11 068 clipped vertices and is the route for anyone without the file. What it forces on 16b's design: neighbouring polygons share their boundaries, so each shared edge arrives twice; the polygons run past the domain and must be clipped; the extract is committed as a fixture with the Copernicus attribution. Placed before 20c because 20c may split constraint segments and should be designed and measured on inputs that have interior ones | designed in 16's R6, no increment file yet; after the serial phase, before 20c | `docs/increments/16-domain-polygon.md` (R6) |
| 20c | Soft quality criterion: a penalty that each Steiner node or constraint split must pay for in angle gained, instead of 20's hard 25°; applied at the start and during DEM refinement; may split constraint segments when that improves the mesh. Ola's rulings on 20's C1-C3 | to design after 16b (`@architect` measures cost against 20 first) | `docs/increments/20-start-quality.md` (Ola's rulings) |
| — | Auto-catchment: the watershed upstream of a coordinate, computed from the DEM and handed to `--domain`, so a catchment no longer has to be supplied as a file (Ola, 2026-09-27: "not far into the future"). The textbook route is depression handling (Priority-Flood, Barnes, Lehman and Mulla 2014), D8 flow directions (O'Callaghan and Mark 1984) and accumulation, the pour point snapped to the strongest flow nearby, the upstream cells traced and their outline turned into a polygon; the literature check is `@architect`'s. Open for its design: whether it runs in the C++ core (a 10 m tile is 25 M cells); how a stair-stepped cell outline becomes a domain polygon, which meets input coarsening; and that a real catchment crosses tile edges, so it needs gap 6 (a DEM in several tiles) first. Legacy has nothing on it (`grep -rliE "watershed|flow.?acc|flow.?dir|pour.?point|catchment" legacy` returns no files) | to design; after gap 6, which it needs; placed after 20c, can move ahead of it on Ola's word | none yet |
Expand Down
3 changes: 2 additions & 1 deletion docs/benchmarks/2026-09-26/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,4 +42,5 @@ This measures the `_core.refine` call alone, with threads forced by
Battery and AC agree to within about 7 % (the largest gap is the quarter circle at 7 threads:
0.258 s on AC against 0.241 s on battery), and they show the same ceiling. Roughly half of the single-thread
refine time does not parallelise. That part is thought to be the serial insert
and flip phase, but it has not been profiled yet.
and flip phase, but it has not been profiled yet. (Later profiled: about a
third, mostly Lawson legalisation; `docs/benchmarks/2026-09-27/serial-profile/README.md`.)
38 changes: 38 additions & 0 deletions docs/benchmarks/2026-09-27/serial-profile-bench/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,38 @@
# Benchmark run `serial-profile-bench`

Generated by `tools/bench.py`.

## Verdict

- NO BASELINE: no comparable stored run

Time threshold: 5.0 %.

## Method

- Started 2026-09-27T10:34:35.755087+02:00; tree `d6d6bebe4425df21ec758fc5377b4fc5f4b780f2` (dirty); bench.py blob `77765b181714c31714da476c2335c799c6a84021`. Build: Release, AppleClang 21.0.0.21000101, `-O3 -DNDEBUG`, _core sha256 `c1d29205fcc4b51a86b315b219de15456001e1837c16248dd380e0e655c4279a`.
- Apple M1 Max, 8 P + 2 E cores, 32 GiB, macOS 27.0, Python 3.14.7, numpy 2.5.3.
- Power **battery** (84%), `pmset -g batt` before and after (in run.json).
- DEM `/Users/skavhaug/projects/rasputin/tests/fixtures/dem_archive/7908_3_10m_z33.tif` (sha256 `aabd0cbc28471ce8593e4c278381811bec3fbf4e1194058b3b4888d00c3af575`), tolerance 1.0, extra mesh args ``; domains: `quarter` (/Users/skavhaug/projects/rasputin/docs/benchmarks/2026-09-26/quarter.geojson).
- One child per sample, `/Users/skavhaug/projects/rasputin/.venv/bin/python /Users/skavhaug/projects/rasputin/tools/bench.py _child --pkg /Users/skavhaug/projects/rasputin/build-bench/pkg --threads <N> -- mesh --dem /Users/skavhaug/projects/rasputin/tests/fixtures/dem_archive/7908_3_10m_z33.tif --tolerance 1 --out <path> --binary`, repeats interleaved over thread counts; t=0 is the CLI's default, other counts are forced into `refine`. Quality: one `--ascii` run per domain at t=0, kept out of the repository; rerun the child with `--ascii --out PATH` to regenerate it.

## `quarter`

| threads | n | median s | min s | max s |
|---:|---:|---:|---:|---:|
| 0 | 5 | 0.2398 | 0.2332 | 0.2423 |
| 1 | 5 | 0.4662 | 0.4636 | 0.4674 |
| 2 | 5 | 0.3354 | 0.3343 | 0.3386 |
| 4 | 5 | 0.2699 | 0.2693 | 0.2762 |
| 6 | 5 | 0.2416 | 0.2408 | 0.2518 |
| 8 | 5 | 0.2320 | 0.2305 | 0.2325 |
| 10 | 5 | 0.2415 | 0.2396 | 0.2494 |
| 12 | 5 | 0.2394 | 0.2382 | 0.2439 |
| 16 | 5 | 0.2306 | 0.2294 | 0.2411 |
| 20 | 5 | 0.2351 | 0.2271 | 0.2369 |

Ceiling: 1.98x at 20 threads over 1, best 2.02x at 16; 2026-09-26: about 2.2x, flat from about 7.

Quality: worst angle 0.3955 deg, median 45.00, share under 1 deg 0.00003, max degree 18, within tolerance True, Delaunay 0 violations of 641791 edges (80136 decided exactly), mesh sha256 `1e531976f33b38dd883f9cf1945f6303dbac94c92f30b52b8bd7b46db7eab9ee`.

<!-- bench.py: generated above this line; hand-written prose below is kept -->
50 changes: 50 additions & 0 deletions docs/benchmarks/2026-09-27/serial-profile-bench/raw.tsv
Original file line number Diff line number Diff line change
@@ -0,0 +1,50 @@
quarter 0 0 0.238149 battery
quarter 1 0 0.464229 battery
quarter 2 0 0.334424 battery
quarter 4 0 0.276165 battery
quarter 6 0 0.241329 battery
quarter 8 0 0.230470 battery
quarter 10 0 0.241540 battery
quarter 12 0 0.238213 battery
quarter 16 0 0.241120 battery
quarter 20 0 0.235054 battery
quarter 0 1 0.239837 battery
quarter 1 1 0.463626 battery
quarter 2 1 0.334267 battery
quarter 4 1 0.269896 battery
quarter 6 1 0.240825 battery
quarter 8 1 0.232145 battery
quarter 10 1 0.249413 battery
quarter 12 1 0.239442 battery
quarter 16 1 0.230285 battery
quarter 20 1 0.236204 battery
quarter 0 2 0.240615 battery
quarter 1 2 0.466478 battery
quarter 2 2 0.337431 battery
quarter 4 2 0.269493 battery
quarter 6 2 0.251845 battery
quarter 8 2 0.231956 battery
quarter 10 2 0.246728 battery
quarter 12 2 0.238936 battery
quarter 16 2 0.232863 battery
quarter 20 2 0.227131 battery
quarter 0 3 0.242317 battery
quarter 1 3 0.467371 battery
quarter 2 3 0.335420 battery
quarter 4 3 0.269319 battery
quarter 6 3 0.242405 battery
quarter 8 3 0.232472 battery
quarter 10 3 0.241090 battery
quarter 12 3 0.239928 battery
quarter 16 3 0.229365 battery
quarter 20 3 0.232811 battery
quarter 0 4 0.233192 battery
quarter 1 4 0.466176 battery
quarter 2 4 0.338617 battery
quarter 4 4 0.270899 battery
quarter 6 4 0.241555 battery
quarter 8 4 0.231848 battery
quarter 10 4 0.239570 battery
quarter 12 4 0.243907 battery
quarter 16 4 0.230588 battery
quarter 20 4 0.236879 battery
Loading
Loading