Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 11 additions & 0 deletions .claude/agents/perf.md
Original file line number Diff line number Diff line change
Expand Up @@ -49,6 +49,17 @@ it is never unmeasured again. You report measured figures only.
machine, the thread counts, the DEM, the domain and the tolerance.
* **Record quality with speed:** worst angle, max vertex degree, the tolerance
check and the Delaunay check. A faster run that loses quality is a regression.
* **Sanitize a scratch build before you run it for numbers.** A simulation or
instrumentation patch that changes C++ runs once sanitized on a small case
before any Release, timed or full-size run: `-fsanitize=address,undefined`
(as `build-san` builds) where it can run, which is a C++ driver or test
binary. ASan cannot be loaded into the Python process here (the harness
strips `DYLD_INSERT_LIBRARIES`), so a run through `_core` uses
`-fsanitize=undefined` with libc++'s extensive hardening mode instead. A run
that crashed produces no numbers. A deliberate crash (a plant) stays out of
a script's default loop: a crash in a Python process shows as a crash dialog
on Ola's screen, and five did on 2026-09-27 from one 21c plant
(`docs/benchmarks/2026-09-27/21c/data/a0eo/plant_eo_stalecommit/outcome.txt`).

## 3. Evidence
* Commit the evidence under `docs/benchmarks/<date>/`: a `README.md` with the
Expand Down
3 changes: 2 additions & 1 deletion ROADMAP.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,7 +42,8 @@ increment that most needs a picture to check against
| 20 | Quality start: before DEM refinement, Steiner points at the DEM node nearest each bad triangle's circumcentre until the start mesh has a 25° minimum angle (`--start-min-angle`, 0 is off); geometry only, input segments not split, serial and deterministic. Removes increment 16's boundary fans | landed with 20b as interim (C1-C3 open, to 20c; C4 -> 20b) | `docs/increments/20-start-quality.md` |
| 20b | Minimum insertion distance from constraints: when refinement's worst node lies within ε = clamp(tol / slope, cell/100, cell/2) of a constraint segment, insert the foot on the segment (off-node, bilinear z) instead; a footed node still above tolerance is inserted after all, so the tolerance guarantee is unchanged. Removes the 0.0117° needle at 1 m (increment 20's C4) | shipped with 20 (#97) | `docs/increments/20b-min-insertion-distance.md` |
| — | `tools/bench.py`: the 1 m benchmark and the thread-scaling sweep from one checked-in command, with power state, quality and commit recorded per run; the one-off scripts in `docs/benchmarks/2026-09-26/` are its specification. Rule 2's acceptance run needs it | shipped with branch `tools-bench`'s PR | `docs/benchmarks/bench-py.md` |
| — | The serial phase: profile refine's serial insert-and-flip phase, then parallelise what the profile blames. Scaling tops out at about 2.0-2.2×. Profiled 2026-09-27: serial part about a third of 1-thread refine, mostly Lawson legalisation; the scan stops near 5× from load imbalance | profiled; designed as increment 21 (Ola's rulings 2026-09-27: L1 determinism, one path, at most 2 % more triangles). **21a shipped with branch `increment21a-quick-wins`'s PR**: dynamic scan blocks, active merge, reused flip stack; mesh bit-identical; on AC refine -10 to -15 % at 8 threads, ceiling 2.1x -> 2.3x (`docs/benchmarks/2026-09-27/21a-acceptance.md`). **21b shipped with branch `increment21b-lattice-incircle`'s PR**: an int64 lattice incircle answers 99.97-100 % of refine's incircle tests; mesh bit-identical; on battery refine -15 % at 8 threads, ceiling 2.3x -> 2.5x (`docs/benchmarks/2026-09-27/21b-acceptance.md`). 21c measurements next, then 21d | `docs/increments/21-parallel-refine.md`, `docs/benchmarks/2026-09-27/serial-profile/README.md` |
| — | The serial phase: profile refine's serial insert-and-flip phase, then parallelise what the profile blames. Scaling tops out at about 2.0-2.2×. Profiled 2026-09-27: serial part about a third of 1-thread refine, mostly Lawson legalisation; the scan stops near 5× from load imbalance | profiled; designed as increment 21 (Ola's rulings 2026-09-27: L1 determinism, one path, at most 2 % more triangles). **21a shipped with branch `increment21a-quick-wins`'s PR**: dynamic scan blocks, active merge, reused flip stack; mesh bit-identical; on AC refine -10 to -15 % at 8 threads, ceiling 2.1x -> 2.3x (`docs/benchmarks/2026-09-27/21a-acceptance.md`). **21b shipped with branch `increment21b-lattice-incircle`'s PR**: an int64 lattice incircle answers 99.97-100 % of refine's incircle tests; mesh bit-identical; on battery refine -15 % at 8 threads, ceiling 2.3x -> 2.5x (`docs/benchmarks/2026-09-27/21b-acceptance.md`). **21c measured** (`docs/benchmarks/2026-09-27/21c/README.md`): option C costs 6.5-10 % more triangles; A1 about 1 %; A0 is bit-identical (on these two inputs, not proven), and with evaluate-once is modelled at only 7-11 % faster refine at 8 threads and slower at 4. **21d is deferred** (Ola, 2026-09-27: "Review and push 21c, then basin-work"): the multi-tile DEM and the computation CRS come first, then domain decomposition, which parallelises scan and split together | `docs/increments/21-parallel-refine.md`, `docs/benchmarks/2026-09-27/serial-profile/README.md` |
| — | Release hardening: measure libc++'s `_LIBCPP_HARDENING_MODE_FAST` (and libstdc++'s assertions on the GCC leg), which bounds-check `std::vector`, `span` and the like in Release, on the 1 m benchmark; switch it on if the cost is small. Today an out-of-range index in shipped code the tests miss is undefined behaviour and crashes Python (Ola, 2026-09-27: "I'm surprised we don't have proper memory control"; CI's ASan/UBSan/TSan cover what the tests reach) | to measure (Ola, 2026-09-27); 21d, which it was placed after, is deferred behind the basin work | none yet |
| 16b | Interior polygons and polylines as constraints ("terrain polygons": lakes, land cover, roads, rivers): `--features PATH`, a GeoJSON `FeatureCollection`, each feature naming a vocabulary property; closed or open `Breakline`s, crossings noded, off-node vertices with bilinear z. **Its working example is real data** (Ola, 2026-09-27): CORINE Land Cover 2018 over the benchmark tile `7908_3_10m_z33.tif`. The source is Ola's local copy, `rasputin_data/corine_sql/.../U2018_CLC2018_V2020_20u1.gpkg` (8.2 GB, EPSG:3035, a sibling of this repository, which also holds the 254-tile DTM10 archive for gap 6). It reads without GDAL: sqlite3 over its R-tree, the GeoPackage blob header stripped, `shapely.wkb`, then `pyproj` to 25833, so CRS stops in Python as before. Probed 2026-09-27 in 0.2 s: 60 polygons in 8 classes (heath, bare rock, sparse vegetation, bogs, intertidal flats, water, sea, urban), 11 068 vertices clipped to the tile, median segment 54 m against 10 m cells. The EEA's public ArcGIS service (`image.discomap.eea.europa.eu`, `Corine/CLC2018_WM`) returns the same 11 068 clipped vertices and is the route for anyone without the file. What it forces on 16b's design: neighbouring polygons share their boundaries, so each shared edge arrives twice; the polygons run past the domain and must be clipped; the extract is committed as a fixture with the Copernicus attribution. Placed before 20c because 20c may split constraint segments and should be designed and measured on inputs that have interior ones | designed in 16's R6, no increment file yet; after the serial phase, before 20c | `docs/increments/16-domain-polygon.md` (R6) |
| 20c | Soft quality criterion: a penalty that each Steiner node or constraint split must pay for in angle gained, instead of 20's hard 25°; applied at the start and during DEM refinement; may split constraint segments when that improves the mesh. Ola's rulings on 20's C1-C3 | to design after 16b (`@architect` measures cost against 20 first) | `docs/increments/20-start-quality.md` (Ola's rulings) |
| — | Auto-catchment: the watershed upstream of a coordinate, computed from the DEM and handed to `--domain`, so a catchment no longer has to be supplied as a file (Ola, 2026-09-27: "not far into the future"). The textbook route is depression handling (Priority-Flood, Barnes, Lehman and Mulla 2014), D8 flow directions (O'Callaghan and Mark 1984) and accumulation, the pour point snapped to the strongest flow nearby, the upstream cells traced and their outline turned into a polygon; the literature check is `@architect`'s. Open for its design: whether it runs in the C++ core (a 10 m tile is 25 M cells); how a stair-stepped cell outline becomes a domain polygon, which meets input coarsening; and that a real catchment crosses tile edges, so it needs gap 6 (a DEM in several tiles) first. Legacy has nothing on it (`grep -rliE "watershed|flow.?acc|flow.?dir|pour.?point|catchment" legacy` returns no files) | to design; after gap 6, which it needs; placed after 20c, can move ahead of it on Ola's word | none yet |
Expand Down
Loading
Loading