From 15241e8ed4831724040c7bef1e6542c90ddb7cd2 Mon Sep 17 00:00:00 2001 From: Ola Skavhaug Date: Sun, 27 Sep 2026 11:53:57 +0200 Subject: [PATCH 1/9] 21: 320 x 1.99 rounds to ~637 Co-Authored-By: Claude Opus 5.5 --- docs/increments/21-parallel-refine.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/increments/21-parallel-refine.md b/docs/increments/21-parallel-refine.md index bb61ac08..1facbcb2 100644 --- a/docs/increments/21-parallel-refine.md +++ b/docs/increments/21-parallel-refine.md @@ -769,7 +769,7 @@ overrun recorded so far is +99 % for a whole increment (6a shipped 467 lines against ~235, `06-cdt-viewer.md:626`) and +116 % for one file (`scene.py`, `05b-noder-driver.md:379-382`); the table gives each at +66 % (increment 17). At +99 % the conclusions hold: A1 (~450) comes to ~895 and is split in any -case; C (~320) comes to ~636, under 700. The per-file factor does not apply to +case; C (~320) comes to ~637, under 700. The per-file factor does not apply to a whole increment, but at +116 % C would be ~691, only just under, so 21d on C is worth counting early. From 09b005bb81c539ed8bf40ef2f0d8db25144ee7b9 Mon Sep 17 00:00:00 2001 From: Ola Skavhaug Date: Sun, 27 Sep 2026 11:54:05 +0200 Subject: [PATCH 2/9] perf.md: wrap the ceiling bullet Co-Authored-By: Claude Opus 5.5 --- .claude/agents/perf.md | 5 +++-- 1 file changed, 3 insertions(+), 2 deletions(-) diff --git a/.claude/agents/perf.md b/.claude/agents/perf.md index 143bacad..19ef44e4 100644 --- a/.claude/agents/perf.md +++ b/.claude/agents/perf.md @@ -31,8 +31,9 @@ it is never unmeasured again. You report measured figures only. * **The scaling ceiling.** Refine speeds up at most about 2.0-2.2× from 1 to 20 threads, flat from about 7-8. The 2026-09-26 sweep (`docs/benchmarks/2026-09-26/scaling/`, medians per thread count) gives 2.2× - on AC and 2.1× on battery; 2026-09-27 gives 2.02-2.03× on battery (`docs/benchmarks/2026-09-27/serial-profile/README.md`). Report - each increment's ceiling against the matching power-state baseline. + on AC and 2.1× on battery; 2026-09-27 gives 2.02-2.03× on battery + (`docs/benchmarks/2026-09-27/serial-profile/README.md`). Report each + increment's ceiling against the matching power-state baseline. ## 2. How a run is made * **Release build, rebuilt.** `bench.py run` builds Release into From 7b54a4e9d417fe979dd84f5b5cabe3d8cd2bf1f9 Mon Sep 17 00:00:00 2001 From: Ola Skavhaug Date: Sun, 27 Sep 2026 12:04:41 +0200 Subject: [PATCH 3/9] red: increment 21a (dynamic scan scheduling, active merge, legalise stack) 21a is bit-identical, so 14's T6 and 18's golden digests stay as they are and remain the output oracle. Three new suites pin the interfaces section 3 left open; the pins are written into 21-parallel-refine.md under "Pinned by the red suite (21a)". test_refinement_chunks_dynamic (invariant-critical, QW1): for_each_block, BlockSchedule and default_block. Every index is visited exactly once, and the calls are exactly the block partition, for n in {0, 1, primes, 16, 17, 128, 1009}, threads {1, 2, 3, 7, 8, 0}, seven block sizes and the inline threshold below, at and above n. One thread or n < inline_below runs on the caller in ascending order. At or above the threshold two blocks rendezvous, which proves they run concurrently. On the exception contract, every block runs and the lowest throwing block index is rethrown, with the lowest thrower delayed so the order of arrival differs from the order of index. Mutation round against a scratch implementation that is not committed. All nine mutants were killed: a dropped last partial block, a doubled block 0, the highest-index exception, the first-in-time exception, always inline, never inline, <= at the threshold, stop after a throw, and a default_block with /8 (a compile-time failure). The scratch implementation passed 30 randomised-order runs and TSan. test_refinement_active (QW3): detail::rebuild_active equals today's collect + sort + unique, copied from 09b005b, on empty, duplicate, interleaved, all-touched and 40 random rounds, including a reused buffer. test_mesh_lawson_stack: legalise_around(..., FlipStack&, on_write) gives the same flip count, on_write sequence and mesh as today's algorithm, copied as the oracle. The comparison runs after each of more than 100 random insertions (inside and edge splits, ring constraints) in three frames. All three are registered only once their header names for_each_block, rebuild_active or FlipStack (the 18 and 20b precedent), so the main build stays green. The headers are configure dependencies. All three are in the TSan job, which stays red until green lands. Co-Authored-By: Claude Opus 5.5 --- .github/workflows/main.yaml | 5 +- docs/increments/21-parallel-refine.md | 38 +++ tests/cpp/CMakeLists.txt | 37 +++ tests/cpp/unit/test_mesh_lawson_stack.cpp | 233 +++++++++++++++++ tests/cpp/unit/test_refinement_active.cpp | 126 ++++++++++ .../unit/test_refinement_chunks_dynamic.cpp | 238 ++++++++++++++++++ 6 files changed, 676 insertions(+), 1 deletion(-) create mode 100644 tests/cpp/unit/test_mesh_lawson_stack.cpp create mode 100644 tests/cpp/unit/test_refinement_active.cpp create mode 100644 tests/cpp/unit/test_refinement_chunks_dynamic.cpp diff --git a/.github/workflows/main.yaml b/.github/workflows/main.yaml index b2fb507d..f105b46e 100644 --- a/.github/workflows/main.yaml +++ b/.github/workflows/main.yaml @@ -102,6 +102,7 @@ jobs: test_mesh_row_spans prop_refinement_scan_equivalence test_mesh_quality prop_refinement_quality test_refinement_constraint_feet + test_refinement_chunks_dynamic test_refinement_active test_mesh_lawson_stack - name: Test env: @@ -112,7 +113,9 @@ jobs: test_refinement_scan_offnode test_refinement_refine \ test_mesh_row_spans prop_refinement_scan_equivalence \ test_mesh_quality prop_refinement_quality \ - test_refinement_constraint_feet; do + test_refinement_constraint_feet \ + test_refinement_chunks_dynamic test_refinement_active \ + test_mesh_lawson_stack; do build-tsan/tests/cpp/$suite done diff --git a/docs/increments/21-parallel-refine.md b/docs/increments/21-parallel-refine.md index 1facbcb2..664041bd 100644 --- a/docs/increments/21-parallel-refine.md +++ b/docs/increments/21-parallel-refine.md @@ -425,6 +425,44 @@ section 5 is measured against the mesh they leave. A smaller item folds into 21a: `legalise_around` allocates its stack per call (0.7 % of refine); a caller-owned buffer removes it. +### Pinned by the red suite (21a) + +Section 3 left these names and signatures open; `@tester` chose them in the +red step, and the suites named below hold them. + +- **QW1** (`include/terrain/parallel_util/chunks.hpp`, + `tests/cpp/unit/test_refinement_chunks_dynamic.cpp`): + `struct BlockSchedule { std::size_t block = 0; std::size_t inline_below; }`, + `constexpr std::size_t default_block(std::size_t n, unsigned threads) noexcept` + returning `max(1, n / (16 * threads))`, and + `template void for_each_block(std::size_t n, unsigned threads, BlockSchedule, Fn&& fn)`. + `block == 0` means `default_block(n, threads)` after `threads == 0` is + resolved as `for_each_chunk` resolves it. Block `k` is + `[k*b, min(n, (k+1)*b))` and `fn` is called exactly once per block, so the + set of calls is the same partition for every thread count. With one thread + or `n < inline_below`, every block runs on the calling thread in ascending + order. Otherwise at most `min(threads, blocks)` threads call `fn`, and blocks + run concurrently. Every block runs even when some throw; the exception of the + lowest block index that threw is rethrown after the join. The default of + `inline_below` is 21a's to set from the sweep and is not pinned. + `BlockSchedule{}` is the scan's schedule. `for_each_chunk` and its suite are + unchanged. +- **QW3** (`include/terrain/refinement/refine.hpp`, + `tests/cpp/unit/test_refinement_active.cpp`): + `void detail::rebuild_active(std::span touched, std::span skipped, std::vector& active)`. + It replaces `active`'s contents with what today's collect, sort and unique + give. Precondition: `skipped` is ascending and below `touched.size()`. +- **The stack** (`include/terrain/mesh/lawson.hpp`, + `tests/cpp/unit/test_mesh_lawson_stack.cpp`): + `using FlipStack = std::vector;` and a `legalise_around` + overload taking `FlipStack& stack` before `on_write`. It must produce the + same flips, `on_write` sequence and mesh as today's algorithm, which the + suite copies as its oracle. Allocation counts are not pinned. + +The three suites are registered in `tests/cpp/CMakeLists.txt` only once the +header names `for_each_block`, `rebuild_active` or `FlipStack` (the 18 and +20b precedent). The review drops those guards. + ## 4. Determinism levels Today's contract (14 R5, 14b R1, tested by 14's T6 and 18's T3 golden diff --git a/tests/cpp/CMakeLists.txt b/tests/cpp/CMakeLists.txt index ad30b640..efd973cf 100644 --- a/tests/cpp/CMakeLists.txt +++ b/tests/cpp/CMakeLists.txt @@ -276,3 +276,40 @@ target_link_libraries(prop_refinement_quality PRIVATE Threads::Threads) add_terrain_backend_test(test_refinement_constraint_feet property/prop_refinement_constraint_feet.cpp) target_link_libraries(test_refinement_constraint_feet PRIVATE Threads::Threads) + +# Parallel refine quick wins, increment 21a (docs/increments/21-parallel-refine.md, +# section 3 and "Pinned by the red suite (21a)"). Bit-identical: 14's T6 and +# 18's golden digests are the output oracle and stay as they are. +# test_refinement_chunks_dynamic (QW1, the dynamic block scheduler) is the +# invariant-critical suite; test_refinement_active (QW3) and +# test_mesh_lawson_stack (legalise_around's caller-owned stack) compare with +# today's code. All three run in the TSan job. +# +# RED UNTIL THE INTERFACES EXIST, keyed on the name each suite needs, because +# 21a adds no header: each suite is registered only once its header names it, +# so the red does not break the rest of the build. The green commit makes them +# build; the review drops these guards. +# The three headers are configure dependencies, so the green commit's edit +# re-runs the configure step and registers the suites without a manual cmake. +set_property(DIRECTORY APPEND PROPERTY CMAKE_CONFIGURE_DEPENDS + "${PROJECT_SOURCE_DIR}/include/terrain/parallel_util/chunks.hpp" + "${PROJECT_SOURCE_DIR}/include/terrain/refinement/refine.hpp" + "${PROJECT_SOURCE_DIR}/include/terrain/mesh/lawson.hpp") +file(READ "${PROJECT_SOURCE_DIR}/include/terrain/parallel_util/chunks.hpp" _chunks_hpp) +string(FIND "${_chunks_hpp}" "for_each_block" _block_at) +if(NOT _block_at EQUAL -1) + add_terrain_test(test_refinement_chunks_dynamic unit/test_refinement_chunks_dynamic.cpp) + target_link_libraries(test_refinement_chunks_dynamic PRIVATE Threads::Threads) +endif() +file(READ "${PROJECT_SOURCE_DIR}/include/terrain/refinement/refine.hpp" _refine_21a) +string(FIND "${_refine_21a}" "rebuild_active" _active_at) +if(NOT _active_at EQUAL -1) + add_terrain_backend_test(test_refinement_active unit/test_refinement_active.cpp) + target_link_libraries(test_refinement_active PRIVATE Threads::Threads) +endif() +file(READ "${PROJECT_SOURCE_DIR}/include/terrain/mesh/lawson.hpp" _lawson_hpp) +string(FIND "${_lawson_hpp}" "FlipStack" _stack_at) +if(NOT _stack_at EQUAL -1) + add_terrain_backend_test(test_mesh_lawson_stack unit/test_mesh_lawson_stack.cpp) + target_link_libraries(test_mesh_lawson_stack PRIVATE Threads::Threads) +endif() diff --git a/tests/cpp/unit/test_mesh_lawson_stack.cpp b/tests/cpp/unit/test_mesh_lawson_stack.cpp new file mode 100644 index 00000000..1c6f50ea --- /dev/null +++ b/tests/cpp/unit/test_mesh_lawson_stack.cpp @@ -0,0 +1,233 @@ +// Increment 21a (docs/increments/21-parallel-refine.md, end of section 3, and +// "Pinned by the red suite (21a)"): legalise_around with a caller-owned stack, +// so refine's serial phase stops allocating one per insertion. 21a is +// bit-identical, so the contract is only what is observable: on the same mesh +// and seeds, the same flips, the same on_write sequence and the same mesh as +// today's legalise_around. Allocation counts are not pinned. +// +// Interface pinned here (include/terrain/mesh/lawson.hpp): +// +// using FlipStack = std::vector; +// template +// std::size_t legalise_around(LatticeMesh&, std::uint32_t q, +// std::span seeds, +// const LatticeFrame&, FlipStack& stack, +// OnWrite&& on_write); +// +// The oracle is today's algorithm (lawson.hpp at 09b005b), copied below, not +// the five-argument overload, which 21a may turn into a wrapper of the new one. + +#include +#include +#include + +#include +#include +#include +#include + +#include "refinement_fixtures.hpp" + +#include +#include +#include +#include +#include +#include +#include + +using terrain::TriangleIndices; +using terrain::mesh::FlipStack; +using terrain::mesh::kNoNeighbour; +using terrain::mesh::LatticeFrame; +using terrain::mesh::LatticeMesh; +using terrain::mesh::LatticeVertex; +using terrain::mesh::legalise_around; +using terrain::pred::DefaultKernel; +using refinement_fixtures::orient; +using refinement_fixtures::RC; + +namespace { + +// legalise_around as of 09b005b, with its own vector. +std::size_t reference_around(LatticeMesh& m, std::uint32_t q, std::span seeds, + const LatticeFrame& f, std::vector& writes) { + std::vector stack(seeds.begin(), seeds.end()); + std::size_t flips = 0; + while (!stack.empty()) { + const auto t = stack.back(); + stack.pop_back(); + const auto& tri = m.triangles()[t]; + unsigned i = 0; + while (i < 3 && tri[i] != q) ++i; + if (i == 3 || !terrain::mesh::detail::must_flip(m, t, (i + 1) % 3, f)) continue; + const auto u = m.neighbours(t)[(i + 1) % 3]; + m.flip(t, (i + 1) % 3); + ++flips; + writes.push_back(t); + writes.push_back(u); + stack.push_back(t); + stack.push_back(u); + } + return flips; +} + +// A (n+1) x (n+1) grid of nodes `step` apart, two triangles per cell, the +// ring constrained. Diagonal direction alternates by cell when `mixed`, so +// the start already has Delaunay ties both ways. +LatticeMesh grid(std::uint32_t n, std::uint32_t step, bool mixed) { + std::vector v; + for (std::uint32_t r = 0; r <= n; ++r) + for (std::uint32_t c = 0; c <= n; ++c) v.push_back(LatticeVertex{r * step, c * step}); + const auto at = [n](std::uint32_t r, std::uint32_t c) { return r * (n + 1) + c; }; + std::vector t; + std::vector bits; + std::vector> masks; + for (std::uint32_t r = 0; r < n; ++r) + for (std::uint32_t c = 0; c < n; ++c) { + const auto tl = at(r, c), tr = at(r, c + 1), bl = at(r + 1, c), br = at(r + 1, c + 1); + if (!mixed || (r + c) % 2 == 0) { + t.push_back({tl, bl, br}); + t.push_back({tl, br, tr}); + } else { + t.push_back({tl, bl, tr}); + t.push_back({bl, br, tr}); + } + } + // Constrain every edge that lies on the ring (one triangle only has it). + for (const auto& tri : t) { + std::uint8_t b = 0; + for (unsigned k = 0; k < 3; ++k) { + const auto p = v[tri[k]], q = v[tri[(k + 1) % 3]]; + const bool on_ring = (p.row == q.row && (p.row == 0 || p.row == n * step)) || + (p.col == q.col && (p.col == 0 || p.col == n * step)); + if (on_ring) b = static_cast(b | (1u << k)); + } + bits.push_back(b); + masks.push_back({(b & 1u) != 0 ? 1u : 0u, (b & 2u) != 0 ? 1u : 0u, (b & 4u) != 0 ? 1u : 0u}); + } + auto m = LatticeMesh::build(std::move(v), std::move(t), std::move(bits), std::move(masks)); + REQUIRE(m.has_value()); + return *m; +} + +RC rc(const LatticeMesh& m, std::uint32_t i) { + const auto v = m.vertices()[i].as_node().value(); + return RC{v.row, v.col}; +} + +// The triangle holding node p, and the edge p lies on if any; nullopt when p +// is already a vertex. Integer orientation, independent of the mesh's own. +struct Hit { + std::uint32_t t; + std::optional edge; +}; +std::optional locate(const LatticeMesh& m, LatticeVertex p) { + const RC x{p.row, p.col}; + for (std::uint32_t t = 0; t < m.triangle_count(); ++t) { + const auto& tri = m.triangles()[t]; + std::array o{}; + for (unsigned k = 0; k < 3; ++k) o[k] = orient(rc(m, tri[k]), rc(m, tri[(k + 1) % 3]), x); + if (o[0] < 0 || o[1] < 0 || o[2] < 0) continue; + const int zeros = (o[0] == 0) + (o[1] == 0) + (o[2] == 0); + if (zeros >= 2) return std::nullopt; + if (zeros == 0) return Hit{t, std::nullopt}; + return Hit{t, static_cast(o[0] == 0 ? 0 : o[1] == 0 ? 1 : 2)}; + } + return std::nullopt; +} + +void require_same_mesh(const LatticeMesh& a, const LatticeMesh& b) { + REQUIRE(a.triangle_count() == b.triangle_count()); + REQUIRE(a.vertices().size() == b.vertices().size()); + for (std::uint32_t t = 0; t < a.triangle_count(); ++t) { + CAPTURE(t); + REQUIRE(a.triangles()[t] == b.triangles()[t]); + REQUIRE(a.neighbours(t) == b.neighbours(t)); + for (unsigned k = 0; k < 3; ++k) { + REQUIRE(a.is_constrained(t, k) == b.is_constrained(t, k)); + REQUIRE(a.mask(t, k) == b.mask(t, k)); + } + } +} + +const std::array kFrames{LatticeFrame{1.0, 1.0}, LatticeFrame{10.0, 5.0}, + LatticeFrame{0.1, 0.3}}; + +} // namespace + +TEST_CASE("legalise_around with a caller-owned stack matches today's, insertion by insertion", + "[lawson][around][stack]") { + // Random nodes inserted as refine inserts them: split_inside for an + // interior node, split_edge for one on an edge (the ring's constrained + // edges included), with refine's seeds. One FlipStack serves every + // insertion, as refine will reuse it. After each insertion the new + // overload must agree with the reference on a copy: flip count, the + // on_write sequence in order, and the whole mesh. + const std::uint32_t seed = GENERATE(range(1u, 9u)); + const std::size_t frame = GENERATE(std::size_t{0}, 1, 2); + const bool mixed = GENERATE(false, true); + CAPTURE(seed, frame, mixed); + const LatticeFrame f = kFrames[frame]; + auto m = grid(4, 6, mixed); // nodes 0..24 both ways + std::mt19937 rng{seed}; + FlipStack stack; + std::size_t inserted = 0, flipped = 0, on_edges = 0; + for (int attempt = 0; attempt < 300; ++attempt) { + const LatticeVertex p{static_cast(rng() % 25), static_cast(rng() % 25)}; + const auto hit = locate(m, p); + if (!hit) continue; + const auto t = hit->t; + const auto before = static_cast(m.triangle_count()); + std::array seeds{t, before, before + 1, before + 1}; + std::size_t n_seeds = 3; + std::uint32_t q = 0; + if (!hit->edge) { + q = m.split_inside(t, p); + } else { + const std::uint32_t u = m.neighbours(t)[*hit->edge]; + q = m.split_edge(t, *hit->edge, p); + n_seeds = u != kNoNeighbour ? 4 : 2; + if (n_seeds == 4) seeds[3] = u; + ++on_edges; + } + const std::span s{seeds.data(), n_seeds}; + LatticeMesh expected = m; + std::vector expected_writes; + const std::size_t expected_flips = reference_around(expected, q, s, f, expected_writes); + + std::vector writes; + const std::size_t flips = legalise_around( + m, q, s, f, stack, [&writes](std::uint32_t w) { writes.push_back(w); }); + CAPTURE(attempt, p.row, p.col); + REQUIRE(flips == expected_flips); + REQUIRE(writes == expected_writes); + require_same_mesh(m, expected); + ++inserted; + flipped += flips; + } + // The run exercised what it claims to: many insertions, flips, and edge + // splits (a run with none of these would compare nothing). + REQUIRE(inserted >= 100); + REQUIRE(flipped >= 20); + REQUIRE(on_edges >= 5); +} + +TEST_CASE("legalise_around with a caller-owned stack and no flip to make", "[lawson][around][stack]") { + // Seeds that do not contain q and a q whose fan is already Delaunay: no + // flip, no write, mesh unchanged, as today. + auto m = grid(2, 4, false); + const auto q = m.split_inside(0, LatticeVertex{3, 1}); + LatticeMesh expected = m; + std::vector expected_writes; + const std::array seeds{1, 2, 3, 0, 1}; + const std::size_t expected_flips = + reference_around(expected, q, std::span{seeds}, kFrames[0], expected_writes); + FlipStack stack; + std::vector writes; + REQUIRE(legalise_around(m, q, std::span{seeds}, kFrames[0], stack, + [&writes](std::uint32_t w) { writes.push_back(w); }) == + expected_flips); + REQUIRE(writes == expected_writes); + require_same_mesh(m, expected); +} diff --git a/tests/cpp/unit/test_refinement_active.cpp b/tests/cpp/unit/test_refinement_active.cpp new file mode 100644 index 00000000..f3149069 --- /dev/null +++ b/tests/cpp/unit/test_refinement_active.cpp @@ -0,0 +1,126 @@ +// Increment 21a, QW3 (docs/increments/21-parallel-refine.md, section 3, and +// "Pinned by the red suite (21a)"): the next round's `active` is rebuilt by a +// merge instead of sort + unique. 21a is bit-identical, so the only contract is +// that the vector is the one today's code builds. End to end that is held by +// 14's T6 and 18's golden digests, untouched; this suite holds it on +// constructed inputs, where the edge cases can be named. +// +// Interface pinned here (include/terrain/refinement/refine.hpp): +// +// namespace terrain::refinement::detail { +// void rebuild_active(std::span touched, +// std::span skipped, +// std::vector& active); +// } +// +// active's previous contents are replaced. Precondition: skipped ascending +// (refine fills it while walking the ascending `active`) and every entry below +// touched.size(). A skipped slot may also be touched; it appears once. + +#include +#include +#include + +#include + +#include +#include +#include +#include +#include + +using terrain::refinement::detail::rebuild_active; + +namespace { + +// Today's code, refine.hpp at 09b005b, lines 360-366, verbatim but for names. +std::vector reference(const std::vector& touched, + const std::vector& skipped) { + std::vector active; + for (std::uint32_t t = 0; t < touched.size(); ++t) + if (touched[t] != 0) + active.push_back(t); + active.insert(active.end(), skipped.begin(), skipped.end()); + std::sort(active.begin(), active.end()); + active.erase(std::unique(active.begin(), active.end()), active.end()); + return active; +} + +std::vector rebuilt(const std::vector& touched, + const std::vector& skipped, + std::vector active = {}) { + rebuild_active(std::span{touched}, std::span{skipped}, active); + return active; +} + +std::vector touched_at(std::size_t size, std::initializer_list at) { + std::vector touched(size, 0); + for (const auto t : at) touched[t] = 1; + return touched; +} + +} // namespace + +TEST_CASE("rebuild_active: both empty gives an empty active", "[refinement][active]") { + REQUIRE(rebuilt({}, {}).empty()); + REQUIRE(rebuilt(std::vector(7, 0), {}).empty()); +} + +TEST_CASE("rebuild_active replaces the previous contents", "[refinement][active]") { + const std::vector stale{0, 1, 2, 3, 4, 5, 6, 99}; + REQUIRE(rebuilt({}, {}, stale).empty()); + const auto touched = touched_at(6, {4}); + REQUIRE(rebuilt(touched, {1}, stale) == std::vector{1, 4}); +} + +TEST_CASE("rebuild_active: only touched, only skipped", "[refinement][active]") { + REQUIRE(rebuilt(touched_at(5, {0, 2, 4}), {}) == std::vector{0, 2, 4}); + REQUIRE(rebuilt(std::vector(9, 0), {1, 3, 8}) == std::vector{1, 3, 8}); +} + +TEST_CASE("rebuild_active: a skipped slot touched later appears once", "[refinement][active]") { + // Duplicates between the two inputs, at the front, middle and back. + const auto touched = touched_at(10, {0, 4, 5, 9}); + const std::vector skipped{0, 3, 5, 9}; + REQUIRE(rebuilt(touched, skipped) == std::vector{0, 3, 4, 5, 9}); + REQUIRE(rebuilt(touched, skipped) == reference(touched, skipped)); +} + +TEST_CASE("rebuild_active: interleaved inputs", "[refinement][active]") { + const auto touched = touched_at(12, {1, 3, 5, 7, 11}); + const std::vector skipped{0, 2, 4, 6, 8, 10}; + REQUIRE(rebuilt(touched, skipped) == std::vector{0, 1, 2, 3, 4, 5, 6, 7, 8, 10, 11}); +} + +TEST_CASE("rebuild_active: every slot touched and every slot skipped", "[refinement][active]") { + std::vector all(33); + for (std::uint32_t i = 0; i < all.size(); ++i) all[i] = i; + REQUIRE(rebuilt(std::vector(33, 1), all) == all); +} + +TEST_CASE("rebuild_active: touched holds values other than 1", "[refinement][active]") { + // refine writes 1, but resize(.., 1) and a char's truth are the contract: + // any nonzero entry is touched. + std::vector touched{0, 2, 0, -1, 1}; + REQUIRE(rebuilt(touched, {2}) == std::vector{1, 2, 3, 4}); +} + +TEST_CASE("rebuild_active equals today's sort + unique on random rounds", "[refinement][active]") { + // Shaped like a round: `size` slots, a touched density, and `skipped` an + // ascending subset drawn independently, so it overlaps touched at random. + const std::uint32_t seed = GENERATE(range(1u, 41u)); + std::mt19937 rng{seed}; + const std::size_t size = rng() % 2000; + const std::uint32_t touched_per_mille = rng() % 1001, skipped_per_mille = rng() % 1001; + std::vector touched(size, 0); + std::vector skipped; + for (std::uint32_t t = 0; t < size; ++t) { + touched[t] = rng() % 1000 < touched_per_mille ? 1 : 0; + if (rng() % 1000 < skipped_per_mille) skipped.push_back(t); + } + CAPTURE(seed, size, touched_per_mille, skipped_per_mille); + REQUIRE(rebuilt(touched, skipped) == reference(touched, skipped)); + // And into a reused, non-empty buffer, as refine reuses `active`. + REQUIRE(rebuilt(touched, skipped, std::vector(size / 2 + 3, 7)) == + reference(touched, skipped)); +} diff --git a/tests/cpp/unit/test_refinement_chunks_dynamic.cpp b/tests/cpp/unit/test_refinement_chunks_dynamic.cpp new file mode 100644 index 00000000..bacb46de --- /dev/null +++ b/tests/cpp/unit/test_refinement_chunks_dynamic.cpp @@ -0,0 +1,238 @@ +// Increment 21a, QW1 (docs/increments/21-parallel-refine.md, section 3, and +// "Pinned by the red suite (21a)"): the dynamic block scheduler that replaces +// for_each_chunk for the scan. test_refinement_chunks extended to it. +// +// INVARIANT-CRITICAL (section 7): a dropped block leaves a stale scan result +// and a doubled one is a second writer to a result slot, and either breaks the +// tolerance guarantee silently. So every index is visited exactly once for +// every n, thread count, block size and side of the inline threshold, and the +// exception contract is decided here. Mutation-tested in the red step. +// +// Interface pinned here (include/terrain/parallel_util/chunks.hpp): +// +// struct BlockSchedule { +// std::size_t block = 0; // items per block; 0: default_block(n, threads) +// std::size_t inline_below = ?; // n < inline_below runs on the calling thread; +// // the default is 21a's, from the sweep +// }; +// constexpr std::size_t default_block(std::size_t n, unsigned threads) noexcept; +// // max(1, n / (16 * threads)), threads >= 1 +// template +// void for_each_block(std::size_t n, unsigned threads, BlockSchedule, Fn&& fn); +// +// threads == 0 means hardware_concurrency, or 1 if that is 0, as for +// for_each_chunk. With b the block size, block k is [k*b, min(n, (k+1)*b)), +// and fn(begin, end) is called exactly once per block: the set of calls is the +// partition, whatever the threads or the timing. With one thread or +// n < inline_below every call is on the calling thread, in ascending order. +// Otherwise at most min(threads, blocks) threads call fn, and two blocks can +// run at once. Every block runs even when some throw; after the join the +// exception of the lowest block index that threw is rethrown. + +#include +#include + +#include + +#include +#include +#include +#include +#include +#include +#include +#include +#include +#include + +using terrain::parallel_util::BlockSchedule; +using terrain::parallel_util::default_block; +using terrain::parallel_util::for_each_block; + +namespace { + +using Range = std::pair; + +unsigned resolved(unsigned threads) { + return threads == 0 ? std::max(1u, std::thread::hardware_concurrency()) : threads; +} + +// The partition the contract promises, built independently of the scheduler. +std::vector partition(std::size_t n, std::size_t block) { + std::vector blocks; + for (std::size_t begin = 0; begin < n; begin += block) + blocks.emplace_back(begin, std::min(n, begin + block)); + return blocks; +} + +// Every call, in the order calls were recorded, with the thread that made it. +struct Trace { + std::vector> hits; + std::vector calls; + std::vector by; + std::mutex m; + explicit Trace(std::size_t n) : hits(n) {} + + auto sink() { + return [this](std::size_t begin, std::size_t end) { + for (std::size_t i = begin; i < end; ++i) hits[i].fetch_add(1); + const std::lock_guard lock{m}; + calls.emplace_back(begin, end); + by.push_back(std::this_thread::get_id()); + }; + } +}; + +struct BlockError { + std::size_t block; +}; + +} // namespace + +TEST_CASE("default_block is n / (16 * threads), at least 1", "[refinement][chunks][dynamic]") { + STATIC_REQUIRE(default_block(0, 1) == 1); + STATIC_REQUIRE(default_block(1, 1) == 1); + STATIC_REQUIRE(default_block(15, 1) == 1); + STATIC_REQUIRE(default_block(16, 1) == 1); + STATIC_REQUIRE(default_block(47, 1) == 2); + STATIC_REQUIRE(default_block(1000, 1) == 62); + STATIC_REQUIRE(default_block(1280, 8) == 10); + STATIC_REQUIRE(default_block(1279, 8) == 9); + STATIC_REQUIRE(default_block(10, 8) == 1); + STATIC_REQUIRE(default_block(100000, 20) == 312); +} + +TEST_CASE("for_each_block calls fn once per block of the partition, every index exactly once", + "[refinement][chunks][dynamic]") { + // n covers 0, 1, primes, powers of two and their neighbours; block covers + // the default (0), 1, blocks that do and do not divide n, and one larger + // than every n. The threshold is put below, at and above n. + const std::size_t n = GENERATE(std::size_t{0}, 1, 2, 3, 7, 13, 16, 17, 97, 128, 1009); + const unsigned threads = GENERATE(1u, 2u, 3u, 7u, 8u, 0u); + const std::size_t block = GENERATE(std::size_t{0}, 1, 2, 3, 7, 64, 5000); + const int where = GENERATE(-1, 0, 1); // threshold at n + where, clamped at 0 + const std::size_t inline_below = where < 0 ? (n == 0 ? 0 : n - 1) : n + static_cast(where); + CAPTURE(n, threads, block, inline_below); + + Trace t{n}; + const std::thread::id caller = std::this_thread::get_id(); + for_each_block(n, threads, BlockSchedule{.block = block, .inline_below = inline_below}, t.sink()); + + for (std::size_t i = 0; i < n; ++i) { + CAPTURE(i); + REQUIRE(t.hits[i].load() == 1); + } + const std::size_t b = block == 0 ? default_block(n, resolved(threads)) : block; + const std::vector expected = partition(n, b); + std::vector sorted = t.calls; + std::sort(sorted.begin(), sorted.end()); + REQUIRE(sorted == expected); + + const bool inline_run = resolved(threads) == 1 || n < inline_below; + if (inline_run) { + REQUIRE(t.calls == expected); // ascending, on the caller + for (const auto id : t.by) REQUIRE(id == caller); + } + const std::set used(t.by.begin(), t.by.end()); + REQUIRE(used.size() <= std::max(1, std::min(resolved(threads), expected.size()))); +} + +TEST_CASE("for_each_block with n == 0 never calls fn", "[refinement][chunks][dynamic]") { + const unsigned threads = GENERATE(1u, 2u, 0u); + const std::size_t inline_below = GENERATE(std::size_t{0}, 1, 100); + int calls = 0; + for_each_block(0, threads, BlockSchedule{.block = 0, .inline_below = inline_below}, + [&calls](std::size_t, std::size_t) { ++calls; }); + REQUIRE(calls == 0); +} + +TEST_CASE("for_each_block at or above the threshold runs blocks concurrently", + "[refinement][chunks][dynamic]") { + // Two blocks, two threads, n == inline_below. Each block waits, up to a + // generous timeout, for the other to have started, which it can only do + // if they run on different threads at once. A scheduler that ran the + // round inline (the threshold ignored, or compared the wrong way) makes + // each wait time out. A correct one never waits more than a thread start. + const std::size_t inline_below = GENERATE(std::size_t{0}, 2); + CAPTURE(inline_below); + std::mutex m; + std::condition_variable cv; + int started = 0; + std::atomic timed_out{0}; + for_each_block(2, 2, BlockSchedule{.block = 1, .inline_below = inline_below}, + [&](std::size_t, std::size_t) { + std::unique_lock lock{m}; + ++started; + cv.notify_all(); + if (!cv.wait_for(lock, std::chrono::seconds{10}, [&] { return started >= 2; })) + timed_out.fetch_add(1); + }); + REQUIRE(started == 2); + REQUIRE(timed_out.load() == 0); +} + +TEST_CASE("for_each_block below the threshold runs on the calling thread even with many threads", + "[refinement][chunks][dynamic]") { + const std::size_t n = GENERATE(std::size_t{1}, 5, 99); + Trace t{n}; + const std::thread::id caller = std::this_thread::get_id(); + for_each_block(n, 8, BlockSchedule{.block = 1, .inline_below = n + 1}, t.sink()); + REQUIRE(t.calls == partition(n, 1)); + for (const auto id : t.by) REQUIRE(id == caller); +} + +TEST_CASE("for_each_block rethrows the lowest throwing block's exception after running every block", + "[refinement][chunks][dynamic]") { + // 64 items in blocks of 4 is 16 blocks. Every block from `first` on throws + // its own index. Block `first` sleeps before throwing, so on the threaded + // path higher blocks throw earlier in time: a scheduler that keeps the + // first exception to arrive, or the last, or the highest index, rethrows + // the wrong one. The inline paths (one thread; n below the threshold) must + // behave the same way. + const unsigned threads = GENERATE(1u, 4u, 8u); + const std::size_t inline_below = GENERATE(std::size_t{0}, 65); + const std::size_t first = GENERATE(std::size_t{0}, 1, 5, 15); + CAPTURE(threads, inline_below, first); + std::atomic calls{0}, short_blocks{0}; + std::size_t caught = 99; + try { + for_each_block(64, threads, BlockSchedule{.block = 4, .inline_below = inline_below}, + [&](std::size_t begin, std::size_t end) { + // No Catch2 assertion off the test thread: count instead. + if (end - begin != 4) short_blocks.fetch_add(1); + calls.fetch_add(1); + const std::size_t k = begin / 4; + if (k == first) std::this_thread::sleep_for(std::chrono::milliseconds{30}); + if (k >= first) throw BlockError{k}; + }); + } catch (const BlockError& e) { + caught = e.block; + } + REQUIRE(caught == first); + REQUIRE(short_blocks.load() == 0); + REQUIRE(calls.load() == 16); // no block is skipped after a throw +} + +TEST_CASE("for_each_block's exception is the same whatever the thread count and threshold", + "[refinement][chunks][dynamic]") { + // The throwing set is scattered, not a suffix: blocks 3, 6 and 11 of 13 + // (the last one short). Only block 3's exception may surface. + const unsigned threads = GENERATE(1u, 2u, 3u, 7u, 0u); + const std::size_t inline_below = GENERATE(std::size_t{0}, 50, 51); + CAPTURE(threads, inline_below); + std::atomic calls{0}; + std::size_t caught = 99; + try { + for_each_block(50, threads, BlockSchedule{.block = 4, .inline_below = inline_below}, + [&](std::size_t begin, std::size_t) { + calls.fetch_add(1); + const std::size_t k = begin / 4; + if (k == 3) std::this_thread::sleep_for(std::chrono::milliseconds{20}); + if (k == 3 || k == 6 || k == 11) throw BlockError{k}; + }); + } catch (const BlockError& e) { + caught = e.block; + } + REQUIRE(caught == 3); + REQUIRE(calls.load() == 13); +} From 8d171b5a0ad6abb609bca1d72cba43dc79cd6255 Mon Sep 17 00:00:00 2001 From: Ola Skavhaug Date: Sun, 27 Sep 2026 12:35:51 +0200 Subject: [PATCH 4/9] green: increment 21a (dynamic scan scheduling, active merge, legalise stack) QW1: parallel_util::for_each_block, blocks of default_block(n, threads) taken from one atomic work counter (the only shared write); inline on the calling thread with one thread or n < inline_below; per-block exception slots, lowest index rethrown. refine's scan uses it with BlockSchedule{}. QW3: detail::rebuild_active merges the ascending touched walk with the ascending skipped list instead of sort + unique. Stack: mesh::FlipStack and a legalise_around overload taking it; the old overload wraps it, and refine owns one stack for every insertion. inline_below = 256, from a thread sweep on the 1 m tile, recorded in docs/increments/21-parallel-refine.md ("21a: inline_below"). Output is bit-identical: T6 and the golden digests pass untouched, and bench.py's mesh sha256 for tile and quarter is unchanged. Co-Authored-By: Claude Opus 5.5 --- docs/increments/21-parallel-refine.md | 24 ++++++++ include/terrain/mesh/lawson.hpp | 18 +++++- include/terrain/parallel_util/chunks.hpp | 73 ++++++++++++++++++++++++ include/terrain/refinement/refine.hpp | 32 +++++++---- 4 files changed, 135 insertions(+), 12 deletions(-) diff --git a/docs/increments/21-parallel-refine.md b/docs/increments/21-parallel-refine.md index 664041bd..802e6e05 100644 --- a/docs/increments/21-parallel-refine.md +++ b/docs/increments/21-parallel-refine.md @@ -463,6 +463,30 @@ The three suites are registered in `tests/cpp/CMakeLists.txt` only once the header names `for_each_block`, `rebuild_active` or `FlipStack` (the 18 and 20b precedent). The review drops those guards. +### 21a: inline_below + +Set to **256** by `@developer` in the green step, from a sweep that is a +developer's quick look, not `@perf`'s acceptance run. Setup: the 1 m +benchmark's tile (`tests/fixtures/dem_archive/7908_3_10m_z33.tif`, tolerance +1, no domain), M1 Max on AC. A throwaway driver wrapped `cli.refine` the way +`tools/bench.py`'s child does and called it 9 times per thread count, taking +the median of `scan_seconds` and of the wall time of refine. `_core` was +rebuilt for each value. Median scan in ms at 8 threads (10 threads within +3 ms of it): + +| inline_below | 0 | 64 | 256 | 512 | 1024 | 2048 | 8192 | 32768 | +|---|---|---|---|---|---|---|---|---| +| scan, 8 threads | 58.4-58.8 | 57.2 | 57.3-57.5 | 57.5-57.6 | 58.8 | 65.9 | 83.3 | 119.7 | + +From 64 to 1024 the values differ by less than the run-to-run noise, and each +is about 1 ms under 0. From 2048 up, the inline small rounds cost more +than their thread starts save. 256 is in the middle of the flat range. Measured +back to back against the red commit (`7b54a4e`, `for_each_chunk`): scan at +8 threads 65.0 -> 57.7 ms and refine 243 -> 219 ms (-10 %). At 1 thread there +is no difference beyond noise (refine 515 vs 519 ms), as QW1 predicts. The +mesh sha256 that `bench.py` prints for the tile and the quarter circle was the same +before and after. + ## 4. Determinism levels Today's contract (14 R5, 14b R1, tested by 14's T6 and 18's T3 golden diff --git a/include/terrain/mesh/lawson.hpp b/include/terrain/mesh/lawson.hpp index 68165b08..27960558 100644 --- a/include/terrain/mesh/lawson.hpp +++ b/include/terrain/mesh/lawson.hpp @@ -78,10 +78,16 @@ template // Legalise around the vertex q after it was inserted. For each seed slot that // contains q, the edge opposite q is tested; a flip leaves q at index 0 of both // slots it writes, and both go back on the stack. +// +// The stack is the caller's, so a loop of insertions reuses one buffer +// (docs/increments/21-parallel-refine.md, 21a); its contents on entry are +// discarded, and it is empty on return. +using FlipStack = std::vector; + template std::size_t legalise_around(LatticeMesh& m, std::uint32_t q, std::span seeds, - const LatticeFrame& f, OnWrite&& on_write) { - std::vector stack(seeds.begin(), seeds.end()); + const LatticeFrame& f, FlipStack& stack, OnWrite&& on_write) { + stack.assign(seeds.begin(), seeds.end()); std::size_t flips = 0; while (!stack.empty()) { const auto t = stack.back(); @@ -103,6 +109,14 @@ std::size_t legalise_around(LatticeMesh& m, std::uint32_t q, std::span +std::size_t legalise_around(LatticeMesh& m, std::uint32_t q, std::span seeds, + const LatticeFrame& f, OnWrite&& on_write) { + FlipStack stack; + return legalise_around(m, q, seeds, f, stack, std::forward(on_write)); +} + // Legalise the whole mesh: every interior edge once on the stack, and each // flip pushes the four outer sides of its quad. A stale entry (its slot was // rewritten since) names some current edge and is simply tested. diff --git a/include/terrain/parallel_util/chunks.hpp b/include/terrain/parallel_util/chunks.hpp index 192702d9..cdbcde9a 100644 --- a/include/terrain/parallel_util/chunks.hpp +++ b/include/terrain/parallel_util/chunks.hpp @@ -16,8 +16,21 @@ // No pool. Threads are created per call and joined at its end, so no state // outlives a call. threads == 0 means hardware_concurrency, or 1 if that is 0. // With fewer items than threads, fewer threads start: a chunk is never empty. +// +// for_each_block is the dynamic variant the refinement scan uses +// (docs/increments/21-parallel-refine.md, QW1): [0, n) is cut into blocks of +// `block` items, block k being [k*b, min(n, (k+1)*b)), and workers take block +// indices from one shared atomic counter until it runs out, so a slow thread +// takes fewer blocks. The counter is the only shared write; which thread runs +// a block cannot change what fn writes, so 14 R7's "no shared result writes" +// holds. fn is called exactly once per block. With one thread or +// n < inline_below every block runs on the calling thread in ascending order, +// which spares small rounds the thread starts. Every block runs even when some +// throw, and the exception of the lowest block index that threw is rethrown, +// fixed by the partition, not by timing. #include +#include #include #include #include @@ -61,4 +74,64 @@ void for_each_chunk(std::size_t n, unsigned threads, Fn&& fn) { std::rethrow_exception(error); } +struct BlockSchedule { + std::size_t block = 0; // items per block; 0: default_block(n, threads) + // Rounds of fewer items run inline. From a thread sweep on the 1 m + // benchmark (docs/increments/21-parallel-refine.md, "21a: inline_below"). + std::size_t inline_below = 256; +}; + +[[nodiscard]] constexpr std::size_t default_block(std::size_t n, unsigned threads) noexcept { + return std::max(1, n / (std::size_t{16} * threads)); +} + +template +void for_each_block(std::size_t n, unsigned threads, BlockSchedule schedule, Fn&& fn) { + if (threads == 0) + threads = std::max(1u, std::thread::hardware_concurrency()); + const std::size_t b = schedule.block == 0 ? default_block(n, threads) : schedule.block; + const std::size_t blocks = n / b + (n % b != 0 ? 1 : 0); + const auto run = [&fn, n, b](std::size_t k) { + const std::size_t begin = k * b; + fn(begin, begin + std::min(b, n - begin)); + }; + const std::size_t workers = std::min(threads, blocks); + if (workers <= 1 || n < schedule.inline_below) { + std::exception_ptr first; // ascending, so the first to throw is the lowest + for (std::size_t k = 0; k < blocks; ++k) { + try { + run(k); + } catch (...) { + if (!first) + first = std::current_exception(); + } + } + if (first) + std::rethrow_exception(first); + return; + } + // Declared before the workers so they outlive the join; slot k is written + // only by the thread that took block k. + std::vector errors(blocks); + std::atomic next{0}; + { + std::vector team; + team.reserve(workers); + for (std::size_t w = 0; w < workers; ++w) + team.emplace_back([&] { + // Relaxed: the counter orders nothing; the join publishes fn's writes. + for (std::size_t k; (k = next.fetch_add(1, std::memory_order_relaxed)) < blocks;) { + try { + run(k); + } catch (...) { + errors[k] = std::current_exception(); + } + } + }); + } // the jthreads join here + for (const std::exception_ptr& error : errors) + if (error) + std::rethrow_exception(error); +} + } // namespace terrain::parallel_util diff --git a/include/terrain/refinement/refine.hpp b/include/terrain/refinement/refine.hpp index 7605f75d..c0c6a256 100644 --- a/include/terrain/refinement/refine.hpp +++ b/include/terrain/refinement/refine.hpp @@ -99,7 +99,7 @@ struct RefineOutcome { std::size_t feet_refused = 0; // 20b R2 step 4: a foot refused, N inserted instead // Wall seconds on the calling thread, steady_clock (17-mesh-stats.md R6). - // for_each_chunk joins its workers before returning, so no worker reads a + // for_each_block joins its workers before returning, so no worker reads a // clock. Not part of the determinism guarantee. double legalise_seconds = 0.0; // the start mesh's legalise_all double scan_seconds = 0.0; // every round's parallel scan, summed @@ -167,6 +167,23 @@ namespace detail { return std::move(*m); } +// The next round's active set: every touched slot and every skipped one, +// ascending and once each. `skipped` is ascending (filled while walking the +// ascending `active`) and below touched.size(), so one linear merge with the +// slot walk gives what collecting, sorting and deduplicating gave (21a, QW3). +inline void rebuild_active(std::span touched, std::span skipped, + std::vector& active) { + active.clear(); + auto s = skipped.begin(); + for (std::uint32_t t = 0; t < touched.size(); ++t) { + bool in = touched[t] != 0; + for (; s != skipped.end() && *s == t; ++s) + in = true; + if (in) + active.push_back(t); + } +} + [[nodiscard]] inline bool needs_split(const ScanResult& r, double tolerance) noexcept { return r.node.has_value() && (r.is_void || r.max_error > tolerance); } @@ -279,6 +296,7 @@ template } std::vector results; std::set> footed; // (row, col), R2 step 5 + mesh::FlipStack flip_stack; // one buffer for every legalise_around std::vector active(m.triangle_count()); for (std::uint32_t t = 0; t < active.size(); ++t) active[t] = t; @@ -287,7 +305,7 @@ template ++out.rounds; results.resize(m.triangle_count()); t0 = clock::now(); - parallel_util::for_each_chunk(active.size(), options.threads, + parallel_util::for_each_block(active.size(), options.threads, parallel_util::BlockSchedule{}, [&](std::size_t begin, std::size_t end) { for (std::size_t i = begin; i < end; ++i) results[active[i]] = scan(dem, m, active[i]); @@ -345,7 +363,7 @@ template if (n_seeds == 4) touched[seeds[3]] = 1; out.flips += mesh::legalise_around( - m, q, std::span{seeds.data(), n_seeds}, frame, + m, q, std::span{seeds.data(), n_seeds}, frame, flip_stack, [&](std::uint32_t s) { touched[s] = 1; }); ++out.inserted; out.carved += r.is_void ? 1 : 0; @@ -357,13 +375,7 @@ template out.split_seconds += since(t0); if (!any) break; - active.clear(); - for (std::uint32_t t = 0; t < touched.size(); ++t) - if (touched[t] != 0) - active.push_back(t); - active.insert(active.end(), skipped.begin(), skipped.end()); - std::sort(active.begin(), active.end()); - active.erase(std::unique(active.begin(), active.end()), active.end()); + detail::rebuild_active(touched, skipped, active); } // By the stopping rule a void triangle holds no valid node, so `uncovered` From 065a2bd69ad767133dfa172f30b5e58959696574 Mon Sep 17 00:00:00 2001 From: Ola Skavhaug Date: Sun, 27 Sep 2026 12:36:59 +0200 Subject: [PATCH 5/9] 21a: remove the red-step scaffold, register the three suites plainly The 16 precedent (07317f5): the guards existed so the red commit did not break the build; with green in, they only hide a suite whose name drifts. Co-Authored-By: Claude Opus 5.5 --- tests/cpp/CMakeLists.txt | 36 +++++++----------------------------- 1 file changed, 7 insertions(+), 29 deletions(-) diff --git a/tests/cpp/CMakeLists.txt b/tests/cpp/CMakeLists.txt index efd973cf..a21cc2c2 100644 --- a/tests/cpp/CMakeLists.txt +++ b/tests/cpp/CMakeLists.txt @@ -284,32 +284,10 @@ target_link_libraries(test_refinement_constraint_feet PRIVATE Threads::Threads) # invariant-critical suite; test_refinement_active (QW3) and # test_mesh_lawson_stack (legalise_around's caller-owned stack) compare with # today's code. All three run in the TSan job. -# -# RED UNTIL THE INTERFACES EXIST, keyed on the name each suite needs, because -# 21a adds no header: each suite is registered only once its header names it, -# so the red does not break the rest of the build. The green commit makes them -# build; the review drops these guards. -# The three headers are configure dependencies, so the green commit's edit -# re-runs the configure step and registers the suites without a manual cmake. -set_property(DIRECTORY APPEND PROPERTY CMAKE_CONFIGURE_DEPENDS - "${PROJECT_SOURCE_DIR}/include/terrain/parallel_util/chunks.hpp" - "${PROJECT_SOURCE_DIR}/include/terrain/refinement/refine.hpp" - "${PROJECT_SOURCE_DIR}/include/terrain/mesh/lawson.hpp") -file(READ "${PROJECT_SOURCE_DIR}/include/terrain/parallel_util/chunks.hpp" _chunks_hpp) -string(FIND "${_chunks_hpp}" "for_each_block" _block_at) -if(NOT _block_at EQUAL -1) - add_terrain_test(test_refinement_chunks_dynamic unit/test_refinement_chunks_dynamic.cpp) - target_link_libraries(test_refinement_chunks_dynamic PRIVATE Threads::Threads) -endif() -file(READ "${PROJECT_SOURCE_DIR}/include/terrain/refinement/refine.hpp" _refine_21a) -string(FIND "${_refine_21a}" "rebuild_active" _active_at) -if(NOT _active_at EQUAL -1) - add_terrain_backend_test(test_refinement_active unit/test_refinement_active.cpp) - target_link_libraries(test_refinement_active PRIVATE Threads::Threads) -endif() -file(READ "${PROJECT_SOURCE_DIR}/include/terrain/mesh/lawson.hpp" _lawson_hpp) -string(FIND "${_lawson_hpp}" "FlipStack" _stack_at) -if(NOT _stack_at EQUAL -1) - add_terrain_backend_test(test_mesh_lawson_stack unit/test_mesh_lawson_stack.cpp) - target_link_libraries(test_mesh_lawson_stack PRIVATE Threads::Threads) -endif() + +add_terrain_test(test_refinement_chunks_dynamic unit/test_refinement_chunks_dynamic.cpp) +target_link_libraries(test_refinement_chunks_dynamic PRIVATE Threads::Threads) +add_terrain_backend_test(test_refinement_active unit/test_refinement_active.cpp) +target_link_libraries(test_refinement_active PRIVATE Threads::Threads) +add_terrain_backend_test(test_mesh_lawson_stack unit/test_mesh_lawson_stack.cpp) +target_link_libraries(test_mesh_lawson_stack PRIVATE Threads::Threads) From 27055a94ca79ff4fba751d9588f721853c9f44a3 Mon Sep 17 00:00:00 2001 From: Ola Skavhaug Date: Sun, 27 Sep 2026 12:48:07 +0200 Subject: [PATCH 6/9] =?UTF-8?q?21a=20review=20fixes:=20scaffold=20sentence?= =?UTF-8?q?,=20inline=5Fbelow=20prose=20vs=20table,=20project=5Fstructure'?= =?UTF-8?q?s=20chunks.hpp,=20the=20TSan=20job=20comment,=20=C2=A77=20suite?= =?UTF-8?q?=20name?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Co-Authored-By: Claude Opus 5.5 --- .github/workflows/main.yaml | 3 ++- docs/increments/21-parallel-refine.md | 12 ++++++------ project_structure.md | 4 ++-- 3 files changed, 10 insertions(+), 9 deletions(-) diff --git a/.github/workflows/main.yaml b/.github/workflows/main.yaml index f105b46e..f1245ed6 100644 --- a/.github/workflows/main.yaml +++ b/.github/workflows/main.yaml @@ -70,7 +70,8 @@ jobs: # Increment 14, R7: refinement is the first code path that starts threads, # and asan+ubsan cannot see a data race. TSan cannot share a build with ASan, - # so this is its own job, building only the four refinement suites. T6 runs + # so this is its own job, building only the suites listed below (refinement and + # the parallel helpers they use). T6 runs # the loop at 1, 2, 7 and all threads; a clean run is the independent check # of the design's "no shared writes" claim. The binaries are run directly # because ctest names the tests by Catch case, not by target. diff --git a/docs/increments/21-parallel-refine.md b/docs/increments/21-parallel-refine.md index 802e6e05..232a42e5 100644 --- a/docs/increments/21-parallel-refine.md +++ b/docs/increments/21-parallel-refine.md @@ -459,9 +459,9 @@ red step, and the suites named below hold them. same flips, `on_write` sequence and mesh as today's algorithm, which the suite copies as its oracle. Allocation counts are not pinned. -The three suites are registered in `tests/cpp/CMakeLists.txt` only once the -header names `for_each_block`, `rebuild_active` or `FlipStack` (the 18 and -20b precedent). The review drops those guards. +At the red commit the three suites were registered only once their header +named `for_each_block`, `rebuild_active` or `FlipStack` (the 18 and 20b +precedent); `065a2bd` removed those guards after green. ### 21a: inline_below @@ -478,8 +478,8 @@ rebuilt for each value. Median scan in ms at 8 threads (10 threads within |---|---|---|---|---|---|---|---|---| | scan, 8 threads | 58.4-58.8 | 57.2 | 57.3-57.5 | 57.5-57.6 | 58.8 | 65.9 | 83.3 | 119.7 | -From 64 to 1024 the values differ by less than the run-to-run noise, and each -is about 1 ms under 0. From 2048 up, the inline small rounds cost more +From 64 to 1024 the values differ by less than the run-to-run noise; 64-512 +are about 1 ms under 0, and 1024 is level with it. From 2048 up, the inline small rounds cost more than their thread starts save. 256 is in the middle of the flat range. Measured back to back against the red commit (`7b54a4e`, `for_each_chunk`): scan at 8 threads 65.0 -> 57.7 ms and refine 243 -> 219 ms (-10 %). At 1 thread there @@ -837,7 +837,7 @@ is worth counting early. | increment | what | est. | at +66 % | determinism | invariant-critical suite (mutation round) | |---|---|---|---|---|---| -| **21a** | QW1 dynamic scan scheduling with a small-round inline threshold (`parallel_util/`), QW3 merge, `legalise_around`'s caller-owned stack; QW4 only if measured worth it (about +20) | ~60 | ~100 | bit-identical | `test_refinement_chunks`, extended to the dynamic scheduler: every index visited exactly once for every n, thread count and block size. A dropped or doubled block leaves a stale scan result, which breaks the tolerance guarantee silently, so this is where the guarantee is decided | +| **21a** | QW1 dynamic scan scheduling with a small-round inline threshold (`parallel_util/`), QW3 merge, `legalise_around`'s caller-owned stack; QW4 only if measured worth it (about +20) | ~60 | ~100 | bit-identical | `test_refinement_chunks_dynamic`, a new suite for the dynamic scheduler (`test_refinement_chunks` is unchanged): every index visited exactly once for every n, thread count and block size. A dropped or doubled block leaves a stale scan result, which breaks the tolerance guarantee silently, so this is where the guarantee is decided | | **21b** | QW2 `lattice_incircle` in `mesh/`, the exact-frame check, the call in `must_flip` | ~50 | ~85 | bit-identical where it answers | a new `test_mesh_lattice_incircle`: agreement with `DetriaExact` on the frame doubles for random and adversarial node quads (grid rectangles, other cocircular lattice quads such as points on a circle of radius 5, near-overflow differences at 2^14, `dx != dy` and inexact `dx` returning `nullopt`, off-node corners returning `nullopt`). Mutants: bound at 2^15, `dx != dy` not refused, the exact-frame check dropped, a sign flip | | **21c** | measurement only, `@perf`; evidence under `docs/benchmarks//` | 0 | 0 | — | none | | **21d** (option C) | thread team (`std::barrier`, per call), batch split with prefix-sum numbering, deterministic parallel flip rounds, the round loop | ~320 | ~530 | L1 | a new `test_mesh_parallel_lawson`: CDT property, conformity and a flip-count bound on random lattice meshes with cocircular ties and constraints, and identical output for threads 1, 2, 7 and hardware concurrency. Plus `prop_refinement_refine`'s tolerance oracle re-run, with the mutant "a flipped slot not marked touched" | diff --git a/project_structure.md b/project_structure.md index c818493b..bfba4743 100644 --- a/project_structure.md +++ b/project_structure.md @@ -47,7 +47,7 @@ include/terrain/ # public C++ headers, header-only where possible window.hpp # window_for: bbox -> index window (planned; unbuilt, # refinement walks lattice nodes directly) parallel_util/ - chunks.hpp # for_each_chunk: contiguous chunks over std::jthread + chunks.hpp # for_each_chunk (contiguous chunks) and for_each_block (dynamic blocks) over std::jthread mesh/ lattice_mesh.hpp # LatticeMesh: flat triangle array over DEM nodes, # neighbour links, the three splits (14), flip (14b) @@ -291,7 +291,7 @@ only TU that includes `detria.hpp`, enforced by CMake privacy, an `#error` guard ### `parallel_util` -Header-only. Today one helper, `for_each_chunk(n, threads, fn)` in `chunks.hpp`: contiguous chunks over `std::jthread`, created per call and joined before it returns, no pool. `std::execution::par` and OpenMP were both ruled out in increment 14 (R7): neither builds on macOS without an experimental flag or an extra runtime. Needs only `Threads::Threads`. +Header-only. Two helpers in `chunks.hpp`, both over `std::jthread` created per call and joined before it returns, no pool: `for_each_chunk(n, threads, fn)`, contiguous equal-count chunks; and `for_each_block(n, threads, BlockSchedule, fn)`, blocks handed out from one atomic counter, run inline below `BlockSchedule::inline_below`, which refine's scan uses (increment 21a). `std::execution::par` and OpenMP were both ruled out in increment 14 (R7): neither builds on macOS without an experimental flag or an extra runtime. Needs only `Threads::Threads`. ### `vector_simplify` From cade7522e2c37c224868762ba86ad4e585c1282e Mon Sep 17 00:00:00 2001 From: Ola Skavhaug Date: Sun, 27 Sep 2026 12:53:13 +0200 Subject: [PATCH 7/9] 21a tests: exception order both ways; non-empty FlipStack on entry Test amendment after green, from @reviewer's findings on 8d171b5. No production code; @developer does not edit tests, so it lands here. test_refinement_chunks_dynamic: in both exception cases the lowest-index thrower was also the last to arrive, so a scheduler that kept the last-to-arrive exception passed (the reviewer's mutant, 5/5). New case: the lowest thrower throws at once and a higher one waits for it, then throws, so arrival order is reversed. The existing case now forces its order with a bounded wait on the higher blocks instead of a bare sleep. Keep-last is killed by the new case, keep-first by the existing ones. The header no longer calls this suite test_refinement_chunks extended. test_mesh_lawson_stack: a case passing a non-empty FlipStack on entry, pinning lawson.hpp's "contents on entry are discarded" and the empty stack on return. Kills seeds-below-stale; seeds-above-valid-stale is observationally equivalent under Lawson's post-condition, and the comment says so. The insertion step is factored into a helper both cases share. Co-Authored-By: Claude Opus 5.5 --- tests/cpp/unit/test_mesh_lawson_stack.cpp | 101 +++++++++++++++--- .../unit/test_refinement_chunks_dynamic.cpp | 86 +++++++++++++-- 2 files changed, 162 insertions(+), 25 deletions(-) diff --git a/tests/cpp/unit/test_mesh_lawson_stack.cpp b/tests/cpp/unit/test_mesh_lawson_stack.cpp index 1c6f50ea..3c47db0a 100644 --- a/tests/cpp/unit/test_mesh_lawson_stack.cpp +++ b/tests/cpp/unit/test_mesh_lawson_stack.cpp @@ -28,6 +28,7 @@ #include "refinement_fixtures.hpp" +#include #include #include #include @@ -137,6 +138,33 @@ std::optional locate(const LatticeMesh& m, LatticeVertex p) { return std::nullopt; } +// Insert p as refine does: split_inside for an interior node, split_edge for +// one on an edge (the ring's constrained edges included), with refine's +// seeds. nullopt when p is already a vertex. +struct Insertion { + std::uint32_t q; + std::array seeds; + std::size_t n_seeds; + bool on_edge; + std::span span() const { return {seeds.data(), n_seeds}; } +}; +std::optional insert(LatticeMesh& m, LatticeVertex p) { + const auto hit = locate(m, p); + if (!hit) return std::nullopt; + const auto t = hit->t; + const auto before = static_cast(m.triangle_count()); + Insertion ins{0, {t, before, before + 1, before + 1}, 3, hit->edge.has_value()}; + if (!hit->edge) { + ins.q = m.split_inside(t, p); + } else { + const std::uint32_t u = m.neighbours(t)[*hit->edge]; + ins.q = m.split_edge(t, *hit->edge, p); + ins.n_seeds = u != kNoNeighbour ? 4 : 2; + if (ins.n_seeds == 4) ins.seeds[3] = u; + } + return ins; +} + void require_same_mesh(const LatticeMesh& a, const LatticeMesh& b) { REQUIRE(a.triangle_count() == b.triangle_count()); REQUIRE(a.vertices().size() == b.vertices().size()); @@ -175,23 +203,11 @@ TEST_CASE("legalise_around with a caller-owned stack matches today's, insertion std::size_t inserted = 0, flipped = 0, on_edges = 0; for (int attempt = 0; attempt < 300; ++attempt) { const LatticeVertex p{static_cast(rng() % 25), static_cast(rng() % 25)}; - const auto hit = locate(m, p); - if (!hit) continue; - const auto t = hit->t; - const auto before = static_cast(m.triangle_count()); - std::array seeds{t, before, before + 1, before + 1}; - std::size_t n_seeds = 3; - std::uint32_t q = 0; - if (!hit->edge) { - q = m.split_inside(t, p); - } else { - const std::uint32_t u = m.neighbours(t)[*hit->edge]; - q = m.split_edge(t, *hit->edge, p); - n_seeds = u != kNoNeighbour ? 4 : 2; - if (n_seeds == 4) seeds[3] = u; - ++on_edges; - } - const std::span s{seeds.data(), n_seeds}; + const auto ins = insert(m, p); + if (!ins) continue; + if (ins->on_edge) ++on_edges; + const std::uint32_t q = ins->q; + const auto s = ins->span(); LatticeMesh expected = m; std::vector expected_writes; const std::size_t expected_flips = reference_around(expected, q, s, f, expected_writes); @@ -231,3 +247,54 @@ TEST_CASE("legalise_around with a caller-owned stack and no flip to make", "[law REQUIRE(writes == expected_writes); require_same_mesh(m, expected); } + +TEST_CASE("legalise_around discards a non-empty FlipStack's contents on entry", "[lawson][around][stack]") { + // lawson.hpp: the stack's "contents on entry are discarded, and it is + // empty on return". Before every insertion the stack is loaded with stale + // entries: the new seeds, reversed so they pop in the opposite order to the + // reference's, on top of triangle 0 and the last triangle. The result must still be the reference's (flip count, on_write + // sequence in order, mesh), and the stack empty on return. + // + // What this kills: an overload that keeps the stale entries on top of the + // seeds (prepends the seeds), which pops them first and so flips in a + // different order; one that leaves anything behind. What it cannot see: + // an overload that appends the seeds above valid stale entries. Those are + // popped after the seeds, when Lawson's post-condition already holds (no + // edge opposite q must flip, and triangles without q are skipped), so the + // result is identical. An out-of-range stale entry would expose it, but + // only as undefined behaviour, which is not a test. + const std::uint32_t seed = GENERATE(range(1u, 9u)); + const std::size_t frame = GENERATE(std::size_t{0}, 1, 2); + CAPTURE(seed, frame); + const LatticeFrame f = kFrames[frame]; + auto m = grid(4, 6, true); + std::mt19937 rng{seed}; + FlipStack stack; + std::size_t inserted = 0, flipped = 0; + for (int attempt = 0; attempt < 300; ++attempt) { + const LatticeVertex p{static_cast(rng() % 25), static_cast(rng() % 25)}; + const auto ins = insert(m, p); + if (!ins) continue; + const auto s = ins->span(); + LatticeMesh expected = m; + std::vector expected_writes; + const std::size_t expected_flips = reference_around(expected, ins->q, s, f, expected_writes); + + stack.assign(s.begin(), s.end()); + stack.push_back(0); + stack.push_back(static_cast(m.triangle_count() - 1)); + std::reverse(stack.begin(), stack.end()); // seeds last, reversed: popped first + std::vector writes; + const std::size_t flips = legalise_around( + m, ins->q, s, f, stack, [&writes](std::uint32_t w) { writes.push_back(w); }); + CAPTURE(attempt, p.row, p.col); + REQUIRE(flips == expected_flips); + REQUIRE(writes == expected_writes); + REQUIRE(stack.empty()); + require_same_mesh(m, expected); + ++inserted; + flipped += flips; + } + REQUIRE(inserted >= 100); + REQUIRE(flipped >= 20); +} diff --git a/tests/cpp/unit/test_refinement_chunks_dynamic.cpp b/tests/cpp/unit/test_refinement_chunks_dynamic.cpp index bacb46de..84b864c7 100644 --- a/tests/cpp/unit/test_refinement_chunks_dynamic.cpp +++ b/tests/cpp/unit/test_refinement_chunks_dynamic.cpp @@ -1,6 +1,7 @@ // Increment 21a, QW1 (docs/increments/21-parallel-refine.md, section 3, and // "Pinned by the red suite (21a)"): the dynamic block scheduler that replaces -// for_each_chunk for the scan. test_refinement_chunks extended to it. +// for_each_chunk for the scan. A new suite; test_refinement_chunks, which pins +// for_each_chunk, is unchanged. // // INVARIANT-CRITICAL (section 7): a dropped block leaves a stale scan result // and a doubled one is a second writer to a result slot, and either breaks the @@ -181,19 +182,28 @@ TEST_CASE("for_each_block below the threshold runs on the calling thread even wi for (const auto id : t.by) REQUIRE(id == caller); } +// The two exception-order cases below make arrival order opposite ways, so +// that neither "keep the first exception to arrive" nor "keep the last" can +// pass both. Arrival order is forced with waits on what the other blocks have +// done, not left to sleeps racing each other; every wait has a timeout, so a +// scheduler that serialises the blocks is slow, never hung. + TEST_CASE("for_each_block rethrows the lowest throwing block's exception after running every block", "[refinement][chunks][dynamic]") { // 64 items in blocks of 4 is 16 blocks. Every block from `first` on throws - // its own index. Block `first` sleeps before throwing, so on the threaded - // path higher blocks throw earlier in time: a scheduler that keeps the - // first exception to arrive, or the last, or the highest index, rethrows - // the wrong one. The inline paths (one thread; n below the threshold) must - // behave the same way. + // its own index. On the threaded path block `first` waits until every + // higher block is about to throw, then a little longer, so it is the LAST + // to arrive. Kills: keep-the-first-to-arrive, keep-the-highest-index. + // (Keep-the-last-to-arrive passes this case; the next one kills it.) The + // inline paths (one thread; n below the threshold) run in ascending order, + // so there block `first` arrives first; they must rethrow the same block. const unsigned threads = GENERATE(1u, 4u, 8u); const std::size_t inline_below = GENERATE(std::size_t{0}, 65); const std::size_t first = GENERATE(std::size_t{0}, 1, 5, 15); CAPTURE(threads, inline_below, first); - std::atomic calls{0}, short_blocks{0}; + const bool threaded = threads > 1 && inline_below <= 64; + const int higher = static_cast(15 - first); + std::atomic calls{0}, short_blocks{0}, about_to_throw{0}, timed_out{0}; std::size_t caught = 99; try { for_each_block(64, threads, BlockSchedule{.block = 4, .inline_below = inline_below}, @@ -202,17 +212,77 @@ TEST_CASE("for_each_block rethrows the lowest throwing block's exception after r if (end - begin != 4) short_blocks.fetch_add(1); calls.fetch_add(1); const std::size_t k = begin / 4; - if (k == first) std::this_thread::sleep_for(std::chrono::milliseconds{30}); + if (k == first && threaded) { + const auto deadline = std::chrono::steady_clock::now() + std::chrono::seconds{10}; + while (about_to_throw.load() < higher) { + if (std::chrono::steady_clock::now() > deadline) { + timed_out.fetch_add(1); + break; + } + std::this_thread::yield(); + } + // Let the higher blocks' throws reach the scheduler. + std::this_thread::sleep_for(std::chrono::milliseconds{30}); + } + if (k > first) about_to_throw.fetch_add(1); if (k >= first) throw BlockError{k}; }); } catch (const BlockError& e) { caught = e.block; } + REQUIRE(timed_out.load() == 0); REQUIRE(caught == first); REQUIRE(short_blocks.load() == 0); REQUIRE(calls.load() == 16); // no block is skipped after a throw } +TEST_CASE("for_each_block rethrows the lowest throwing block's exception when it arrives first", + "[refinement][chunks][dynamic]") { + // The reverse arrival order: block `low` throws at once, and block `high` + // waits until `low` is about to throw, then a little longer, so `low` is + // the FIRST to arrive and `high` the LAST, on every path (the inline paths + // run `low` first anyway). Kills: keep-the-last-to-arrive, and + // keep-the-highest-index. `high` is block 15, the last one handed out, or a + // middle one, so the slow block is not always the scheduler's last. + const unsigned threads = GENERATE(1u, 2u, 4u, 8u); + const std::size_t inline_below = GENERATE(std::size_t{0}, 65); + const std::size_t low = GENERATE(std::size_t{0}, 5); + const std::size_t high = GENERATE(std::size_t{9}, 15); + CAPTURE(threads, inline_below, low, high); + std::atomic calls{0}, timed_out{0}; + std::atomic low_thrown{false}; + std::size_t caught = 99; + try { + for_each_block(64, threads, BlockSchedule{.block = 4, .inline_below = inline_below}, + [&](std::size_t begin, std::size_t) { + calls.fetch_add(1); + const std::size_t k = begin / 4; + if (k == low) { + low_thrown.store(true); + throw BlockError{k}; + } + if (k == high) { + const auto deadline = std::chrono::steady_clock::now() + std::chrono::seconds{10}; + while (!low_thrown.load()) { + if (std::chrono::steady_clock::now() > deadline) { + timed_out.fetch_add(1); + break; + } + std::this_thread::yield(); + } + // Let block `low`'s throw reach the scheduler first. + std::this_thread::sleep_for(std::chrono::milliseconds{30}); + throw BlockError{k}; + } + }); + } catch (const BlockError& e) { + caught = e.block; + } + REQUIRE(timed_out.load() == 0); + REQUIRE(caught == low); + REQUIRE(calls.load() == 16); +} + TEST_CASE("for_each_block's exception is the same whatever the thread count and threshold", "[refinement][chunks][dynamic]") { // The throwing set is scattered, not a suffix: blocks 3, 6 and 11 of 13 From 2fc54a1edaa38a7c69b5cfc37eee24a607f082b4 Mon Sep 17 00:00:00 2001 From: Ola Skavhaug Date: Sun, 27 Sep 2026 13:07:23 +0200 Subject: [PATCH 8/9] 21a acceptance: ACCEPTED on AC, refine -11 % (tile) / -15 % (quarter) at 8 threads vs 227522c, mesh identical Back-to-back tools/bench.py runs of master 227522c (worktree, --tree) and 21a's HEAD cade752, twice each, all on AC: threads 1..20, 5 repeats, tile and quarter circle, tolerance 1. Ceiling 2.07-2.14x -> 2.28-2.29x. Mesh sha256, worst angle, max degree, tolerance and Delaunay checks identical. Median-of-5 spread between runs: max 4.6 % (base), 0.85 % (21a). Co-Authored-By: Claude Opus 5.5 --- docs/benchmarks/2026-09-27/21a-acceptance.md | 145 + .../2026-09-27/21a-base-227522c-r2/README.md | 79 + .../2026-09-27/21a-base-227522c-r2/raw.tsv | 210 ++ .../2026-09-27/21a-base-227522c-r2/run.json | 2955 +++++++++++++++++ .../2026-09-27/21a-base-227522c/README.md | 79 + .../2026-09-27/21a-base-227522c/raw.tsv | 210 ++ .../2026-09-27/21a-base-227522c/run.json | 2955 +++++++++++++++++ docs/benchmarks/2026-09-27/21a-r2/README.md | 81 + docs/benchmarks/2026-09-27/21a-r2/raw.tsv | 210 ++ docs/benchmarks/2026-09-27/21a-r2/run.json | 2955 +++++++++++++++++ docs/benchmarks/2026-09-27/21a/README.md | 81 + docs/benchmarks/2026-09-27/21a/raw.tsv | 210 ++ docs/benchmarks/2026-09-27/21a/run.json | 2955 +++++++++++++++++ 13 files changed, 13125 insertions(+) create mode 100644 docs/benchmarks/2026-09-27/21a-acceptance.md create mode 100644 docs/benchmarks/2026-09-27/21a-base-227522c-r2/README.md create mode 100644 docs/benchmarks/2026-09-27/21a-base-227522c-r2/raw.tsv create mode 100644 docs/benchmarks/2026-09-27/21a-base-227522c-r2/run.json create mode 100644 docs/benchmarks/2026-09-27/21a-base-227522c/README.md create mode 100644 docs/benchmarks/2026-09-27/21a-base-227522c/raw.tsv create mode 100644 docs/benchmarks/2026-09-27/21a-base-227522c/run.json create mode 100644 docs/benchmarks/2026-09-27/21a-r2/README.md create mode 100644 docs/benchmarks/2026-09-27/21a-r2/raw.tsv create mode 100644 docs/benchmarks/2026-09-27/21a-r2/run.json create mode 100644 docs/benchmarks/2026-09-27/21a/README.md create mode 100644 docs/benchmarks/2026-09-27/21a/raw.tsv create mode 100644 docs/benchmarks/2026-09-27/21a/run.json diff --git a/docs/benchmarks/2026-09-27/21a-acceptance.md b/docs/benchmarks/2026-09-27/21a-acceptance.md new file mode 100644 index 00000000..1cd3a575 --- /dev/null +++ b/docs/benchmarks/2026-09-27/21a-acceptance.md @@ -0,0 +1,145 @@ +# Increment 21a acceptance run (@perf, 2026-09-27) + +**Verdict: ACCEPTED.** On AC, against the previous merge (`227522c`) measured +back to back: refine at 8 threads is 11 % faster on the tile and 15 % faster on +the quarter circle. The mesh sha256 is identical on both domains. No measure +regressed. + +## Method + +- **Why a re-measured base.** The only stored `bench.py` run + (`serial-profile-bench/`) is on battery, so there is no AC baseline. As + `docs/increments/README.md` ("Acceptance") says to, master `227522c` (the + previous merge, #101) was checked out in a temporary `git worktree` and + measured with `--tree`, back to back with 21a's HEAD `cade752`. The worktree + was removed afterwards. +- **Two pairs.** Each tree was measured twice, in the order base, 21a, base, + 21a, between 12:54 and 13:05 CEST. The second pair (`-r2`) is there to + measure how much the median of 5 moves between runs (see "Spread" below). + `bench.py`'s verdicts: `21a` against `21a-base-227522c` (`--baseline`), + `21a-r2` against `21a-base-227522c-r2` (`--baseline`), and + `21a-base-227522c-r2` against `21a-base-227522c` (found by the baseline + search). All three are ACCEPTED. +- **Command** (the main tree's `tools/bench.py`, blob `77765b1`, for all four + runs; each run did its own Release build into `/build-bench`): + + ``` + .venv/bin/python tools/bench.py run --label