Skip to content

Increment 21a: dynamic scan scheduling, active merge, reused flip stack - #102

Merged
skavhaug merged 9 commits into
masterfrom
increment21a-quick-wins
Sep 27, 2026
Merged

skavhaug merged 9 commits into
masterfrom
increment21a-quick-wins

Conversation

@skavhaug

Copy link
Copy Markdown
Member

What

Increment 21a (docs/increments/21-parallel-refine.md, §3 QW1/QW3 and the legalise stack). Bit-identical output: T6 variants and 18's golden digests pass untouched; bench mesh sha256 equal to master on both domains.

  • parallel_util::for_each_block: blocks handed out from one atomic counter, inline below inline_below = 256 (swept on AC); per-block exception slots, lowest block index rethrown. refine's scan uses it. for_each_chunk unchanged.
  • refinement::detail::rebuild_active: next round's active by linear merge, not sort+unique.
  • mesh::FlipStack: one stack owned by refine and reused by legalise_around.
  • 78 production lines (est. ~60).

Red 7b54a4e before green 8d171b5; scaffold removed 065a2bd; test amendment cade752 (exception order both ways, non-empty FlipStack). Invariant-critical suite test_refinement_chunks_dynamic: mutation round, 9/9 plus the reviewer's keep-last mutant.

Acceptance (@Perf, AC, back to back with master 227522c, run twice)

refine, median of 5 master 21a
tile, 8 threads 249.5 ms 222.4 ms (-10.9 %)
quarter, 8 threads 230.4 ms 196.5 ms (-14.7 %)
ceiling (best) 2.1x 2.3x

Evidence: docs/benchmarks/2026-09-27/21a-acceptance.md. Verdict ACCEPTED.

Review

@Reviewer: CHANGES REQUESTED once (prose + a test gap), APPROVED at d3ee2ce. Local: ctest 762/762, TSan clean on all 15 CI suites, pytest 1719 passed.

Follow-ups (non-blocking): bench.py's dirty flag should ignore untracked files and record the baseline used; a few wording nits.

🤖 Generated with Claude Code

skavhaug and others added 9 commits September 27, 2026 11:53
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tack)

21a is bit-identical, so 14's T6 and 18's golden digests stay as they are
and remain the output oracle. Three new suites pin the interfaces section 3
left open; the pins are written into 21-parallel-refine.md under "Pinned by
the red suite (21a)".

test_refinement_chunks_dynamic (invariant-critical, QW1): for_each_block,
BlockSchedule and default_block. Every index is visited exactly once, and
the calls are exactly the block partition, for n in {0, 1, primes, 16, 17,
128, 1009}, threads {1, 2, 3, 7, 8, 0}, seven block sizes and the inline
threshold below, at and above n. One thread or n < inline_below runs on the
caller in ascending order. At or above the threshold two blocks rendezvous,
which proves they run concurrently. On the exception contract, every block
runs and the lowest throwing block index is rethrown, with the lowest
thrower delayed so the order of arrival differs from the order of index.

Mutation round against a scratch implementation that is not committed. All
nine mutants were killed: a dropped last partial block, a doubled block 0,
the highest-index exception, the first-in-time exception, always inline,
never inline, <= at the threshold, stop after a throw, and a default_block
with /8 (a compile-time failure). The scratch implementation passed 30
randomised-order runs and TSan.

test_refinement_active (QW3): detail::rebuild_active equals today's
collect + sort + unique, copied from 09b005b, on empty, duplicate,
interleaved, all-touched and 40 random rounds, including a reused buffer.

test_mesh_lawson_stack: legalise_around(..., FlipStack&, on_write) gives the
same flip count, on_write sequence and mesh as today's algorithm, copied as
the oracle. The comparison runs after each of more than 100 random
insertions (inside and edge splits, ring constraints) in three frames.

All three are registered only once their header names for_each_block,
rebuild_active or FlipStack (the 18 and 20b precedent), so the main build
stays green. The headers are configure dependencies. All three are in the
TSan job, which stays red until green lands.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… stack)

QW1: parallel_util::for_each_block, blocks of default_block(n, threads) taken
from one atomic work counter (the only shared write); inline on the calling
thread with one thread or n < inline_below; per-block exception slots, lowest
index rethrown. refine's scan uses it with BlockSchedule{}.
QW3: detail::rebuild_active merges the ascending touched walk with the
ascending skipped list instead of sort + unique.
Stack: mesh::FlipStack and a legalise_around overload taking it; the old
overload wraps it, and refine owns one stack for every insertion.

inline_below = 256, from a thread sweep on the 1 m tile, recorded in
docs/increments/21-parallel-refine.md ("21a: inline_below"). Output is
bit-identical: T6 and the golden digests pass untouched, and bench.py's mesh
sha256 for tile and quarter is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The 16 precedent (07317f5): the guards existed so the red commit did not
break the build; with green in, they only hide a suite whose name drifts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ject_structure's chunks.hpp, the TSan job comment, §7 suite name

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Test amendment after green, from @Reviewer's findings on 8d171b5. No
production code; @developer does not edit tests, so it lands here.

test_refinement_chunks_dynamic: in both exception cases the lowest-index
thrower was also the last to arrive, so a scheduler that kept the
last-to-arrive exception passed (the reviewer's mutant, 5/5). New case:
the lowest thrower throws at once and a higher one waits for it, then
throws, so arrival order is reversed. The existing case now forces its
order with a bounded wait on the higher blocks instead of a bare sleep.
Keep-last is killed by the new case, keep-first by the existing ones.
The header no longer calls this suite test_refinement_chunks extended.

test_mesh_lawson_stack: a case passing a non-empty FlipStack on entry,
pinning lawson.hpp's "contents on entry are discarded" and the empty
stack on return. Kills seeds-below-stale; seeds-above-valid-stale is
observationally equivalent under Lawson's post-condition, and the
comment says so. The insertion step is factored into a helper both
cases share.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… at 8 threads vs 227522c, mesh identical

Back-to-back tools/bench.py runs of master 227522c (worktree, --tree) and
21a's HEAD cade752, twice each, all on AC: threads 1..20, 5 repeats, tile
and quarter circle, tolerance 1. Ceiling 2.07-2.14x -> 2.28-2.29x. Mesh
sha256, worst angle, max degree, tolerance and Delaunay checks identical.
Median-of-5 spread between runs: max 4.6 % (base), 0.85 % (21a).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-power, not battery only

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@skavhaug
skavhaug merged commit 071e4f3 into master Sep 27, 2026
8 checks passed
@skavhaug
skavhaug deleted the increment21a-quick-wins branch September 29, 2026 21:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant