Increment 21c: measurements for choosing 21d (C, A1, A0, evaluate-once); 21d deferred - #104
Merged
Merged
Conversation
…+1.3 % Measurement only (docs/increments/21-parallel-refine.md sections 5, 6 and 7 item 2), on the 1 m benchmark at 7f688aa, battery. Scratch simulations in scripts/sim.patch (unapplied); every simulated mesh passes a full rescan and the constrained Delaunay check, and each check was shown to fail on a planted defect. - C (batch insert + parallel flip rounds) misses Ola's 2 % rule on the quarter circle: +10.0 %, +7.5 % with section 5's thinning, +6.5 % with two-hop thinning; worst angle worse in every variant. - A1 (hashed-priority reservations): +0.86 to +1.27 %, max degree no worse; worst angle better or worse depending on the hash seed. A1 gave the same mesh as the serial loop in hash order. - 7.4 % of footprints leave a 32-node halo (1.0 % at 128). - Today's index order has dependence depth 13-19 in the big rounds. - Split phase at 7f688aa: 92.6 ms at 1 thread, 96-98 ms at 8. - bench.py's tolerance check trusts refine's reported max_error and misses a stale-scan defect that a full rescan catches. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t 8 threads only C beats today clearly A0 (reservations, priority = slot index, today's skip rules, end-of-round renumbering) gives today's mesh on the quarter circle and the tile: same sha256, counts, and mesh hash after every round, 0 order violations. Its sub-rounds are close to A1's on the quarter (469 vs 447) but reach 101 in the tile's first rounds. Each mark is evaluated about 4 times, so the model puts A0 and A1 at 81-102 ms of split at 8 threads with a barrier team (today's serial split 92.6 ms), C at 50-52 ms; with fresh threads per step none of the reservation variants beats the serial split. Model, battery, labelled. Scratch only: scripts/sim_a0.patch on top of scripts/sim.patch, never in include/. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… hardening after 21d (Ola) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…led refine 7-11 % faster than today at 8 threads Simulated the rule "compute a mark's footprint once; re-evaluate only when a slot of it was written (write versions) or its node's footed status changed" in three dirty-set variants (scripts/sim_a0_eo.patch, unapplied, on top of sim.patch and sim_a0.patch). Today's `touched` set is an exact dirty set (argued, and 0 stale re-uses measured). All variants are bit-identical to today on quarter and tile (sha256, per-round lattice hash, A0's sub-round lists), with plants shown to fail. Evaluations per mark fall from 3.5-4.0 to 2.1, not 1; the model at 8 threads gives split 80 / 84 ms and refine 144 / 169 ms against today's measured 161 / 181. Scan and rest measured at 4, 8 and 16 threads. The five crash reports were the deliberate eo_stalecommit plant; the modes are clean under UBSan with libc++ hardening. Battery throughout. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…f.md's sanitizer rule fits what runs here Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e 20:57 sleep; labels, pointers and the fifth crash report Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Measurement only, no production code (
docs/benchmarks/2026-09-27/21c/README.md). Simulations are unapplied patches kept as evidence..claude/agents/perf.md: sanitize a scratch build before running it for numbers; keep deliberate crash plants out of default loops (Ola's yes). ROADMAP: a Release-hardening row (Ola's yes).docs/increments/21-parallel-refine.mdstatus line brought up to date.All figures battery; no AC baseline. Model vs measurement labelled throughout.
Review
@Reviewer: CHANGES REQUESTED once (three prose claims), APPROVED at
3c71c86. The analysis scripts reproduce their tables byte for byte; patches apply on the base they name.🤖 Generated with Claude Code