Skip to content

Increment 21b: exact integer incircle for DEM-node quads - #103

Merged
skavhaug merged 12 commits into
masterfrom
increment21b-lattice-incircle
Sep 27, 2026
Merged

skavhaug merged 12 commits into
masterfrom
increment21b-lattice-incircle

Conversation

@skavhaug

Copy link
Copy Markdown
Member

What

Increment 21b (docs/increments/21-parallel-refine.md, §3 QW2). Bit-identical output: T6 variants and 18's golden digests untouched; bench mesh sha256 equal to master on both domains.

  • Premise measured first (docs/benchmarks/2026-09-27/21b-ties/): every exact-path incircle tie is a four-node lattice-cocircular quad under QW2's conditions (146,962/146,962 quarter, 154,502/154,502 tile); QW2 answers 99.97-100 % of refine's incircle calls.
  • mesh::lattice_frame / lattice_incircle: int64 determinant on (col, -row) differences from d, only on an enabling frame (dx == dy, exact node coordinates, finite extent), four node corners, spread ≤ 2^14; refuses non-finite corners via the spread test. must_flip asks it first. refine builds the frame. 55 production lines (est. ~50).

Red 042d3e4 before green f707322; test amendment 24c2801 (both sides of the exact-frame bound, overflowing extent, non-finite corners); a03ae65 + 4be156c the non-finite refusal, folded to zero cost. Invariant-critical suite test_mesh_lattice_incircle: 18/18 mutants plus the reviewer's two.

Acceptance (@Perf, battery, back to back with master 071e4f3, run twice)

refine, median of 5 master 21b
tile, 8 threads 226.4 ms 193.0 ms (-14.8 %)
quarter, 8 threads 200.7 ms 171.1 ms (-14.8 %)
ceiling (best) 2.3x 2.5x

Measured on a03ae65; the addendum shows 4be156c recovers its 4 % over that. Evidence: docs/benchmarks/2026-09-27/21b-acceptance.md. Verdict ACCEPTED.

Review

@Reviewer: CHANGES REQUESTED twice (a suite gap at the exact-frame bound, then prose), APPROVED at 7f688aa. Local: ctest 781/781 (Release, ASan/UBSan, -ffp-contract=off), TSan clean, pytest 1719 passed.

🤖 Generated with Claude Code

skavhaug and others added 12 commits September 27, 2026 13:30
…ircular under QW2's conditions; QW2 admits 99.97 % of all incircle calls

Measurement only; no production code. Instrumented d3ee2ce in a scratch
worktree (scripts/instrument.patch, kept unapplied), 1 m benchmark, quarter
circle and tile, tolerance 1, AC. Refine loop: 146,962 / 154,502 exact-path
calls, every one four node corners, int64 lattice det on (col, -row) = 0,
dx == dy, spread <= 2^14, frame exact. Only 38 % are axis-aligned
rectangles. Of all refine-loop incircle calls, 99.973 % (quarter) and
100 % (tile) meet QW2's conditions; lattice sign agreed with the kernel on
every four-node call. Output identical to 21a's acceptance meshes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
21b is bit-identical, so 14's T6 and 18's golden digests stay untouched as
the output oracle. The new suite test_mesh_lattice_incircle is 21b's
invariant-critical suite (section 7). It pins the interface QW2 left open; the
pins are under "Pinned by the red suite (21b)" in 21-parallel-refine.md:
lattice_frame(dx, dy, rows, cols) decides once per refine call whether the
integer path may answer, and lattice_incircle(a, b, c, d, frame) returns an
optional Incircle. A directly built LatticeFrame{dx, dy} never enables the
path.

Agreement with DetriaExact::incircle_ccw on the frame doubles, for every quad
it answers: an exhaustive 5 x 5 neighbourhood at the origin and at the
benchmark's far corner in four exact frames; the tie shapes 21b-ties lists, in
every rotation; every quad on the radius-5, radius-sqrt(65) and radius-8085
lattice circles, and those quads with d moved one node; random uniform and
near-circle quads up to the spread bound. It answers at a spread of exactly
2^14 and refuses at 2^14 + 1, for each corner, axis and direction. A named
test fails on a (col, row) determinant. It refuses an off-node corner in each
position; dx != dy; dx that is zero, negative, infinite or NaN; an inexact
frame (0.1, with rows and cols tested separately, and a 53-bit dx); and a
direct frame. It must answer where QW2's sufficient condition holds, including
exactly 53 bits. A search finds a lattice tie that the 0.1 frame breaks, which
shows why that refusal is needed.

must_flip against a copy of today's version, on random lattice meshes with
ties and off-node feet, in five frames: the same decisions edge by edge, and
the same legalise_all flips, writes and mesh. A counting kernel shows that
must_flip asks the kernel nothing when the integer path answers.

Mutation round against a scratch implementation that is not committed. All 18
mutants were killed: sign inverted; (col, row); >= at the bound; a bound of
2^14 + 1; a bound of 2^15; the bound checked on one side only; the bound
checked on columns only; no exact-frame check; the exact-frame test at < 53;
the exact-frame test on columns only; on rows only; no node check; a node
check that skips d; dx != dy accepted; dx <= 0 accepted; a direct frame
enabling the path; no call in must_flip; a determinant in doubles. The
scratch implementation also passes test_mesh_lawson, test_mesh_lawson_stack
and test_mesh_quality.

The suite is registered only once lawson.hpp names lattice_incircle, as in
21a, and the header is a configure dependency. That the guard fires was
checked by building with the scratch header in place (the suite built and
passed), then restoring the header. The suite starts no threads, so it is not
in the TSan job.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
QW2 (docs/increments/21-parallel-refine.md, section 3, and "Pinned by the red
suite (21b)"), in lawson.hpp:

- LatticeFrame becomes a class with a private flag that only lattice_frame()
  sets, so LatticeFrame{dx, dy} keeps compiling and never enables the integer
  path. A constructor rather than a third aggregate member keeps GCC's
  -Wmissing-field-initializers out of every existing two-value brace init.
- lattice_frame(dx, dy, rows, cols) enables it when dx == dy, dx is a
  positive normal double, the significant bits of dx plus
  bit_width(max(rows, cols) - 1) are at most 53 (QW2's sufficient condition,
  used as the test), and the largest col * dx is finite.
- lattice_incircle(a, b, c, d, f): four node corners, every difference from d
  at most 2^14 nodes, then the int64 incircle determinant on (col, -row)
  translated to d (below 3 * 2^58).
- must_flip calls it first; when it answers, the answer is used and the kernel
  is not asked, not even the frame orient2d.

refine builds its frame with lattice_frame(g.delta_x(), g.delta_y(), g.rows(),
g.cols()). No test file and not the CMake guard are touched.

Bit-identical: T6 and test_refine_golden.py pass unchanged, and the 1 m
benchmark's binary meshes hash as 21a's. Timed back to back against d3ee2ce on
battery: split phase -26 to -27 %, refine -7 to -8 % at 1 thread and -17 % at
8 (recorded in the 21 file; @Perf's AC acceptance is still owed).

54 production lines added, 3 removed (CLAUDE.md section 2's unit), against
an estimate of ~50.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…, the 10x10 tie

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…on-finite corners

Test amendment after green, in its own commit as docs/increments/README.md
allows; no production code. @Reviewer found two surviving mutants in
lattice_frame and one unguarded input in lattice_incircle:

- The limit taken one bit loose (bit_width(top - 1) - 1) survived: the 0.1
  refusals sat only at 7-bit spans. Now pinned on both sides: 0.1 and
  1.5 + 2^-51 (52 bits, 3 * dx rounds) must refuse on 4x4, 4x2 and 2x4, and
  must answer on 2x2; 1.5 + 2^-50 (51 bits) must answer on 4x4. A sweep of
  every ordered quad of the 4x4 grid requires any answer to match
  DetriaExact on the 0.1 frame, and the reviewer's quad (a lattice tie the
  0.1 frame calls Inside) is pinned by name.
- Dropping the isfinite extent clause survived. 2^1023 on a 3x3 grid (1 + 2
  bits, extent inf) must refuse; on 2x2 it must answer.
- is_node() is true for +-inf, and inf - inf = NaN passes `fabs > bound` into
  the int64 cast. lattice_incircle must refuse non-finite corners. This case
  is RED at this commit: with a, b, c and d on the same infinity it returns a
  value (4 failing assertions); the developer adds the check next.

Also fixes the stale "as every caller builds it today" comment, and records
the finite-corner and extent pins and the two mutants in the increment file.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…el preconditions stated

is_node() is true for +-inf, and a corner and d on the same infinity give
inf - inf = NaN, which passed the spread bound and reached the int64 cast
(UB; UBSan fired on [nonfinite]). Every corner is now checked finite before
any difference is taken.

Comments only, no behaviour change: the DetriaExact-sign claim holds for
nodes inside the rows x cols grid the frame was built for; must_flip answering
without K assumes K's incircle sign is exact.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…identical

Back to back against master 071e4f3 (21a merge), two pairs, bench.py
blob 77765b1, threads 1..20, 5 repeats, quarter and tile, 1 m DEM,
tolerance 1. Split phase -21..-24 %, scan unchanged; ceiling 2.45-2.55x.
Head's split phase is 4-6 % slower than green f707322 (non-finite
refusal in lattice_incircle, attributed by elimination).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
a03ae65 checked std::isfinite on all eight coordinates before is_node(),
and the split phase got 4-6 % slower than f707322 (21b-acceptance.md,
"Green against head"). The explicit checks are gone. The spread test is
now written as !(fabs <= bound), which is true for NaN as well as for a
large value. NaN is not a node. Every coordinate of a, b, c and d enters
some difference from d, and a difference with an infinite operand is
either +-inf or NaN (inf - inf). So every non-finite corner is still
refused before its int64 cast. Against f707322 the only code change is
that comparison. The [nonfinite] test passes, and UBSan is quiet on it.
With the old `fabs > bound` put back, UBSan fires at the cast
(lawson.hpp:111).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nce run pointed to; TSan rule for threaded suites

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@skavhaug
skavhaug merged commit 4a0ba35 into master Sep 27, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant