Skip to content

fix(coupling): a report describes the step that ran; tests for verdicts over a span; diagnostics cost docs - #256

Merged
NicholasEhsanRoy merged 8 commits into
release/0.4.0from
fix/p4-40-coupling-r8
Oct 7, 2026
Merged

NicholasEhsanRoy merged 8 commits into
release/0.4.0from
fix/p4-40-coupling-r8

Conversation

@NicholasEhsanRoy

@NicholasEhsanRoy NicholasEhsanRoy commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

The coupling batch after audit round 8: one library fix (item 2), three test gaps closed (item 1 and its sweep), one item not reproducible without a GPU (item 3, registered), one docs paragraph (item 4).

Item 2: a report describes the step that ran

Cause. coupling_diagnostics() took the _meta slots from the step but measured the residual's float floor on the node states the graph held at report time. After set_node_state the same step read another spectral_error_bound / precision_limited / spectral_usable: 0.0, flag set, where every member was written to zero (reproduced on the base: every stepping entry point, also after load_state and after a recompile).

Choice: (a), the report is computed from what the step left. GraphManager._keep_state_for_reports is called by each door that writes node states in place (set_node_state, so PUT /graph/state/{node} and load_state; add_node; remove_node) before its write. The first write after a step keeps a shallow copy of the state dict (references, no array data) beside the report slots it goes with. The reader (_state_a_report_describes) uses the copy only while the group's iterations and residual slots are still the very objects kept with it; otherwise it reads the live state. So a door that brings its own slots (a step, a loaded checkpoint, reset_state, a recompile that restarts a group) needs no bookkeeping and cannot be described by a stale copy; a missed write door degrades to the old behaviour, never to a report measured on an older step's state. _store_state drops the copy (untraced), so the replaced arrays are referenced only until the next step. No _meta slot, nothing in the compiled step.

Why not (b): every write door would have to mark the report stale, a missed door would show a moved report as fresh, and users would lose a report that the graph can still give exactly.

Checkpoints: a checkpoint is a copy of the state, _meta included. (The first version of this PR saved a written group's iterations counter as 0; the REST write-sequence oracle caught that the archive then no longer reproduced the live _meta. Redone.)

  • The state and _meta are saved exactly as they are, always.
  • The fact "this group's state was written after its last step" is carried beside the state: an optional archive member _reports/<group>/written_after_step, present only for a group whose members were written to other values since its step (or that was itself loaded so marked and has not stepped). Every other archive has the members it had before. _reports joins the reserved node names.
  • What a report reads after load_state:
    • checkpoint saved straight after a step (no member): the saved step's report, key for key; _meta identical; it survives a later write; further steps and reports bit-identical.
    • checkpoint saved after a member was written (member present): state and _meta identical to the saver's; the group's entry reads spectral_error_bound NaN, spectral_usable, gradient_bound_usable and precision_limited False, and a not_usable_reason saying the checkpoint was saved after the state was written; every other key is the slots' own. It stays so under further writes and is saved marked again; the group's next step brings a whole report. coupling_report() prints "no bound reported: ..." and keeps the estimate's caveats.
    • an archive without the member (an older build's, or the member stripped): loads as before, report measured on the loaded state. A member for a group the graph does not have is ignored. A load that fails leaves no mark.
  • A group that stores its floor in the step (CPL-188) is never marked.
  • The live graph is unchanged: the pre-write copy.

Paths newly enabled or changed. set_node_state, add_node, remove_node keep a dict copy when the graph has a compiled coupling group and holds no tracers; checkpoint.save_state adds one archive member in the case above and load_state reads it; profiler.compile_counts and _one_iteration_variant hand the kept copy back with the state they restore. Nothing under a trace is kept (tested: a traced write, and a transform that stepped the graph).

Not covered: an assignment into the private gm._state dict.

Report-time reads (sibling sweep)

Entry Computed from Moves under a later write
iterations, total_iterations, residual, amplification, error_estimate, ratio_usable, gradient_error_estimate, converged the step's _meta slots; the criterion of the group committed at compile() no
rho_spectral, gradient_relative_error_bound slots no
spectral_error_bound, precision_limited, spectral_usable, gradient_bound_usable slots and the float floor did (floor from the live state); now the kept state
the floor of a group whose interface norm reads a mapped edge (CPL-188) the reading_floor slot no
the floor's mappings argument live params["mappings"], only where that slot is absent not after a step of this build (probed on the mapped seed cell: a mappings write does not move the report)
evaluation count, declared flag, internal edges _committed_floor_inputs (compile snapshot) no
whether there is an entry member present in the live state; iterations > 0 by design (a removed member or a reset ends the report)
coupling_report() / print_coupling_report() coupling_diagnostics() only follows it; no new reason key, printers unchanged
profiler coupling statistics, calibration _meta read right after their own steps no write in between

Tests: tests/core/test_coupling_report_describes_the_step_that_ran.py (27: six write shapes, six stepping entry points, next step, load, a marked checkpoint, an unmarked archive, the reserved prefix, failed load, recompile, reset/removal, tracers twice, profiler, references), tests/api/test_a_coupling_report_survives_a_rest_state_write.py (2), and the call sequence step / write / read as a seed shape of the coupling targeted search (test_a_report_on_a_seed_shape_does_not_move_when_the_state_is_written_afterwards, 7 seeds, float32 and float64, mapped, hub; 6 of 7 fail on the base). 17 of the 25 fail on the base.

Mutants (scratch copy, plans/tools/mutants.py): 14 seeded, 14 caught on the live-graph mechanism, and 7 of 7 on the checkpoint marker (one after strengthening a test: a write inside a transform must not replace the copy kept before it).

Item 1: verdicts taken over a span (tests; the library is right)

tests/core/test_sysid_mask_reads_every_step_of_the_window.py: a float32 linear Gauss-Seidel pair (0.81 per pass, cap 4, tolerance 1e-6) restarted far from its fixed point; early steps exit at the cap, the last converges. Loss and gradient exactly zero under the mask; the same window kept (equal to the unmasked loss and gradient) when the cap lets every step converge; three (window, sample_every) shapes so the recovery crosses each of the two scans; under multiple shooting the recovering window and its continuity term go and the converged window stays, equal to that window alone. CPL-160 and CPL-161 cite them.

Site Reduction Swap mutant
sysid mask, windowed_loss._simulate.inner and over every base step last step only: survived every test before, caught now
strict_convergence in step / run / run_scan none: each step's in-graph check raises; the firing gate (_gate_on_firing) is reachable only under vmap, since unbatched the cond does not run a discarded solve gate dropped, inverted, silenced: 3 caught by two new direct tests (before: one incidental batched test)
run_adaptive, host (_raise_if_a_kept_solve_failed) any of the two kept half steps last half only: caught by an existing test
run_adaptive_scan, in graph or of the two halves last half only: caught (existing)
_fold_kept_half_step_reports the first half where it alone failed never the first: caught (existing)
profiler converged_fraction, at_cap_fraction mean over the window last step only: both survived (every window in the tests was uniform); caught by a new test on a recovering window
waveform sweeps the report is the last sweep's, strict raises about any (documented in CPL-160's conditions) existing test, not re-mutated
POST /sim/run reply carries no verdict; a strict raise is a 400 with steps_run n/a

9 mutants, 9 caught.

Item 3: NaN on GPU, a number on CPU (not fixed; MADD-ANO-232, open, minor)

Not reproduced on CPU. Read: every quotient in _gradient_error_bound_body and ift_gradient_error_bound is guarded by a select; the documented NaN paths are "nothing responds", "range not captured" and a non-finite tangent. On the stock pair (four state entries, a square range basis) captured holds with or without one-ulp noise on the Jacobian products; forcing it False gives NaN with the flag False, the GPU's reading. Which intermediate differs on the GPU needs a GPU run with the bound's parts printed. A report-time rewrite of not-usable values was not made: the three kinds (finite, inf, NaN) are documented by cause and pinned by tests, and spectral_error_bound keeps its not-usable value too. The user guide now says not to compare not-usable values between backends.

Item 4: docs

docs/user_guide/inspection.md: what diagnostics=True costs (the product counts of CPL-013, the dense factorisations and the eigenvalue solve, the GPU measurement), that a report describes the step that ran, and the not-usable value above. The measurement is committed as benchmarks/results/gpu_eigvals_probe/ (the maintainer's local run; no GPU was used here).

Records

MADD-ANO-231 (resolved, major, never released), MADD-ANO-232 (open, minor). CHANGELOG: one line under Fixed, one under Verification. Claims: CPL-160, 161 (new tests and domain cells), CPL-088, 092, 097 (the write test), CPL-188 (conditions), new CPL-189.

Checks

  • scripts/capture_step_programs.py --check, fresh process each: jax 0.10.2, 0.11.0, 0.11.2: 24 graphs unchanged.
  • scripts/compile_counts.py --check on jax 0.10.2: counts match.
  • All scripts/check_*.py; the changed and added test files; the three source scans; the registry, release-notes, claims, guard-mutation and slow-only compliance tests.
  • After the checkpoint rework, on the merged tree: the slow REST machine test_any_sequence_of_rest_writes_leaves_a_graph_that_runs_as_its_reload at the ci profile, once: passed (21 min, three cores, jax 0.11.0); the per-push lane of that file, of 16 checkpoint test files, of the step-layout guard together with this PR's tests; digests and compile counts again unchanged.
  • pyright was not run locally (not installed in the shared environment); CI runs it.
  • Slow tests run once locally (-m slow, three cores, jax 0.11.0), in every file that reads a coupling report or profiles a graph and also writes or checkpoints state: tests/core/test_compile_counts.py (2), test_coupling_claims_in_every_domain.py (28), test_coupling_groups_in_every_domain.py (1), test_coupling_interface_reading_is_what_the_edge_delivers.py (2), test_coupling_spectral_bound_on_the_transformed_reading.py (32), tests/property/test_coupling_convergence_at_every_field_scale.py (1), test_differential_coupling_solvers.py (8), test_differential_coupling_topologies.py (259), test_differential_fixed_point.py (8), test_differential_geometry_edges.py (127), test_coupling_targeted_search.py (28), the five coupling examples of tests/test_examples_smoke.py: all pass. tests/property/test_coupling_error_bound.py: 3 pass, 1 fails, and fails the same way on the base (below). Not run: the rest of tests/test_examples_smoke.py (it covers the cloud examples' dry-run modes).

The slow normal-contraction property (red on a draw on the base; settled here)

tests/property/test_coupling_error_bound.py::test_the_spectral_bound_holds_on_random_normal_contractions asserted spectral_usable is True on every drawn normal contraction, and one random draw read False on the base (1f66801c) and on this branch alike.

The draw. float32, n = 6, eigenvalues 0 and 0.5 five times, acceleration "none"; 12 iterations, residual 2.6e-4, rho_spectral 0.5000001 (right), spectral_error_bound 6.03e-4 against a true distance of 5.21e-4 (holds).

Why the flag is False. Not the componentwise certificate: the "Krylov space still growing at the cap is never settled" rule of PR #252. The start vector's Krylov space closes after two steps (one eigenspace). The estimator continues from the residual, whose part outside that eigenspace is the float32 iterate's rounding: 1e-4 of the residual, far above the breakdown test's eight ulps of a product, so it is carried as a direction. Each later step grows by 1e-5 to 5e-4 and the space is still growing at step eight. The stored spectral_residual reads 0.02656, which is exactly 1.0625 x the margin (0.05 x (1 - rho)), the value that rule writes.

What the claims say: outcome (b). No row promises the flag on normal contractions. CPL-088 claims the bound where spectral_usable is True; CPL-092 says the flag is False "for a group whose Krylov space is still growing at the cap"; CPL-087 claims the radius where the flag is set and no more than eight scalars cross the group's edges (twelve cross here). The assertion was the test's own, written before that rule had no exceptions. No claims row changed; the certificate was not touched.

The test now (same name, still slow, same per-push witnesses): the draw is pinned as an explicit example and must read not usable; where the flag is set the radius and the bound are asserted as before; where it is not, the stored Arnoldi residual must be over the settle margin (withdrawn for its documented reason); and the flag must be set on at least 0.9 of the drawn examples. Measured over the test's own draws: 300 of 300 on fifteen seeded runs of twenty, plus the one pinned draw from one unseeded run. Run once to green in its slow profile. 3 mutants, 3 caught (the flag withdrawn everywhere; the bound without its amplification; the growing-at-the-cap rule removed, which sets the flag on the pinned draw).

🤖 Generated with Claude Code

https://claude.ai/code/session_013UkCde7g23gTziUjYAvnKD

NicholasEhsanRoy and others added 4 commits October 7, 2026 04:13
…hed strict gate each read every step

The mask and-s every base step's verdict over a window; no test had a window
that recovers (early steps at the cap, the last one converged), so a mask
reading the last step alone passed.  Same gap for the profiler's two
fractions (every window was uniform) and for the strict check's firing gate,
reachable only under vmap and held by one incidental test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UkCde7g23gTziUjYAvnKD
… state is written afterwards

coupling_diagnostics() measured the residual's float floor on the live state
at report time, so set_node_state after a step gave the same step another
spectral_error_bound / precision_limited / spectral_usable (a bound of 0.0
with its flag set once every member was written to zero).

The graph keeps a shallow copy of what the step left from the first later
write, beside the report slots it goes with, and the report is measured on it
while those slots are still the ones in the graph.  A checkpoint saved after a
member was written carries no report for that group.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UkCde7g23gTziUjYAvnKD
…MADD-ANO-231 and 232

The inspection guide states the diagnostics' per-step work and the one GPU
measurement of it; the registry records the report that moved under a later
state write (resolved, never released) and a not-usable gradient bound that
read NaN on a GPU backend and a number on CPU (open: cause not established
without a GPU).  Claims CPL-160/161 cite the recovering-window tests and
CPL-189 holds the new report claim.

[skip ci]

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UkCde7g23gTziUjYAvnKD
The previous commit of this branch was pushed as work in progress with CI
skipped; nothing has changed since.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UkCde7g23gTziUjYAvnKD
@NicholasEhsanRoy
NicholasEhsanRoy marked this pull request as ready for review October 7, 2026 03:41
NicholasEhsanRoy and others added 4 commits October 7, 2026 06:05
…d where the flag is set

It asserted spectral_usable is True on every drawn normal contraction.  No
claim promises that: the bound is claimed where the flag is set, and the flag
is False for a Krylov space still growing at the cap.  A drawn 6 x 6 matrix
with five equal eigenvalues does that in float32 (the space continues from
the residual's rounding), so the slow lane could go red on a draw.  The draw
is pinned, the bound and the radius are asserted where usable, a withdrawn
flag is checked against its stored residual, and the share of draws the flag
is set on is held to a measured floor.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UkCde7g23gTziUjYAvnKD
) into fix/p4-40-coupling-r8

# Conflicts:
#	CHANGELOG.md
#	docs/validation/soup_package.md
…r the step is carried beside it

The previous commit saved a written group's iterations counter as 0, so the
archive no longer reproduced the live _meta (the REST write-sequence oracle
caught it).  The state and _meta are saved as they are again.  Where a
group's members were written after its last step the archive carries the
optional member _reports/<group>/written_after_step, and the graph that
loads it reports that group's bound as NaN and its *_usable flags and
precision_limited as False, with a not_usable_reason, until the group steps.
An archive without the member loads as before; _reports is a reserved node
name.

[skip ci]

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UkCde7g23gTziUjYAvnKD
The previous commit was pushed as work in progress with CI skipped; nothing
has changed since.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UkCde7g23gTziUjYAvnKD
@NicholasEhsanRoy
NicholasEhsanRoy merged commit 8828874 into release/0.4.0 Oct 7, 2026
25 of 27 checks passed
NicholasEhsanRoy added a commit that referenced this pull request Oct 7, 2026
) into fix/p4-41-iqn-nonfinite-and-gradient-bound

# Conflicts:
#	docs/release_notes/v0.4.0.md
#	docs/validation/known_anomalies.yaml
#	docs/validation/soup_package.md
#	tests/compliance/test_soup_evidence.py
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant