Skip to content

fix(sysid): fit_lm reads a small jump of the residual from a long rejected step - #258

Merged
NicholasEhsanRoy merged 3 commits into
release/0.4.0from
fix/p4-42-sysid-jump-residual
Oct 7, 2026
Merged

NicholasEhsanRoy merged 3 commits into
release/0.4.0from
fix/p4-42-sysid-jump-residual

Conversation

@NicholasEhsanRoy

@NicholasEhsanRoy NicholasEhsanRoy commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

What was wrong

tests/property/test_sysid_targeted_search.py::test_no_wrong_fit_found_by_the_search[converged-beside-a-lower-loss] went red in the slow lane on jaxlib 0.10.2: fit_lm on the stock bouncing ball returned converged=True at a loss of 6.0e-3 where the loss still falls along the piece the point is on (score 11.8 against a threshold of 1).

Reproduced on both jax versions, identical to the last digit (jaxlib 0.10.2 and 0.11.0; BallCase(0.6133024381551229, 0.7699776319888987, 0.5, 1.0)): converged=True, loss 0.006037198938429356, 18 iterations, score 11.815881429033254. It is not a 0.10.2 effect; the hunt is random and no PR's CI runs it.

Classification: (a), a jump the detection missed. At the last iteration:

  • the floor rule declared convergence (not the proposal test, not tol); the conditional retry from lam0 did not run (_gain_in_the_gap false);
  • three rejected candidates were read by the one-sided test:
candidate's move departure ahead at the mirror image linear change ratio
5.9e-6 (35 float spacings of the elasticity) 1.47e-2 7.3e-6 2.5e-5 456
5.7e-7 7.5e-6 3.0e-6 2.6e-6 1.3
6.3e-8 2.5e-7 3.1e-7 2.5e-7 0.44
  • a line through the returned point, one float spacing at a time: in the elasticity there is one jump, 22 float spacings below the iterate (relative 1.3e-6), where the loss rises by 2.0e-4; in gravity the loss falls smoothly, by 8.76e-6 at 1e-4 (the gradient there, 0.0119, times the step). So the score's statement is true: the loss falls along the piece, and the Levenberg-Marquardt steps cannot follow it because they cross the edge.

The ratio was under 2**10 because the candidate was too long for a small jump: the jump (0.0147 in the residual's norm, a late bounce that changes 19 of 200 samples) is divided by the linearisation's change over the whole candidate, which grows with its length. The next candidate down the ladder (4 spacings) did not reach the edge. PR #245 measured 2.2e4 to 4.4e5 on 55 jump endings; over 4,500 drawn fits the first reading is 1.9e3 to 5.8e5 at 261 jump endings, and 341 at one more (also converged=True before, score 5.6).

The fix

A rejected candidate whose ratio is above 2**5 and not above 2**10 is read a second time (_excess_across_a_spacing): its step is halved, keeping the half over which the residual departs further from the linearisation, down to two neighbouring points of the coordinates' grid; the departure across that one spacing is compared with the same mirror-image departure of the whole candidate plus the linearisation's change over one spacing. Above 2**8 it is a jump: converged=False with the existing warning (which prints the reading that decided). A candidate at or under 2**5, or over 2**10, is read exactly as before. Cost: one residual evaluation a halving (7 on the example), at the end of the run, only for such a candidate.

Measured separation (CPU; jaxlib 0.10.2 and 0.11.0 give the same numbers on the ball):

endings read by the floor rule first reading second reading
smooth residuals: every ending existing per-push and slow sysid tests reach (float32 105, x64 37 incl. two of a new unit test) 140 (1,186 candidates) at most 5.7, but for the wide-logit fit at 21 (unchanged from #245) not asked (all under 2**5); computed anyway: at most 13, the wide logit 52
ball, at a piece's own minimum 6 of 4,500 draws at most 3.2 at most 5.2
ball, the two small jumps the hunt's example, and 1 of 4,500 draws 456, 341 1.95e3, 1.77e3
ball, the other jump endings 261 of 4,500 draws 1.9e3 to 5.8e5 (decided by the first) 3.4e4 or more

2**8 is 4.9 times above the largest smooth second reading and 6.9 times below the smallest jump's; 2**5 is 1.5 times above the largest smooth first reading (which only decides whether a second reading is taken) and 10 times below the smallest jump's.

A variant I measured and rejected: taking the mirror image across one spacing too (two neighbouring points against their own mirror images) reads the example at 3.0e4, but is no measure of rounding: neighbouring residuals often agree exactly, and smooth floors read 4.1e3 (wide logit), 2.5e4 and 1.0e6 (a noiseless x64 fit at its floor). Mirroring through the nearer end of the pair fails on the ball itself: at the scale of one spacing the two sides of the edge interleave.

Proof

  • Deterministic per-push tests (hold on both jax versions; run on both):
    • test_fit_lm_on_the_ball_is_not_converged_where_the_loss_falls_along_its_piece[cell2, cell3]: the hunt's example and the drawn one. Fail on the base (origin/release/0.4.0 src, both jax versions), pass now.
    • test_fit_lm_reads_a_small_jump_of_the_ball_from_a_long_candidate[cell0, cell1]: converged=False, one warning naming the second reading; with the second reading switched off, converged=True at the same point with the score over its threshold; a second reading between 2**8 and 2**10 ends the run the same way.
    • Unit tests of the reading in tests/core/test_sysid_non_differentiable_residual.py: a small jump beside the iterate, inside the step and at its far end; a jump reported beside a larger smooth reading, in either order; a jump met by one on the other side; a smooth residual whose slope grows by exp(6.5) over the candidate (second reading under 1); nothing read twice outside the two thresholds; a residual that is not finite inside the step.
  • The slow hunt of this score, seeded, 800 examples a seed: seeds 1 to 8 on jaxlib 0.10.2 and on 0.11.0: nothing over the threshold (16 hunts). And the unseeded slow test once on 0.11.0: passed.
  • 4,500 uniform draws of the same ranges on each jax version: 0 over the threshold (1 before, plus the hunt's).
  • Honest fits: 0 flipped. Every floor-rule ending of the per-push lane of the 23 test files that call fit_lm and of the slow tests below, with the rule instrumented: 250 endings in 47 existing tests (140 of smooth fits in 39 tests, 110 of the ball), verdict changed in 0. No smooth ending is read a second time.
  • Mutants (plans/tools/mutants.py, scratch copy of src/): 19 seeded, 17 caught, 2 documented equivalents (stationary = True on the verdict line, as in fix(sysid): fit_lm tells a jump of the residual from the rounding floor, and warns when excited_rank reads full on a degenerate fit #245: the warning's break ends the run either way; dropping the finiteness test of the second reading: a NaN compares false against the threshold).
mutant caught by
_JUMP_SUSPECT to 2**10, to 0; the second reading never taken the ball pin; ..._is_read_once, the smooth fit's max(read) < _JUMP_SUSPECT
_JUMP_ACROSS to 2**11, to 2**5 the ball pin; ..._met_by_one_on_the_other_side...
one halving only; the wrong half kept; halving from the iterate the ball pin, ..._read_across_one_spacing[1.00019]
the scale without the mirror image, or with 0 passed for it ..._the_one_reported, ..._met_by_one...
the linear change over the whole step; the departure from the iterate the ball pin; ..._the_one_reported
a second reading over its threshold not a jump; the first reading reported for it; the larger number reported instead of the jump the ball pin; ..._read_across_one_spacing; ..._the_one_reported
the first reading never a jump test_the_one_sided_excess_of_a_jump...
the warning put to the first threshold the ball pin (a second reading of 2**9)

Code paths this enables or changes

  • fit_lm, floor rule, both branches (the damped candidates and the undamped Gauss-Newton candidate are in the same list): a candidate between the thresholds now costs up to 64 residual evaluations (measured 7 to 21) and can end the run converged=False with the warning. Nothing else in the loop changes; no compiled program changes (scripts/compile_counts.py --check on jax 0.10.2: counts match).
  • _one_sided_excess returns (excess, move, jumped) (private; was a pair); fit_lm reads the verdict from it.

Siblings checked (every place that declares convergence, as in #245): the floor rule (fixed, both branches); _gain_in_the_gap (the other user of 2**10: a promised reduction against the loss's rounding, with no step length in the denominator, so nothing dilutes it; unchanged); the proposal test and tol (no rejected candidates to read; documented as not detecting a jump); fit and fit_multiple_shooting (tol only). Still not detected, as documented: a kink, a jump met by its like on the other side, a jump whose mirror image the bounds do not allow, an iterate the proposal test passes, a jump under 2**5 times the linear change of the step that crossed it.

Records

  • MADD-ANO-216 (never released, resolved): the description gains the residual case with its numbers, the residual risk states both thresholds, three tests added. No new ID: it is the same defect in the same unreleased rule (so no release-notes line and no change of the high-water mark).
  • SYS-080: the claim quotes the second reading, the condition says "over its whole length or across one float spacing inside it", seven tests added. SYS-086 unchanged. No guard-mutant anchor cites a sysid row.
  • CHANGELOG: one line under Fixed.

Local runs

Gates (scripts/check_*.py, all nine) and generate_soup_tables.py --check: pass. Per push, jax 0.11.0: the two changed test files, the three source scans, the soup, release-notes, claims and sysid-docs compliance files: 469 passed. jax 0.10.2: the two changed test files: 35 passed. pyright on sysid.py: no new error (one in _ExcitationTracker.projector on this box's numpy stubs, the same on the base).

Slow tests run once (jax 0.11.0, cores 12-14; 69 passed in 41 min): every slow test of tests/property/test_sysid_targeted_search.py (the four hunts), test_sysid_truth_recovery.py (41), test_sysid_one_evaluation.py, test_sysid_contract.py (14), test_differential_precision.py (5), tests/core/test_sysid.py, test_sysid_claims_edges.py, test_cross_feature_edge_cases.py (2).

Also here: a property-test strategy that drew a spec the library refuses

tests/property/test_sysid_contract.py::spec_and_value drew ParamSpec(bounds=(1.0027701900272462e-38, 1.0), transform='logit'): a subnormal float32 lower bound, which ParamSpec has refused by name for a float32 leaf since PR #249. The strategy predates the refusal, so test_constrain_inverts_unconstrain_inside_the_bounds and test_constrain_lands_inside_the_bounds_from_any_coordinate (the two users of the strategy) failed at random: 2 of 3 per-push runs here, the same against the base's src/.

  • The strategy now excludes exactly one thing: a logit bound that is a subnormal float32 (non-zero and below 1.17549e-38 once rounded to float32). Such a draw is moved to 0.0, the bound the refusal's message recommends, not rejected. Left as drawn: zero, the smallest normal, bounds no float32 holds exactly, and subnormal bounds under transform=None and log (legal; the existing @examples with 1.4e-45 stay). The upper bound lo + span cannot be subnormal (span >= 1e-2). No library change, no assertion loosened.
  • The refusal is held by tests/core/test_constrain_lands_inside_bounds_near_the_smallest_normal.py::test_a_logit_interval_narrower_than_four_smallest_normals_is_refused_by_name, which already had the shape (tiny / 2, 1.0); the drawn spec is now an explicit case there (float32: to_constrained, to_unconstrained and check refuse it by name; x64: an ordinary bound, mapped).
  • test_the_spec_strategy_moves_only_a_subnormal_logit_bound pins what is moved and what is not, and that the drawn spec is refused where its neighbour is mapped. Both properties gain @examples on the accepted side of the boundary ((0.0, 1.0) and (smallest normal, 1.0)).
  • Siblings: the only other strategy that draws logit bounds near zero, in test_differential_param_acceptance.py, treats a refusal as an outcome it compares across doors; the slow fit test of the class draws bounds of 0.5 and up.
  • Run: TestBoundsAndTransforms per push with --hypothesis-seed 1 to 6 on jax 0.10.2 and on 0.11.0 (12 runs, all pass, including the replay of the stored failing draw); its slow test once (pass); the refusal's file (pass).

jaxlib 0.11.2

CI's second lane is 0.11.2 (the shared venv's 0.11.0 is a third version). There the ball's trajectory rounds differently: the first pinned ending has two candidates between the thresholds, with second readings 4.4e3 and 2.1e3, where 0.10.2 and 0.11.0 have one at 2.0e3. The detection is the same; test_fit_lm_reads_a_small_jump_of_the_ball_from_a_long_candidate had counted the candidates. It now asserts what holds on every version: converged=False with one warning; the first reading alone finds no jump and its largest ratio is in (2**5, 2**10]; the verdict is one of the second readings, over 2**8, and the warning prints it. The two changed test files, the strategy's class and the refusal's file pass on 0.10.2, 0.11.0 and 0.11.2 (64 passed each); the seeded hunt on 0.11.2, seeds 1 to 3: nothing over the threshold; the ten mutants that this test catches were re-run against it: 10 caught.

🤖 Generated with Claude Code

https://claude.ai/code/session_013UkCde7g23gTziUjYAvnKD

…ected step

The floor rule's one-sided test divides a rejected candidate's departure
from the linearisation by the mirror image's departure plus the
linearisation's own change over the candidate.  That change grows with
the candidate's length and a jump does not, so a small jump (a late
bounce that moves by one time step changes few samples) read from a long
candidate fell under the 2**10 that reads a jump: two fits of the stock
ball ended converged=True at ratios of 341 and 456, where the loss still
fell along the piece (found by the slow random hunt on jaxlib 0.10.2;
the same on 0.11.0).

A candidate above 2**5 and not above 2**10 is now read a second time:
its step is halved down to the one float spacing across which the
residual departs most, and that departure is compared with the same
mirror-image departure plus the linearisation's change over one spacing,
against 2**8.  Measured: 1.8e3 and 2.0e3 at the two endings, 3.4e4 or
more at the other jump endings of 4,500 drawn fits; at most 52 on smooth
endings, none of which is asked (their first reading is at most 21).

The mirror image is kept from the whole candidate because one pair of
neighbouring points is no measure of rounding: with it taken across one
spacing too, smooth floors read 4e3 to 1e6.

MADD-ANO-216 (never released) gains the residual case; SYS-080 quotes
the second reading.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UkCde7g23gTziUjYAvnKD
@NicholasEhsanRoy
NicholasEhsanRoy force-pushed the fix/p4-42-sysid-jump-residual branch from 603e301 to 2e6f3ac Compare October 7, 2026 10:03
@NicholasEhsanRoy
NicholasEhsanRoy marked this pull request as ready for review October 7, 2026 10:04
NicholasEhsanRoy and others added 2 commits October 7, 2026 12:07
…nnot have

spec_and_value drew ParamSpec(bounds=(1.0027701900272462e-38, 1.0),
transform='logit'): a subnormal float32 lower bound, which ParamSpec
refuses by name for a float32 leaf.  The strategy predates that refusal,
so two properties of TestBoundsAndTransforms failed at random.  A
subnormal logit bound is now moved to 0.0 (the bound the message
recommends) rather than rejected; the drawn spec joins the refusal's own
test, and the accepted side of the boundary is an explicit example.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UkCde7g23gTziUjYAvnKD
…s were read twice

On jaxlib 0.11.2 the ball's trajectory rounds differently and the ending
of the first cell has two candidates between the thresholds (second
readings 4.4e3 and 2.1e3) where 0.10.2 and 0.11.0 have one (2.0e3).  The
test counted them.  It now asserts what the fix guarantees on every
version: converged=False with one warning; the first reading alone finds
no jump and is in the band read again; the verdict is a second reading
over its threshold, and the warning prints it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013UkCde7g23gTziUjYAvnKD
@NicholasEhsanRoy
NicholasEhsanRoy merged commit d508cd4 into release/0.4.0 Oct 7, 2026
15 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant