Repository navigation
fix(sysid): fit_lm reads a small jump of the residual from a long rejected step - #258
Merged
Merged
Conversation
…ected step The floor rule's one-sided test divides a rejected candidate's departure from the linearisation by the mirror image's departure plus the linearisation's own change over the candidate. That change grows with the candidate's length and a jump does not, so a small jump (a late bounce that moves by one time step changes few samples) read from a long candidate fell under the 2**10 that reads a jump: two fits of the stock ball ended converged=True at ratios of 341 and 456, where the loss still fell along the piece (found by the slow random hunt on jaxlib 0.10.2; the same on 0.11.0). A candidate above 2**5 and not above 2**10 is now read a second time: its step is halved down to the one float spacing across which the residual departs most, and that departure is compared with the same mirror-image departure plus the linearisation's change over one spacing, against 2**8. Measured: 1.8e3 and 2.0e3 at the two endings, 3.4e4 or more at the other jump endings of 4,500 drawn fits; at most 52 on smooth endings, none of which is asked (their first reading is at most 21). The mirror image is kept from the whole candidate because one pair of neighbouring points is no measure of rounding: with it taken across one spacing too, smooth floors read 4e3 to 1e6. MADD-ANO-216 (never released) gains the residual case; SYS-080 quotes the second reading. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UkCde7g23gTziUjYAvnKD
NicholasEhsanRoy
force-pushed
the
fix/p4-42-sysid-jump-residual
branch
from
October 7, 2026 10:03
603e301 to
2e6f3ac
Compare
NicholasEhsanRoy
marked this pull request as ready for review
October 7, 2026 10:04
…nnot have spec_and_value drew ParamSpec(bounds=(1.0027701900272462e-38, 1.0), transform='logit'): a subnormal float32 lower bound, which ParamSpec refuses by name for a float32 leaf. The strategy predates that refusal, so two properties of TestBoundsAndTransforms failed at random. A subnormal logit bound is now moved to 0.0 (the bound the message recommends) rather than rejected; the drawn spec joins the refusal's own test, and the accepted side of the boundary is an explicit example. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UkCde7g23gTziUjYAvnKD
…s were read twice On jaxlib 0.11.2 the ball's trajectory rounds differently and the ending of the first cell has two candidates between the thresholds (second readings 4.4e3 and 2.1e3) where 0.10.2 and 0.11.0 have one (2.0e3). The test counted them. It now asserts what the fix guarantees on every version: converged=False with one warning; the first reading alone finds no jump and is in the band read again; the verdict is a second reading over its threshold, and the warning prints it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013UkCde7g23gTziUjYAvnKD
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What was wrong
tests/property/test_sysid_targeted_search.py::test_no_wrong_fit_found_by_the_search[converged-beside-a-lower-loss]went red in the slow lane on jaxlib 0.10.2:fit_lmon the stock bouncing ball returnedconverged=Trueat a loss of 6.0e-3 where the loss still falls along the piece the point is on (score 11.8 against a threshold of 1).Reproduced on both jax versions, identical to the last digit (jaxlib 0.10.2 and 0.11.0;
BallCase(0.6133024381551229, 0.7699776319888987, 0.5, 1.0)):converged=True, loss 0.006037198938429356, 18 iterations, score 11.815881429033254. It is not a 0.10.2 effect; the hunt is random and no PR's CI runs it.Classification: (a), a jump the detection missed. At the last iteration:
tol); the conditional retry fromlam0did not run (_gain_in_the_gapfalse);The ratio was under
2**10because the candidate was too long for a small jump: the jump (0.0147 in the residual's norm, a late bounce that changes 19 of 200 samples) is divided by the linearisation's change over the whole candidate, which grows with its length. The next candidate down the ladder (4 spacings) did not reach the edge. PR #245 measured 2.2e4 to 4.4e5 on 55 jump endings; over 4,500 drawn fits the first reading is 1.9e3 to 5.8e5 at 261 jump endings, and 341 at one more (alsoconverged=Truebefore, score 5.6).The fix
A rejected candidate whose ratio is above
2**5and not above2**10is read a second time (_excess_across_a_spacing): its step is halved, keeping the half over which the residual departs further from the linearisation, down to two neighbouring points of the coordinates' grid; the departure across that one spacing is compared with the same mirror-image departure of the whole candidate plus the linearisation's change over one spacing. Above2**8it is a jump:converged=Falsewith the existing warning (which prints the reading that decided). A candidate at or under2**5, or over2**10, is read exactly as before. Cost: one residual evaluation a halving (7 on the example), at the end of the run, only for such a candidate.Measured separation (CPU; jaxlib 0.10.2 and 0.11.0 give the same numbers on the ball):
logitfit at 21 (unchanged from #245)2**5); computed anyway: at most 13, the widelogit522**8is 4.9 times above the largest smooth second reading and 6.9 times below the smallest jump's;2**5is 1.5 times above the largest smooth first reading (which only decides whether a second reading is taken) and 10 times below the smallest jump's.A variant I measured and rejected: taking the mirror image across one spacing too (two neighbouring points against their own mirror images) reads the example at 3.0e4, but is no measure of rounding: neighbouring residuals often agree exactly, and smooth floors read 4.1e3 (wide
logit), 2.5e4 and 1.0e6 (a noiseless x64 fit at its floor). Mirroring through the nearer end of the pair fails on the ball itself: at the scale of one spacing the two sides of the edge interleave.Proof
test_fit_lm_on_the_ball_is_not_converged_where_the_loss_falls_along_its_piece[cell2, cell3]: the hunt's example and the drawn one. Fail on the base (origin/release/0.4.0src, both jax versions), pass now.test_fit_lm_reads_a_small_jump_of_the_ball_from_a_long_candidate[cell0, cell1]:converged=False, one warning naming the second reading; with the second reading switched off,converged=Trueat the same point with the score over its threshold; a second reading between2**8and2**10ends the run the same way.tests/core/test_sysid_non_differentiable_residual.py: a small jump beside the iterate, inside the step and at its far end; a jump reported beside a larger smooth reading, in either order; a jump met by one on the other side; a smooth residual whose slope grows byexp(6.5)over the candidate (second reading under 1); nothing read twice outside the two thresholds; a residual that is not finite inside the step.fit_lmand of the slow tests below, with the rule instrumented: 250 endings in 47 existing tests (140 of smooth fits in 39 tests, 110 of the ball), verdict changed in 0. No smooth ending is read a second time.plans/tools/mutants.py, scratch copy ofsrc/): 19 seeded, 17 caught, 2 documented equivalents (stationary = Trueon the verdict line, as in fix(sysid): fit_lm tells a jump of the residual from the rounding floor, and warns when excited_rank reads full on a degenerate fit #245: the warning'sbreakends the run either way; dropping the finiteness test of the second reading: a NaN compares false against the threshold)._JUMP_SUSPECTto2**10, to 0; the second reading never taken..._is_read_once, the smooth fit'smax(read) < _JUMP_SUSPECT_JUMP_ACROSSto2**11, to2**5..._met_by_one_on_the_other_side......_read_across_one_spacing[1.00019]..._the_one_reported,..._met_by_one......_the_one_reported..._read_across_one_spacing;..._the_one_reportedtest_the_one_sided_excess_of_a_jump...2**9)Code paths this enables or changes
fit_lm, floor rule, both branches (the damped candidates and the undamped Gauss-Newton candidate are in the same list): a candidate between the thresholds now costs up to 64 residual evaluations (measured 7 to 21) and can end the runconverged=Falsewith the warning. Nothing else in the loop changes; no compiled program changes (scripts/compile_counts.py --checkon jax 0.10.2: counts match)._one_sided_excessreturns(excess, move, jumped)(private; was a pair);fit_lmreads the verdict from it.Siblings checked (every place that declares convergence, as in #245): the floor rule (fixed, both branches);
_gain_in_the_gap(the other user of2**10: a promised reduction against the loss's rounding, with no step length in the denominator, so nothing dilutes it; unchanged); the proposal test andtol(no rejected candidates to read; documented as not detecting a jump);fitandfit_multiple_shooting(tolonly). Still not detected, as documented: a kink, a jump met by its like on the other side, a jump whose mirror image the bounds do not allow, an iterate the proposal test passes, a jump under2**5times the linear change of the step that crossed it.Records
Local runs
Gates (
scripts/check_*.py, all nine) andgenerate_soup_tables.py --check: pass. Per push, jax 0.11.0: the two changed test files, the three source scans, the soup, release-notes, claims and sysid-docs compliance files: 469 passed. jax 0.10.2: the two changed test files: 35 passed. pyright onsysid.py: no new error (one in_ExcitationTracker.projectoron this box's numpy stubs, the same on the base).Slow tests run once (jax 0.11.0, cores 12-14; 69 passed in 41 min): every slow test of
tests/property/test_sysid_targeted_search.py(the four hunts),test_sysid_truth_recovery.py(41),test_sysid_one_evaluation.py,test_sysid_contract.py(14),test_differential_precision.py(5),tests/core/test_sysid.py,test_sysid_claims_edges.py,test_cross_feature_edge_cases.py(2).Also here: a property-test strategy that drew a spec the library refuses
tests/property/test_sysid_contract.py::spec_and_valuedrewParamSpec(bounds=(1.0027701900272462e-38, 1.0), transform='logit'): a subnormal float32 lower bound, whichParamSpechas refused by name for a float32 leaf since PR #249. The strategy predates the refusal, sotest_constrain_inverts_unconstrain_inside_the_boundsandtest_constrain_lands_inside_the_bounds_from_any_coordinate(the two users of the strategy) failed at random: 2 of 3 per-push runs here, the same against the base'ssrc/.logitbound that is a subnormal float32 (non-zero and below 1.17549e-38 once rounded to float32). Such a draw is moved to0.0, the bound the refusal's message recommends, not rejected. Left as drawn: zero, the smallest normal, bounds no float32 holds exactly, and subnormal bounds undertransform=Noneandlog(legal; the existing@examples with1.4e-45stay). The upper boundlo + spancannot be subnormal (span >= 1e-2). No library change, no assertion loosened.tests/core/test_constrain_lands_inside_bounds_near_the_smallest_normal.py::test_a_logit_interval_narrower_than_four_smallest_normals_is_refused_by_name, which already had the shape(tiny / 2, 1.0); the drawn spec is now an explicit case there (float32:to_constrained,to_unconstrainedandcheckrefuse it by name; x64: an ordinary bound, mapped).test_the_spec_strategy_moves_only_a_subnormal_logit_boundpins what is moved and what is not, and that the drawn spec is refused where its neighbour is mapped. Both properties gain@examples on the accepted side of the boundary ((0.0, 1.0)and(smallest normal, 1.0)).logitbounds near zero, intest_differential_param_acceptance.py, treats a refusal as an outcome it compares across doors; the slow fit test of the class draws bounds of 0.5 and up.TestBoundsAndTransformsper push with--hypothesis-seed1 to 6 on jax 0.10.2 and on 0.11.0 (12 runs, all pass, including the replay of the stored failing draw); its slow test once (pass); the refusal's file (pass).jaxlib 0.11.2
CI's second lane is 0.11.2 (the shared venv's 0.11.0 is a third version). There the ball's trajectory rounds differently: the first pinned ending has two candidates between the thresholds, with second readings 4.4e3 and 2.1e3, where 0.10.2 and 0.11.0 have one at 2.0e3. The detection is the same;
test_fit_lm_reads_a_small_jump_of_the_ball_from_a_long_candidatehad counted the candidates. It now asserts what holds on every version:converged=Falsewith one warning; the first reading alone finds no jump and its largest ratio is in(2**5, 2**10]; the verdict is one of the second readings, over2**8, and the warning prints it. The two changed test files, the strategy's class and the refusal's file pass on 0.10.2, 0.11.0 and 0.11.2 (64 passed each); the seeded hunt on 0.11.2, seeds 1 to 3: nothing over the threshold; the ten mutants that this test catches were re-run against it: 10 caught.🤖 Generated with Claude Code
https://claude.ai/code/session_013UkCde7g23gTziUjYAvnKD