Repository navigation
docs(spec): re-measure every EXPERIMENTS.md sweep at one commit, and correct what no longer holds - #87
Merged
Conversation
…dout's stated history rd-gate scores tier code 1, the 32-byte default, but its header, its measure() doc, the baseline note it writes, the mise task comment and the CI step name all said "tier 0", the pre-reorder numbering. holdout-images named a Wikimedia fallback its URL list does not hold. holdout-candidates described the holdout as Kodak24 + 4 curated photos; it is + 8. The two v07-holdout configs called themselves the single, run-once holdout validation, which the split's history does not support. Config descriptions change here, before the §6 re-run, so every committed result records one config digest.
…ive it the baseline §4.5 compares against Every tuned arm set aniso, scale_fit and ac_nearest but not sel_hv, so it inherited the adopted 0.15 and measured the round-2 recipe: at 21-411 B its rows equal budget-ladder-optimized's to the digit. The arms now pin sel_hv=0, which is what the description says the recipe is. §4.5 and §7.12 compare the tuned and optimized ladders against a "pre-adoption shipped" row no committed sweep produced, and whose numbers were the retired corpus's. The config gains that row as fourteen arms: the v0.6-derived format resized to each budget exactly as budget-ladder does, with the four selection/encoder knobs pinned off.
Every sweep EXPERIMENTS.md §6 lists, re-run from a clean tree at e310475 with iqa-cli 1.2.1 on the pinned corpus. 35 results replace those recorded at c070e17 and 323fdf2; 17 are committed for the first time, for the sweeps §6 cites that no table bound: allocation-grid, precision-by-budget, encoder-compute, retune-32b, refine-objective, refine-grid, selection-hv, cfl-range, combined-optimizer, alpha-layout-control, alpha-encoder, alpha-balance, quant-ranges, scalefactor-bands, compact-tier-alpha, and v07-holdout-photo on both splits. Every re-run result is per-image identical to the one it replaces except budget-ladder-tuned, whose config changed in the previous commit. compact-tier-alpha, which §6 described as never run, now has a result.
…mmit Per-image identical to the result it replaces; only the recorded revision moves, to e310475 with the rest of the §6 set.
…correct what no longer holds Every table verify:experiments binds agrees with the re-recorded results, and the tables those results made bindable are now bound (1040 → 1229 checked cells): - §1 labels its 32 B column as the ladder's 26:9 shape and adds the shipped L28@4 C15@3 default beside it (11.47 / 11.30); §9.5's row is relabelled the same way. - §4.5 and §7.12's pre-adoption rows were the retired corpus's §2 ladder, so every Δ under them compared two corpora; they are now measured arms of budget-ladder-tuned, bound with their Δ rows. The 28-byte claim survives on both splits, and §7.12/§8.3 restate the 20% equal-quality saving, which holds on holdout and not on tune. - §11.12 binds both columns of its photographic table to v07-holdout-photo on each split, fixes the positioning table's resolver (it read the format name as the byte count), states which rule admitted the compact tier at −2.78%, and replaces "consulted once" with the holdout's actual history. - §11.3 replaces a FAIL in the αMAE column with the value, and binds the subgroup table on the post-hoc 35% cut (now disclosed as such) and on a median cut fixed by the corpus; the adopted row ranks first on both. - §11.10 reports compact-tier-alpha, which §6 called never run. - §11.14 gains the ~1.6 kB rows; spec/README §14.1's RGB565 claim is restated from them (pixels beat code 4 on holdout ΔE00, not on tune) and registered in verify-claims, and §2's note and §9.5's row say the same. - §12.2 states the rule as "no guard regresses"; §13.3/§13.4 stop reading a relationship into correlations that are not significant at n = 31 and stop calling codes 1 and 4 equal-fidelity. - §4.6, §4.8 and §7.10 carried retired-corpus figures; re-derived. - §6 lists the tune commands, drops the "never run" claim, and says where the results live; §9.5 and §13.5 carry current counts, and §13.5's resample change is #88. - spec/README §15 records the measured outcomes of weighting matrices and progressive tiers.
…attributable holdout figure RATIONALE still sent readers to the gitignored output/sweeps/ for results that are now committed under tools/comparison/results/. And its account of the v0.6 tune/holdout gap quotes 11.36 (CI 10.3–12.5) as v0.6's holdout score, which matches the v1 row of its own version table (11.364 [10.30, 12.45]) rather than the v0.6 row (11.312 [10.25, 12.39]). No record says which run produced it and the corpus is gone, so the figure is flagged rather than changed.
… which label is not today's default §6 no longer claims the one alpha holdout run that cannot be recorded (#83) was re-run. §7.13, like §4.9, reads coeff/entropy figures off the L26@5 C9@4 row the tools call shipped, which is not today's default; both say so, and both reproduced to the digit. §9.5's list of what the Wikimedia re-run left stale gains the two it missed, §4.6 and §4.8. §1 points at its new table by position correctly.
justin13888
marked this pull request as ready for review
September 25, 2026 07:24
This was referenced Sep 25, 2026
…ble on holdout The holdout pairing (adopted 32 B minus pre-adoption 40 B) is -0.009 with paired CI [-0.098, +0.082] and 15/32 wins, while tune separates the other way (+0.182, CI [+0.066, +0.308]). Restate §7.12, §8.3 and spec/README.md accordingly: the saving is at most 20%, a match on holdout rather than an improvement.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Every sweep
spec/EXPERIMENTS.md§6 lists, plusrd-budgeton both splits, was re-run from a clean tree at one commit,e310475, with iqa-cli 1.2.1 on the pinned corpus. The 54 results are committed undertools/comparison/results/(#73's layout).verify:experiments --strictnow checks 1229 cells against them, up from 1040. Each finding #75 lists was then corrected in place, or shown to hold, as set out below.The re-run. Of the 36 results #73 committed, 35 re-ran per-image identical. Only the recorded revision moved, and in two cases a config description. The exception is
budget-ladder-tuned, whose config this PR fixes (see below). 18 results are committed for the first time. These are the §6 sweeps no table had bound:allocation-grid,precision-by-budget,encoder-compute,retune-32b,refine-objective,refine-grid,selection-hv,cfl-range,combined-optimizer,alpha-layout-control,alpha-encoder,alpha-balance,quant-ranges,scalefactor-bandsandcompact-tier-alpha, plusv07-holdout-photoon both splits andrd-budgeton tune. The one §6 line that cannot run isv07-holdout-alpha --split holdout, because an alpha holdout image was deleted from Commons (#83). The three probes that are not sweeps (cfl-probe,coeff-stats,entropy-budget) and the §12.1/§12.4 report run were also re-run, and checked by hand.Each finding in #75
L28@4 C15@3default (11.47 / 11.30) beside the ladder arm (11.65 / 11.54). §9.5's row is relabelled.pre-adoption shippedrows were §2's ladder on the retired corpus, set beside rows from the current one. There was a second defect as well: everytunedarm omittedsel_hv, so it inherited the adopted 0.15 and measured the round-2 recipe (at 21–411 B it equalled §7.12'soptimizedrow to the digit).budget-ladder-tuned.jsonnow pinssel_hv=0and gains 14 pre-adoption arms, measured in the same run. All six rows are bound, the Δ rows included. The 28-byte claim holds on both splits, with the numbers re-derived (tune 11.60 vs 11.65, holdout 11.61 vs 11.74). §7.12 t1's pre-adoption rows are bound to the same arms. §7.12 and §8.3 restate the 20% equal-quality saving that §8.3 had withdrawn: it holds on holdout (11.298 at 32 B against 11.31 at 40 B) and not on tune.rd-budgetwas committed. §11.14 gains the ~1.6 kB rows. The spec sentence is restated from them, with the "20–40%" gap and the "roughly 4%" entropy figure re-derived. The three figures are registered inverify-claims. §2's note and §9.5's row now say the result depends on the split.v07-holdout-photoon each split. The positioning table's resolver read the format name as the byte count, so it matched nothing and every row was found by the label fallback. It now has a resolver for its column order.FAILcell, the post-hoc 35% cut, and §8.1's noteRATIONALE.md's rule governs changing a shipped constant. The compact tier was a new code with no incumbent, and it was admitted on §8.1/§8.6's criterion: beat ThumbHash on all four metrics out of sample. That criterion was never written intoRATIONALE.md, which is stated as a gap.sel_hv = 0.15was selected on holdout, and that the split's composition changed twice. §8's intro stops calling the split unread. The twov07-holdout-*config descriptions also stop calling themselves the single, run-once validation.stats.ts:71"1892-line document"output/sweeps/is updated too.holdout-images.ts:75: "Wikimedia fallback"holdout-candidates.json: "Kodak24 + 4"rd-gate.ts: "tier 0"TIER = 1, andDEFAULT_TIER = 1is the 32-byte default. So "tier 0" was the old numbering. The header, themeasure()doc, the baseline note it writes, the mise task comment and the CI step name now say "default tier (code 1)".compact-tier-alpha"never run"v07-holdout-alphaandv07-holdout-photoare added.compact-tier-alphawas run. Its result reproduces §11.12's tune figure for the adopted compact alpha row (−13.00%) to the digit, so §11.10 reports it in a new bound table, and §6's "never run" claim is withdrawn.Found by the re-measure, beyond #75's list
L26@5 C9@4as "shipped". §4.9 and §7.13 readcoeff-statsandentropy-budgetoff that label, so both now name the row. Every figure in both reproduced.Changed paths
tools/comparison/results/*.json: the 54 results (35 re-recorded, 1 re-recorded after its config fix, 18 new).tools/comparison/sweeps/budget-ladder-tuned.json: pinssel_hv=0, adds the pre-adoption arms and description.holdout-candidates.json,v07-holdout-photo.json,v07-holdout-alpha.json: descriptions only. All four changed before the re-run, so every result records one config digest. The merge of master then changedv07-holdout-alpha.json's description once more (decision 12). The recorded digest is of the pre-merge file, and no arm changed.tools/comparison/src/verify-experiments.ts: a per-series resolver,byBudgetIn,byFormatThenBytes, per-column image subsets (subsetDeltaPct, the transparency cuts), the new bindings, the register notes andEXPECTED_CELLS.tools/comparison/src/verify-claims.ts: three §11.14 claims.tools/comparison/src/rd-gate.ts,tools/comparison/src/holdout-images.ts,.mise.toml,.github/workflows/ci-comparison.yml: comments, the baseline note string, and one CI step name.spec/EXPERIMENTS.md,spec/README.md,spec/RATIONALE.md: as above.Validation
All of these were first run at
62568d5. Master was then merged in atf332727(#86, #89, #90, #94, #96, #97), and every command below was re-run atf332727, the pushed head, with the same outcomes.verify-claimsstill passes.verify:experiments --strictstill checks 1229 cells.determinism-checkandrd-gatepass on master's encoder changes, which are byte-identical. The attribution check's failure-path step also passes.pnpm --prefix tools/comparison run format:check: passpnpm --prefix tools/comparison run lint: passpnpm --prefix tools/comparison run build: passnode tools/comparison/dist/metric-selftest.js: passnode tools/comparison/dist/verify-claims.js: pass (23 claims, up from 20)node tools/comparison/dist/verify-experiments.js --strict: pass. 1229 cells,EXPECTED_CELLSasserts 1229, 0 SKIP, 0 provenance problems.node tools/comparison/dist/verify-sweep-labels.js: passnode tools/comparison/dist/corpus-licenses.js --check: passnode tools/comparison/dist/metric-reference-test.js: passcargo build --manifest-path rust/Cargo.toml --features research-render --example encode_stdin, and the same with--release: passnode tools/comparison/dist/determinism-check.js: pass. It reproduces the re-recordedresults/synthesis-window.json.node tools/comparison/dist/rd-gate.js: pass, 0.00% driftpnpm --prefix tools/comparison run compare --skip-harnesses --output output/index.html: passpython3 spec/validate.py: pass.bash tools/ci/check-versions.sh: pass.convco check origin/master..HEAD: no errors.node tools/comparison/dist/verify-benchmark.js: exit 1, pre-existing. These are the six TBD PERFORMANCE.md cells (perf: fill PERFORMANCE.md's six deferred cells with benchmark:full on a quiet host #78), which CI runs withcontinue-on-error. This PR touches neither the document nor the runs.ci-comparison.yml. None of their inputs change.Coverage gaps
verify-experiments.tshas no selftest:byBudgetIn,byFormatThenBytes, per-seriesresolve, and the subset Δ%. The file runs on import, as comparison: commit sweep results with provenance, and run verify:experiments --strict in CI #73 noted. The only thing exercising it is--strictover the committed document. Its cells were computed independently by a scratch script over the same results before binding, and every one agreed.UNBOUND_NOTESreason: §4.2 t1, §4.7, §7.2, §7.4, §7.9, the probe tables §4.9/§4.10/§7.13, and the report-run tables §12.1/§12.4. They were re-checked by hand in this re-measure, and a later edit to them is not gated.synthesis-window, by CI's determinism step.v07-holdout-alphaholdout column stays unreproducible (comparison: the alpha holdout split lost cutout-wordmark-aflac, deleted from Commons as a copyright violation #83).spec/README.md,spec/RATIONALE.mdandrust/src/constants.rs. EXPERIMENTS.md now says what that phrase can and cannot mean. Retiring the split is comparison: a new sealed holdout (holdout2), with the spent one folded into tune #76, and rewording every quote of it belongs there.Risks and rollout
selftest:determinismfails in CI when they stop reproducing.Issue
Closes #75
Decisions taken
Taken: keep the ladder on one 26:9 shape, unbold the 32 B column, and add a bound two-row table with the shipped
L28@4 C15@3arm beside it. The curve and its slope row stay one experiment.Rejected: rebinding the 32 B cell to the shipped arm. That would put an off-curve point inside the curve, and inside §1 t2's 16→32 and 32→64 slopes.
Reverses: point §1 t0's 32 column at
32 B t0 SHIPPED(a resolver that prefers the incumbent), and drop §1 t3 and its binding.Taken: 14 pre-adoption arms inside
budget-ladder-tuned, the v0.6-derived format resized asbudget-ladderresizes it with the four knobs pinned off, so baseline and candidate are paired in one run. Thesel_hv=0pin goes in the same change.Rejected: a new config, which would make two runs and unpaired rows. Also rejected: dropping the rows and their Δs, which spec: re-measure every EXPERIMENTS.md sweep at one commit, and correct what no longer holds #75 asks to correct, not remove.
Reverses: revert
sweeps/budget-ladder-tuned.json, re-run it, restore the twoUNBOUND_COLUMN_NOTESentries, and drop the pre-adoption and Δ series.Taken: the upper tier's (
tier: 2), exactly asbudget-ladderand the shipped format render those budgets. The tuned arms render at 32 px, and §4.5 now says this understates the recipe by up to §4.1's 0.92% at 411 B.Rejected: rendering them at 32 px to match the tuned arms. That would be a format that never shipped, labelled "pre-adoption shipped".
Reverses: drop
"tier": 2from the four ≥108 B pre-adoption arms and re-run.compact-tier-alpha: run or remove.Taken: run. Its result reproduces the adopted compact alpha row's tune figure, so it is the evidence for a shipped constant, and §11.10 reports it.
Rejected: removing the §6 line, which would leave that row with no reproducible tune evidence.
Reverses: delete its result, the §11.10 table and binding, and the §6 line.
Taken: disclose it, and add a median cut fixed by the corpus beside it. Both are bound.
Rejected: replacing it with the median. That would erase the record of what the decision actually used.
Reverses: drop the two median columns from the table and from
ALPHA_SUBSETS.Taken: state that the ≥3% retune rule does not cover a new tier code with no incumbent. It was admitted on §8.1/§8.6's pre-stated criterion (beat ThumbHash on all four, out of sample), and that criterion's absence from
RATIONALE.mdis named as a gap.Rejected: reporting it as a pass, or as a failed rule applied anyway. Neither is what happened.
Reverses: edit the paragraph under §11.12's photographic table.
Taken: split to comparison: raise the paired-interval resamples above 1,000, as §13.5 asks, and re-transcribe every quoted interval #88. This re-measure keeps the method fixed, so its intervals are comparable with comparison: paired CIs for every metric, Holm correction, and a CI on stratify's r #74's.
Rejected: folding it in, which would move every quoted interval in a PR whose point is that the numbers did not move.
Reverses: close comparison: raise the paired-interval resamples above 1,000, as §13.5 asks, and re-transcribe every quoted interval #88 and raise
bootstrapCI's default here.Taken:
.mise.tomlandci-comparison.yml(the rd-gate "tier 0" wording, the same finding at its other sites);spec/RATIONALE.md(its results path and the unattributable figure, which spec: re-measure every EXPERIMENTS.md sweep at one commit, and correct what no longer holds #75 lists);verify-claims.ts(the registration spec: re-measure every EXPERIMENTS.md sweep at one commit, and correct what no longer holds #75 asks for). The manifest'ssweeps/*.jsonfor results is comparison: commit sweep results with provenance, and run verify:experiments --strict in CI #73'sresults/directory.Rejected: leaving the same stale wording at the sites the manifest did not name.
Reverses: revert those hunks.
Taken: change the note string the tool writes, and leave
baselines/rd-gate.jsonuntouched. A--updateon this host rewrote five values in their last binary digit, which is not a measurement change.Rejected: committing ulp-only drift. Also rejected: hand-editing a recorded file's note.
Reverses: run
mise run rd:gate:updateand commit it.Taken: flag that it matches the v1 row rather than v0.6's, and say that it cannot be re-derived.
Rejected: swapping in 11.312 [10.25, 12.39], which assumes which run it came from.
Reverses: edit the bullet under "Evaluation methodology".
Taken: left for comparison: a new sealed holdout (holdout2), with the spent one folded into tune #76, which retires the split. EXPERIMENTS.md records the history.
Rejected: rewording every quote here, in a PR whose job is the measurements.
Reverses: reword the three files.
Taken: in
spec/EXPERIMENTS.md, keep master's new tune-only paragraph at the end of §11.11 and this PR's §11.12 heading ("and how often it has been read"). Insweeps/v07-holdout-alpha.json, keep this PR's history ("run again once the mislabelled incumbent below was fixed", in place of master's "single ... run once") and add master's retirement text (comparison: the alpha holdout split lost cutout-wordmark-aflac, deleted from Commons as a copyright violation #83:--split holdoutrefuses to run, and tune plus §11.3's −17.10% is the reproducible evidence). The result is not re-recorded for a description-only change. Master's fix(comparison): retire the alpha holdout split, and probe every pinned source's availability and licence #90 set that precedent for this same file.Rejected: taking either side whole. Master's heading and "run once" wording repeat the claim spec: re-measure every EXPERIMENTS.md sweep at one commit, and correct what no longer holds #75 corrects. This PR's side would drop the retirement facts fix(comparison): retire the alpha holdout split, and probe every pinned source's availability and licence #90 established.
Reverses: edit the §11.12 heading and the config's
description.1e0895f).Taken: state it as not separable on holdout and less than 20% on tune, with no split quoted alone. Paired per image with
bootstrapCI's default seed (adopted 32 B minus pre-adoption 40 B, ΔE00): holdout −0.009, CI [−0.098, +0.082], 15/32 wins; tune +0.182, CI [+0.066, +0.308], 9/31 wins; tune adopted vs pre-adoption 32 B −0.182, CI [−0.282, −0.085]. §7.12's paragraph, §8.3's note andspec/README.md§7 now say the saving is at most 20%, and on holdout a match, not an improvement. This supersedes the §4.5 t1 row's "it holds on holdout ... and not on tune" wording above: on holdout it holds only as a tie.Rejected: keeping "better than" and "the holdout figure is the one to quote" and recording the choice here. A 0.009 gap inside a CI that spans zero does not show "better", and the only split that separates points the other way.
Reverses: revert
1e0895f.Validation at
1e0895f:format:check,lint,build,metric-selftest,verify-claims,verify-experiments --strict(1229 cells),verify-sweep-labels,corpus-licenses --checkandspec/validate.pypass. The change is prose in two spec files, so the encoder, determinism and rd-gate steps were not re-run. The new prose figures are not bound by any gate; they were computed from the committedfinal-candidates-holdout,budget-ladder-tuned-holdoutandbudget-ladder-tunedresults withdist/stats.js.