Skip to content

docs(spec): re-measure every EXPERIMENTS.md sweep at one commit, and correct what no longer holds - #87

Merged
justin13888 merged 10 commits into
masterfrom
fix/75-reverify-experiments-sweeps
Sep 25, 2026
Merged

justin13888 merged 10 commits into
masterfrom
fix/75-reverify-experiments-sweeps

Conversation

@justin13888

@justin13888 justin13888 commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Every sweep spec/EXPERIMENTS.md §6 lists, plus rd-budget on both splits, was re-run from a clean tree at one commit, e310475, with iqa-cli 1.2.1 on the pinned corpus. The 54 results are committed under tools/comparison/results/ (#73's layout). verify:experiments --strict now checks 1229 cells against them, up from 1040. Each finding #75 lists was then corrected in place, or shown to hold, as set out below.

The re-run. Of the 36 results #73 committed, 35 re-ran per-image identical. Only the recorded revision moved, and in two cases a config description. The exception is budget-ladder-tuned, whose config this PR fixes (see below). 18 results are committed for the first time. These are the §6 sweeps no table had bound: allocation-grid, precision-by-budget, encoder-compute, retune-32b, refine-objective, refine-grid, selection-hv, cfl-range, combined-optimizer, alpha-layout-control, alpha-encoder, alpha-balance, quant-ranges, scalefactor-bands and compact-tier-alpha, plus v07-holdout-photo on both splits and rd-budget on tune. The one §6 line that cannot run is v07-holdout-alpha --split holdout, because an alpha holdout image was deleted from Commons (#83). The three probes that are not sweeps (cfl-probe, coeff-stats, entropy-budget) and the §12.1/§12.4 report run were also re-run, and checked by hand.

Each finding in #75

Finding Outcome
§1: the 32 B point is the L26/C9 arm Changed. The ladder keeps its single 26:9 shape, the 32 B cell is no longer bolded as the default, the intro says which tiers are on the curve, and a new bound table puts the shipped L28@4 C15@3 default (11.47 / 11.30) beside the ladder arm (11.65 / 11.54). §9.5's row is relabelled.
§4.5 t1: the Δ rows are wrong; "28 bytes" contradicts the table Changed. Root cause: the pre-adoption shipped rows were §2's ladder on the retired corpus, set beside rows from the current one. There was a second defect as well: every tuned arm omitted sel_hv, so it inherited the adopted 0.15 and measured the round-2 recipe (at 21–411 B it equalled §7.12's optimized row to the digit). budget-ladder-tuned.json now pins sel_hv=0 and gains 14 pre-adoption arms, measured in the same run. All six rows are bound, the Δ rows included. The 28-byte claim holds on both splits, with the numbers re-derived (tune 11.60 vs 11.65, holdout 11.61 vs 11.74). §7.12 t1's pre-adoption rows are bound to the same arms. §7.12 and §8.3 restate the 20% equal-quality saving that §8.3 had withdrawn: it holds on holdout (11.298 at 32 B against 11.31 at 40 B) and not on tune.
spec/README:1639: RGB565 "reads inverted" Changed, and the audit's reading was half right. At ~1.6 kB RGB565 beats code 4 on holdout ΔE00 (6.570 vs 6.768) and loses on tune (6.960 vs 6.726). Only the tune run had been read, and only holdout rd-budget was committed. §11.14 gains the ~1.6 kB rows. The spec sentence is restated from them, with the "20–40%" gap and the "roughly 4%" entropy figure re-derived. The three figures are registered in verify-claims. §2's note and §9.5's row now say the result depends on the split.
spec/README §15 rows Changed. The weighting-matrices and progressive rows now carry their measured outcomes: §11.9's −0.13%, and §7.11's +4.21% at the 32-byte cut.
§11.12: the holdout column is unbound; check the table-2 lookup Changed. Both columns of the photographic table are bound to v07-holdout-photo on each split. The positioning table's resolver read the format name as the byte count, so it matched nothing and every row was found by the label fallback. It now has a resolver for its column order.
§11.3: the FAIL cell, the post-hoc 35% cut, and §8.1's note Changed. The αMAE cell holds the value, 0.1901. The failing guard (Butteraugli) is named on the row. The 35% cut is disclosed as chosen after the sweep, and it holds only 4 images. A median cut, fixed by the corpus alone, is added. Both cuts are bound. The adopted row ranks first in the more-opaque half on both, by 0.18 pp over the next arm on the median cut, and that is stated as a ranking, not a separation. §8.1 points forward to §11.1/§11.3.
§12.2: "all guards improving" Changed to "no guard regressing (the guards must hold, not improve)".
Compact tier at −2.78%, below the 3% rule Changed. It is now stated that RATIONALE.md's rule governs changing a shipped constant. The compact tier was a new code with no incumbent, and it was admitted on §8.1/§8.6's criterion: beat ThumbHash on all four metrics out of sample. That criterion was never written into RATIONALE.md, which is stated as a gap.
§13.3 / §13.4: r is not significant at n=31; codes 1 and 4 are not at the same ΔE00 Changed. No detail-axis coefficient reaches 0.355, so the text reads a lean, not a relationship. Chroma is the one axis that reaches significance, and only code 2 survives Holm. ΔE00 vs detail is significant only at codes 3–4. §13.4's "two defects" becomes "possibly two", and its 1:1-vs-1:7 claim notes that the two ratios sit at different fidelity (6.726 vs 11.473).
"Consulted once" (:1927 and the §11.12 heading) Changed. Both now record the history: six readings, what each decided, that sel_hv = 0.15 was selected on holdout, and that the split's composition changed twice. §8's intro stops calling the split unread. The two v07-holdout-* config descriptions also stop calling themselves the single, run-once validation.
stats.ts:71 "1892-line document" Already fixed on master by #84; the comment reads "the whole document". Nothing to change.
§9.5's config and arm counts Changed. A dated note gives the current counts (831 arms in 51 configs: 481 enforceable, 103 exempt, 247 unreadable) and says where the 24 new arms came from.
RATIONALE CI rounding difference Flagged, not changed. RATIONALE's "11.36 (CI 10.3–12.5)", given as v0.6's holdout score, matches the v1 row of its own table (11.364 [10.30, 12.45]), not the v0.6 row (11.312 [10.25, 12.39]). No record says which run produced it, and that corpus is gone, so the text says so rather than guessing. RATIONALE's pointer to output/sweeps/ is updated too.
holdout-images.ts:75: "Wikimedia fallback" Changed. The comment now says the plain-HTTP URL is the only fallback.
holdout-candidates.json: "Kodak24 + 4" Changed to 8, which is verified against the result's image list.
rd-gate.ts: "tier 0" Confirmed, then changed. The gate scores TIER = 1, and DEFAULT_TIER = 1 is the 32-byte default. So "tier 0" was the old numbering. The header, the measure() doc, the baseline note it writes, the mise task comment and the CI step name now say "default tier (code 1)".
§6: missing tune command; compact-tier-alpha "never run" Changed. The tune lines for v07-holdout-alpha and v07-holdout-photo are added. compact-tier-alpha was run. Its result reproduces §11.12's tune figure for the adopted compact alpha row (−13.00%) to the digit, so §11.10 reports it in a new bound table, and §6's "never run" claim is withdrawn.

Found by the re-measure, beyond #75's list

  • §4.6 and §4.8 still carried retired-corpus figures. §4.6's µ plateau was 10.229–10.286 and is now 11.406–11.415. §4.8's CfL probe was |ρ| 0.457/0.504 and is now 0.463/0.579, and the two example images it named are not in the current corpus. Both are re-derived, and so are §7.10's figures that quote §4.8. §9.5's "not re-measured" list gains both.
  • The tools label L26@5 C9@4 as "shipped". §4.9 and §7.13 read coeff-stats and entropy-budget off that label, so both now name the row. Every figure in both reproduced.
  • §13.5's seed-sensitivity list printed two separate maxima as one interval's. It is re-measured over the 54 results: 839 intervals, 23 interval flips and 14 Holm flips. The deferred resample increase is comparison: raise the paired-interval resamples above 1,000, as §13.5 asks, and re-transcribe every quoted interval #88.
  • More tables bound. §4.2, §4.4 and §11.2 are now bound (their sources became committed results). §4.2's two cells in the last digit are corrected.
  • Re-checked by hand, not bound. §4.2 t1, §4.7, §7.2, §7.4/§7.9's collision claims, §7.14's U16 winners, §11.3's prose figures, §11.8 and §11.9 all reproduce. So do §12.1 and §12.4, re-run from the report command.

Changed paths

  • tools/comparison/results/*.json: the 54 results (35 re-recorded, 1 re-recorded after its config fix, 18 new).
  • tools/comparison/sweeps/budget-ladder-tuned.json: pins sel_hv=0, adds the pre-adoption arms and description. holdout-candidates.json, v07-holdout-photo.json, v07-holdout-alpha.json: descriptions only. All four changed before the re-run, so every result records one config digest. The merge of master then changed v07-holdout-alpha.json's description once more (decision 12). The recorded digest is of the pre-merge file, and no arm changed.
  • tools/comparison/src/verify-experiments.ts: a per-series resolver, byBudgetIn, byFormatThenBytes, per-column image subsets (subsetDeltaPct, the transparency cuts), the new bindings, the register notes and EXPECTED_CELLS.
  • tools/comparison/src/verify-claims.ts: three §11.14 claims.
  • tools/comparison/src/rd-gate.ts, tools/comparison/src/holdout-images.ts, .mise.toml, .github/workflows/ci-comparison.yml: comments, the baseline note string, and one CI step name.
  • spec/EXPERIMENTS.md, spec/README.md, spec/RATIONALE.md: as above.

Validation

All of these were first run at 62568d5. Master was then merged in at f332727 (#86, #89, #90, #94, #96, #97), and every command below was re-run at f332727, the pushed head, with the same outcomes. verify-claims still passes. verify:experiments --strict still checks 1229 cells. determinism-check and rd-gate pass on master's encoder changes, which are byte-identical. The attribution check's failure-path step also passes.

  • pnpm --prefix tools/comparison run format:check: pass
  • pnpm --prefix tools/comparison run lint: pass
  • pnpm --prefix tools/comparison run build: pass
  • node tools/comparison/dist/metric-selftest.js: pass
  • node tools/comparison/dist/verify-claims.js: pass (23 claims, up from 20)
  • node tools/comparison/dist/verify-experiments.js --strict: pass. 1229 cells, EXPECTED_CELLS asserts 1229, 0 SKIP, 0 provenance problems.
  • node tools/comparison/dist/verify-sweep-labels.js: pass
  • node tools/comparison/dist/corpus-licenses.js --check: pass
  • node tools/comparison/dist/metric-reference-test.js: pass
  • cargo build --manifest-path rust/Cargo.toml --features research-render --example encode_stdin, and the same with --release: pass
  • node tools/comparison/dist/determinism-check.js: pass. It reproduces the re-recorded results/synthesis-window.json.
  • node tools/comparison/dist/rd-gate.js: pass, 0.00% drift
  • pnpm --prefix tools/comparison run compare --skip-harnesses --output output/index.html: pass
  • python3 spec/validate.py: pass. bash tools/ci/check-versions.sh: pass.
  • convco check origin/master..HEAD: no errors.
  • node tools/comparison/dist/verify-benchmark.js: exit 1, pre-existing. These are the six TBD PERFORMANCE.md cells (perf: fill PERFORMANCE.md's six deferred cells with benchmark:full on a quiet host #78), which CI runs with continue-on-error. This PR touches neither the document nor the runs.
  • Not run locally: the Go, C# and Python harness build steps of ci-comparison.yml. None of their inputs change.

Coverage gaps

Risks and rollout

  • Tooling and documentation only. No format, encoder, constant or test vector changes.
  • The results directory grows from 1.7 MB to 3.1 MB. The results must be re-recorded whenever the encoder, the scoring path or the pinned iqa-cli changes; selftest:determinism fails in CI when they stop reproducing.

Issue

Closes #75

Decisions taken

  1. How §1 stops presenting the L26/C9 arm as the default.
    Taken: keep the ladder on one 26:9 shape, unbold the 32 B column, and add a bound two-row table with the shipped L28@4 C15@3 arm beside it. The curve and its slope row stay one experiment.
    Rejected: rebinding the 32 B cell to the shipped arm. That would put an off-curve point inside the curve, and inside §1 t2's 16→32 and 32→64 slopes.
    Reverses: point §1 t0's 32 column at 32 B t0 SHIPPED (a resolver that prefers the incumbent), and drop §1 t3 and its binding.
  2. Where §4.5/§7.12's pre-adoption baseline comes from.
    Taken: 14 pre-adoption arms inside budget-ladder-tuned, the v0.6-derived format resized as budget-ladder resizes it with the four knobs pinned off, so baseline and candidate are paired in one run. The sel_hv=0 pin goes in the same change.
    Rejected: a new config, which would make two runs and unpaired rows. Also rejected: dropping the rows and their Δs, which spec: re-measure every EXPERIMENTS.md sweep at one commit, and correct what no longer holds #75 asks to correct, not remove.
    Reverses: revert sweeps/budget-ladder-tuned.json, re-run it, restore the two UNBOUND_COLUMN_NOTES entries, and drop the pre-adoption and Δ series.
  3. The raster of the pre-adoption arms at ≥108 B.
    Taken: the upper tier's (tier: 2), exactly as budget-ladder and the shipped format render those budgets. The tuned arms render at 32 px, and §4.5 now says this understates the recipe by up to §4.1's 0.92% at 411 B.
    Rejected: rendering them at 32 px to match the tuned arms. That would be a format that never shipped, labelled "pre-adoption shipped".
    Reverses: drop "tier": 2 from the four ≥108 B pre-adoption arms and re-run.
  4. compact-tier-alpha: run or remove.
    Taken: run. Its result reproduces the adopted compact alpha row's tune figure, so it is the evidence for a shipped constant, and §11.10 reports it.
    Rejected: removing the §6 line, which would leave that row with no reproducible tune evidence.
    Reverses: delete its result, the §11.10 table and binding, and the §6 line.
  5. §11.3's post-hoc 35% cut.
    Taken: disclose it, and add a median cut fixed by the corpus beside it. Both are bound.
    Rejected: replacing it with the median. That would erase the record of what the decision actually used.
    Reverses: drop the two median columns from the table and from ALPHA_SUBSETS.
  6. Which rule the −2.78% compact adoption was under.
    Taken: state that the ≥3% retune rule does not cover a new tier code with no incumbent. It was admitted on §8.1/§8.6's pre-stated criterion (beat ThumbHash on all four, out of sample), and that criterion's absence from RATIONALE.md is named as a gap.
    Rejected: reporting it as a pass, or as a failed rule applied anyway. Neither is what happened.
    Reverses: edit the paragraph under §11.12's photographic table.
  7. The resample increase §13.5 deferred "to the next full re-measure".
    Taken: split to comparison: raise the paired-interval resamples above 1,000, as §13.5 asks, and re-transcribe every quoted interval #88. This re-measure keeps the method fixed, so its intervals are comparable with comparison: paired CIs for every metric, Holm correction, and a CI on stratify's r #74's.
    Rejected: folding it in, which would move every quoted interval in a PR whose point is that the numbers did not move.
    Reverses: close comparison: raise the paired-interval resamples above 1,000, as §13.5 asks, and re-transcribe every quoted interval #88 and raise bootstrapCI's default here.
  8. Paths outside the planned manifest.
    Taken: .mise.toml and ci-comparison.yml (the rd-gate "tier 0" wording, the same finding at its other sites); spec/RATIONALE.md (its results path and the unattributable figure, which spec: re-measure every EXPERIMENTS.md sweep at one commit, and correct what no longer holds #75 lists); verify-claims.ts (the registration spec: re-measure every EXPERIMENTS.md sweep at one commit, and correct what no longer holds #75 asks for). The manifest's sweeps/*.json for results is comparison: commit sweep results with provenance, and run verify:experiments --strict in CI #73's results/ directory.
    Rejected: leaving the same stale wording at the sites the manifest did not name.
    Reverses: revert those hunks.
  9. The rd-gate baseline's note.
    Taken: change the note string the tool writes, and leave baselines/rd-gate.json untouched. A --update on this host rewrote five values in their last binary digit, which is not a measurement change.
    Rejected: committing ulp-only drift. Also rejected: hand-editing a recorded file's note.
    Reverses: run mise run rd:gate:update and commit it.
  10. RATIONALE's 11.36 (CI 10.3–12.5).
    Taken: flag that it matches the v1 row rather than v0.6's, and say that it cannot be re-derived.
    Rejected: swapping in 11.312 [10.25, 12.39], which assumes which run it came from.
    Reverses: edit the bullet under "Evaluation methodology".
  11. "Never-tuned" wording outside EXPERIMENTS.md.
    Taken: left for comparison: a new sealed holdout (holdout2), with the spent one folded into tune #76, which retires the split. EXPERIMENTS.md records the history.
    Rejected: rewording every quote here, in a PR whose job is the measurements.
    Reverses: reword the three files.
  12. How the merge of master resolves its two conflicts.
    Taken: in spec/EXPERIMENTS.md, keep master's new tune-only paragraph at the end of §11.11 and this PR's §11.12 heading ("and how often it has been read"). In sweeps/v07-holdout-alpha.json, keep this PR's history ("run again once the mislabelled incumbent below was fixed", in place of master's "single ... run once") and add master's retirement text (comparison: the alpha holdout split lost cutout-wordmark-aflac, deleted from Commons as a copyright violation #83: --split holdout refuses to run, and tune plus §11.3's −17.10% is the reproducible evidence). The result is not re-recorded for a description-only change. Master's fix(comparison): retire the alpha holdout split, and probe every pinned source's availability and licence #90 set that precedent for this same file.
    Rejected: taking either side whole. Master's heading and "run once" wording repeat the claim spec: re-measure every EXPERIMENTS.md sweep at one commit, and correct what no longer holds #75 corrects. This PR's side would drop the retirement facts fix(comparison): retire the alpha holdout split, and probe every pinned source's availability and licence #90 established.
    Reverses: edit the §11.12 heading and the config's description.
  13. How the restored 20% equal-quality saving is stated (review note D1, fixed at 1e0895f).
    Taken: state it as not separable on holdout and less than 20% on tune, with no split quoted alone. Paired per image with bootstrapCI's default seed (adopted 32 B minus pre-adoption 40 B, ΔE00): holdout −0.009, CI [−0.098, +0.082], 15/32 wins; tune +0.182, CI [+0.066, +0.308], 9/31 wins; tune adopted vs pre-adoption 32 B −0.182, CI [−0.282, −0.085]. §7.12's paragraph, §8.3's note and spec/README.md §7 now say the saving is at most 20%, and on holdout a match, not an improvement. This supersedes the §4.5 t1 row's "it holds on holdout ... and not on tune" wording above: on holdout it holds only as a tie.
    Rejected: keeping "better than" and "the holdout figure is the one to quote" and recording the choice here. A 0.009 gap inside a CI that spans zero does not show "better", and the only split that separates points the other way.
    Reverses: revert 1e0895f.
    Validation at 1e0895f: format:check, lint, build, metric-selftest, verify-claims, verify-experiments --strict (1229 cells), verify-sweep-labels, corpus-licenses --check and spec/validate.py pass. The change is prose in two spec files, so the encoder, determinism and rd-gate steps were not re-run. The new prose figures are not bound by any gate; they were computed from the committed final-candidates-holdout, budget-ladder-tuned-holdout and budget-ladder-tuned results with dist/stats.js.

…dout's stated history

rd-gate scores tier code 1, the 32-byte default, but its header, its
measure() doc, the baseline note it writes, the mise task comment and the
CI step name all said "tier 0", the pre-reorder numbering. holdout-images
named a Wikimedia fallback its URL list does not hold. holdout-candidates
described the holdout as Kodak24 + 4 curated photos; it is + 8. The two
v07-holdout configs called themselves the single, run-once holdout
validation, which the split's history does not support.

Config descriptions change here, before the §6 re-run, so every committed
result records one config digest.
…ive it the baseline §4.5 compares against

Every tuned arm set aniso, scale_fit and ac_nearest but not sel_hv, so it
inherited the adopted 0.15 and measured the round-2 recipe: at 21-411 B its
rows equal budget-ladder-optimized's to the digit. The arms now pin
sel_hv=0, which is what the description says the recipe is.

§4.5 and §7.12 compare the tuned and optimized ladders against a
"pre-adoption shipped" row no committed sweep produced, and whose numbers
were the retired corpus's. The config gains that row as fourteen arms: the
v0.6-derived format resized to each budget exactly as budget-ladder does,
with the four selection/encoder knobs pinned off.
Every sweep EXPERIMENTS.md §6 lists, re-run from a clean tree at e310475
with iqa-cli 1.2.1 on the pinned corpus. 35 results replace those
recorded at c070e17 and 323fdf2; 17 are committed for the first time,
for the sweeps §6 cites that no table bound: allocation-grid,
precision-by-budget, encoder-compute, retune-32b, refine-objective,
refine-grid, selection-hv, cfl-range, combined-optimizer,
alpha-layout-control, alpha-encoder, alpha-balance, quant-ranges,
scalefactor-bands, compact-tier-alpha, and v07-holdout-photo on both
splits.

Every re-run result is per-image identical to the one it replaces
except budget-ladder-tuned, whose config changed in the previous commit.
compact-tier-alpha, which §6 described as never run, now has a result.
…mmit

Per-image identical to the result it replaces; only the recorded revision
moves, to e310475 with the rest of the §6 set.
§6 has listed the tune run since round 1, and #73 committed only the
holdout one. Recorded from a clean tree at e310475, the re-measure commit,
alongside the holdout run: §2's tune figures at ~1.6 kB and §11.14's tune
cross-check (8.785 against WebP's 8.944 at ~190 B) come from it.
…correct what no longer holds

Every table verify:experiments binds agrees with the re-recorded results,
and the tables those results made bindable are now bound (1040 → 1229
checked cells):

- §1 labels its 32 B column as the ladder's 26:9 shape and adds the
  shipped L28@4 C15@3 default beside it (11.47 / 11.30); §9.5's row is
  relabelled the same way.
- §4.5 and §7.12's pre-adoption rows were the retired corpus's §2 ladder,
  so every Δ under them compared two corpora; they are now measured arms of
  budget-ladder-tuned, bound with their Δ rows. The 28-byte claim survives
  on both splits, and §7.12/§8.3 restate the 20% equal-quality saving,
  which holds on holdout and not on tune.
- §11.12 binds both columns of its photographic table to v07-holdout-photo
  on each split, fixes the positioning table's resolver (it read the format
  name as the byte count), states which rule admitted the compact tier at
  −2.78%, and replaces "consulted once" with the holdout's actual history.
- §11.3 replaces a FAIL in the αMAE column with the value, and binds the
  subgroup table on the post-hoc 35% cut (now disclosed as such) and on a
  median cut fixed by the corpus; the adopted row ranks first on both.
- §11.10 reports compact-tier-alpha, which §6 called never run.
- §11.14 gains the ~1.6 kB rows; spec/README §14.1's RGB565 claim is
  restated from them (pixels beat code 4 on holdout ΔE00, not on tune) and
  registered in verify-claims, and §2's note and §9.5's row say the same.
- §12.2 states the rule as "no guard regresses"; §13.3/§13.4 stop reading
  a relationship into correlations that are not significant at n = 31 and
  stop calling codes 1 and 4 equal-fidelity.
- §4.6, §4.8 and §7.10 carried retired-corpus figures; re-derived.
- §6 lists the tune commands, drops the "never run" claim, and says where
  the results live; §9.5 and §13.5 carry current counts, and §13.5's
  resample change is #88.
- spec/README §15 records the measured outcomes of weighting matrices and
  progressive tiers.
…attributable holdout figure

RATIONALE still sent readers to the gitignored output/sweeps/ for results
that are now committed under tools/comparison/results/. And its account of
the v0.6 tune/holdout gap quotes 11.36 (CI 10.3–12.5) as v0.6's holdout
score, which matches the v1 row of its own version table (11.364
[10.30, 12.45]) rather than the v0.6 row (11.312 [10.25, 12.39]). No record
says which run produced it and the corpus is gone, so the figure is flagged
rather than changed.
… which label is not today's default

§6 no longer claims the one alpha holdout run that cannot be recorded
(#83) was re-run. §7.13, like §4.9, reads coeff/entropy figures off the
L26@5 C9@4 row the tools call shipped, which is not today's default; both
say so, and both reproduced to the digit. §9.5's list of what the
Wikimedia re-run left stale gains the two it missed, §4.6 and §4.8. §1
points at its new table by position correctly.
@justin13888 justin13888 changed the title spec: re-measure every EXPERIMENTS.md sweep at one commit, and correct what no longer holds docs(spec): re-measure every EXPERIMENTS.md sweep at one commit, and correct what no longer holds Sep 25, 2026
@justin13888
justin13888 marked this pull request as ready for review September 25, 2026 07:24
…ble on holdout

The holdout pairing (adopted 32 B minus pre-adoption 40 B) is -0.009 with paired CI [-0.098, +0.082] and 15/32 wins, while tune separates the other way (+0.182, CI [+0.066, +0.308]). Restate §7.12, §8.3 and spec/README.md accordingly: the saving is at most 20%, a match on holdout rather than an improvement.
@justin13888
justin13888 merged commit b09905c into master Sep 25, 2026
5 checks passed
@justin13888
justin13888 deleted the fix/75-reverify-experiments-sweeps branch September 25, 2026 16:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

1 participant