diff --git a/spec/README.md b/spec/README.md index 9331bd5..276afce 100644 --- a/spec/README.md +++ b/spec/README.md @@ -1673,11 +1673,11 @@ is a few hundred bytes and the target is intentionally low-pass. | XYB opsin color | Perceptual LMS-based color space | **Already covered** — OKLAB is the modern peer; no change. | | Variable DCT block sizes (2×2…32×32, incl. rectangular) | Per-block adaptive transform size | **N/A** — a single global DCT is correct at this bitrate; the *tier* is our "variable" axis. Per-block side-info is unaffordable here. | | Adaptive (spatial) quantization | Per-region quant field from a perceptual heuristic | **Defer** — a per-region quant map is too much side-info for a sub-2 KB payload; possibly justified only at code 4. | -| **Chroma-from-luma (CfL)** | Predict X/B chroma from Y luma with per-group multipliers | **Built; the predictor works and cannot pay for its own side-info** — `EXPERIMENTS.md` §7.10 implemented it as a signalled per-channel least-squares gain and measured it at every tier. Given away free the prediction helps, and grows with tier: −0.09% ΔE00 at 32 B to −0.90% at code 3. What refuses it is the bill for signalling the gains — charged against the AC budget at 32 B the field costs **+2.18%**, more than the prediction returns. This row read "strong v0.8 candidate" until that experiment was run, and is kept as the prediction it scored against. | +| **Chroma-from-luma (CfL)** | Predict X/B chroma from Y luma with per-group multipliers | **Built; the predictor works and cannot pay for its own side-info** — `EXPERIMENTS.md` §7.10 implemented it as a signalled per-channel least-squares gain and measured it at every tier. Given away free the prediction helps, and grows with tier: −0.10% ΔE00 at 32 B to −0.90% at code 3. What refuses it is the bill for signalling the gains — charged against the AC budget at 32 B the field costs **+2.11%**, more than the prediction returns. This row read "strong v0.8 candidate" until that experiment was run, and is kept as the prediction it scored against. | | Gaborish | Small post-decode smoothing convolution | **Re-evaluated, still off** — the decode-side synthesis window (`window_weights`, a Hann taper, disabled by default). `EXPERIMENTS.md` §12.2–§12.3 measured it at codes 1 and 2 with artifact metrics that did not exist when v0.6 rejected it: it removes up to 73% of the invented structure and costs ΔE00 and SSIMULACRA2 monotonically, failing the guards at every strength. At code 2 the lightest taper is statistically free on ΔE00 and fails on SSIMULACRA2 alone. | | Edge-preserving filter (EPF) | Adaptive deringing loop filter | **Reject** — a blurred placeholder has few edges to preserve. | | DC image + DC predictors | Separate DC plane with spatial predictors | **N/A** — chromahash has a single average-color DC per channel, already chosen by the decode-aware DC search (§10.3). | -| **Quantization weighting matrices (HVS/CSF)** | Frequency-dependent quant step | **Built as scalefactor bands; below threshold, not adopted** — `EXPERIMENTS.md` §11.9 measured a frequency-shaped scale split on the current corpus: the best arm is −0.13% ΔE00, far under the ≥3% retune rule and unable to pay for signalling it. This row read "evaluate / adopt" until that sweep ran, and is kept as the prediction it scored against. | +| **Quantization weighting matrices (HVS/CSF)** | Frequency-dependent quant step | **Built as scalefactor bands; below threshold, not adopted** — `EXPERIMENTS.md` §11.9 measured a frequency-shaped scale split on the current corpus: the best arm is −0.14% ΔE00, far under the ≥3% retune rule and unable to pay for signalling it. This row read "evaluate / adopt" until that sweep ran, and is kept as the prediction it scored against. | | **Entropy coding (rANS + context modeling + clustering, HybridUint tokens)** | Adaptive entropy coding of quantized coefficients | **Highest-impact v0.8+** — fixed-width µ-law leaves the most on the table; many high-frequency coefficients quantize to zero and would cost almost nothing under an entropy coder, raising the quality ceiling per byte. Heaviest to make bit-exact across all language bindings (incl. the hand-written TS decoder) and it trades away the fixed-per-tier length, so it is deferred deliberately. | | Coefficient ordering / scan + nonzero context | Frequency-ordered scan, run/EOB modeling | **Already frequency-ordered** — the top-K isotropic selection is exactly this; pairs naturally with entropy coding when added. | | Patches / splines / dots | Reference repeated elements / smooth gradients / point sources | **Reject** — no repeated elements or point sources in a placeholder; the DCT already models smooth gradients compactly. | @@ -1690,7 +1690,7 @@ named here were (1) **entropy coding**, (2) **chroma-from-luma**, (3) **frequenc quantization**, (4) **embedded/progressive tiers**. All four were subsequently built and measured in `EXPERIMENTS.md` §7, and the ordering did not survive: chroma-from-luma predicts, but cannot pay for the gain field that signals it (§7.10), -frequency-weighted quantization is −0.13% and below threshold (§11.9), embedded tiers +frequency-weighted quantization is −0.14% and below threshold (§11.9), embedded tiers cost ~4% against a native encode at the same 32 bytes and are an operational feature rather than a quality one (§7.11), and entropy coding buys −1.6% at 32 B and −4.8% at 108 B, at the cost of the O(1) length check that *is* this format's validity check diff --git a/spec/V0.8-DECISIONS.md b/spec/V0.8-DECISIONS.md index e5b2c69..bf389a4 100644 --- a/spec/V0.8-DECISIONS.md +++ b/spec/V0.8-DECISIONS.md @@ -145,7 +145,7 @@ config), and `Verdict` (the outcome and the committed result path). freeze**: the metric cannot price the loss of the check, so passing R1 is necessary for adoption, but not sufficient. - **Prior evidence (tune only).** `EXPERIMENTS.md` §7.13, with the coder scored - leave-one-image-out: −7.4% of the AC bits at 32 B, which buys **−1.6% ΔE00** + leave-one-image-out: −7.5% of the AC bits at 32 B, which buys **−1.6% ΔE00** at code 1 and **−4.8%** at code 2. Both are against the best fixed-field layout at the same budget, not the shipped one. `RATIONALE.md` ("No entropy coding") says the trade reverses at codes 3–4, where there are 3–13 kbit, but diff --git a/tools/comparison/src/verify-claims.ts b/tools/comparison/src/verify-claims.ts index df9ce74..48058e7 100644 --- a/tools/comparison/src/verify-claims.ts +++ b/tools/comparison/src/verify-claims.ts @@ -3,8 +3,8 @@ * * `verify-experiments.ts` closes the loop between the sweeps and the workbench * log. It does not close the one after it: `README.md`, `spec/README.md`, - * `spec/RATIONALE.md`, `rust/src/constants.rs` and `spec/constants.py` all - * restate figures from that log by hand, and nothing has ever checked them. The + * `spec/RATIONALE.md`, `spec/V0.8-DECISIONS.md`, `rust/src/constants.rs` and + * `spec/constants.py` all restate figures from that log by hand, and nothing has ever checked them. The * drift that follows is not hypothetical — it is what the 2026-09 Wikimedia * re-baseline left behind, and it is the second time. After it: * @@ -489,6 +489,82 @@ const REGISTER: Claim[] = [ cellPattern: PARENTHESISED_DELTA, transform: abs, }, + + // ── The v1 roadmap's outcomes, and the v0.8 register's prior evidence ──── + // Four prose figures that restate §7.10, §7.13 and §11.9 by hand. The #102 + // re-score moved all four cells and none of the quotes, because none was + // registered here: this gate stayed at exit 0 while they drifted. + { + file: "spec/V0.8-DECISIONS.md", + what: "the leave-one-image-out coder's AC-bit saving at 32 B", + pattern: /leave-one-image-out: −([\d.]+)% of the AC bits at 32 B/, + section: "7.13", + row: "**per-index context backing off to order-0, LOO**", + column: "vs fixed", + transform: abs, + }, + { + file: "spec/README.md", + what: "CfL with free gains at 32 B", + pattern: /prediction helps, and grows with tier: −([\d.]+)% ΔE00 at 32 B/, + section: "7.10", + row: "CfL free (gains not paid for)", + column: "vs its own control", + transform: abs, + }, + // The same sentence's other end. Not named in the issue that added this + // block, but it is the same cell column one row down, and leaving it unbound + // would gate half a range. + { + file: "spec/README.md", + what: "CfL with free gains at code 3", + pattern: /ΔE00 at 32 B to −([\d.]+)% at code 3/, + section: "7.10", + row: "tier 3 free", + column: "vs its own control", + transform: abs, + }, + { + file: "spec/README.md", + what: "CfL with paid gains at 32 B, against the shipped layout", + // Signed and without `abs`, like the 108 B entry above: the sentence's + // whole content is that paying for the gains makes the hash *worse*. + pattern: /at 32 B the field costs \*\*(\+[\d.]+)%\*\*/, + section: "7.10", + row: "CfL paid, L24@5 C9@4", + column: "ΔE00", + // §7.10's table carries the paid row's score and measures its delta against + // a size-matched control; the +2.11% is against the `shipped` row above it, + // which only §7.10's prose states. Recomputed from the two cells, for the + // same reason as §11.14's WebP margin: the cells are what the sweep emits. + transform: (paid) => { + const shipped = cellOf("7.10", 0, "shipped", "ΔE00"); + if (shipped === null) throw new Error("§7.10 has no shipped row"); + return ((paid - shipped) / shipped) * 100; + }, + }, + // §11.9 states its best arm only in prose; the one table cell carrying it is + // §9.5's re-source summary, whose `Wikimedia` column is the current figure. + // Binding there is binding to the same measurement, one table removed. + { + file: "spec/README.md", + what: "the best scalefactor-band arm, in the roadmap table", + pattern: /the best arm is −([\d.]+)% ΔE00, far under/, + section: "9.5", + row: "Scalefactor bands, best arm", + column: "Wikimedia", + transform: abs, + }, + { + file: "spec/README.md", + what: "the best scalefactor-band arm, in the roadmap summary", + pattern: + /frequency-weighted quantization is −([\d.]+)% and below threshold/, + section: "9.5", + row: "Scalefactor bands, best arm", + column: "Wikimedia", + transform: abs, + }, ]; // ─── Resolving a cell ───────────────────────────────────────────────────────