Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 3 additions & 2 deletions .github/workflows/ci-comparison.yml
Original file line number Diff line number Diff line change
Expand Up @@ -275,7 +275,8 @@ jobs:

# The comparison report below proves the harness still runs; it does not
# prove the format still performs. This gate (spec/EXPERIMENTS.md U18)
# encodes a handful of content-pinned corpus photos at tier 0 and fails if
# encodes a handful of content-pinned corpus photos at the 32-byte default
# (tier code 1; EXPERIMENTS.md's old numbering calls it tier 0) and fails if
# mean ΔE00 has drifted from tools/comparison/baselines/rd-gate.json, so a
# silent rate-distortion regression cannot land unnoticed. It needs a
# release build — the debug encoder is the same bytes but far slower.
Expand All @@ -293,7 +294,7 @@ jobs:
- name: A concurrent run scores identically to a serial one, and a sweep reproduces its committed result
run: node tools/comparison/dist/determinism-check.js

- name: R-D regression gate (tier 0)
- name: R-D regression gate (default tier, code 1)
run: node tools/comparison/dist/rd-gate.js

- uses: actions/setup-go@v5
Expand Down
2 changes: 1 addition & 1 deletion .mise.toml
Original file line number Diff line number Diff line change
Expand Up @@ -634,7 +634,7 @@ run = [

# ─── Quality gate ───────────────────────────────────────────────────────────

# Tier-0 R-D regression gate: encode a fixed handful of content-pinned corpus
# Default-tier (code 1, 32 B) R-D regression gate: encode a fixed handful of content-pinned corpus
# photos, score mean ΔE00, and compare it against the checked-in baseline
# tools/comparison/baselines/rd-gate.json. Deliberately small — a few images, no
# codec baselines — so CI can run it on every push; `compare:rd` and the sweeps
Expand Down
552 changes: 402 additions & 150 deletions spec/EXPERIMENTS.md

Large diffs are not rendered by default.

14 changes: 11 additions & 3 deletions spec/RATIONALE.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,8 +5,10 @@ considered and rejected. Normative text lives in [`README.md`](README.md); this
file records the *evidence*.

Sweep numbers come from the decision tables produced by `mise run sweep <config>`
(configs in `tools/comparison/sweeps/`, results in
`tools/comparison/output/sweeps/`) on the **tune split** of the expanded corpus
(configs in `tools/comparison/sweeps/`; results were written to the gitignored
`tools/comparison/output/sweeps/` when this file's figures were taken, and are
now committed under `tools/comparison/results/`, for the current corpus only)
on the **tune split** of the expanded corpus
(74 images: 43 synthetic + 31 curated photos; the 32-image holdout —
Kodak24 + held-out curated photos — is reserved for validating winners, per the
pre-registered rule below). The curated set grew from 26 to 39 photographs in
Expand Down Expand Up @@ -458,7 +460,13 @@ encoding.
corpus. The split immediately quantified the damage: mean ΔE00 6.48
(CI 4.6–8.4) on the tune split vs **11.36 (CI 10.3–12.5) on the
never-tuned holdout**. All v1 experiments tune on `tune` and validate on
`holdout` under the pre-registered ≥3%-with-guards rule.
`holdout` under the pre-registered ≥3%-with-guards rule. (The holdout
figure and its interval match the **v1** row of the version table above,
11.364 [10.30, 12.45], rather than its v0.6 row, 11.312 [10.25, 12.39],
which is what this bullet describes. No record says which run produced
them, and the corpus both were taken on has since been replaced, so neither
can be re-derived; read the pair as the size of the gap, not as v0.6's
score. `EXPERIMENTS.md` §11.12 records how the holdout has been used since.)

## Future work / open questions
Explicitly unresolved, so nothing evaluated-in-thought silently disappears:
Expand Down
27 changes: 16 additions & 11 deletions spec/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -1508,10 +1508,12 @@ compatibility** with the v0.6 bitstream. The framing changes are:

Together the four constants-level and encoder-side changes above are worth **−3.72% mean
ΔE00** at the default tier on a never-tuned holdout split, with SSIMULACRA2, Butteraugli and DSSIM
all improving. `spec/EXPERIMENTS.md` §8 records the measurements and what was rejected —
including the equal-quality byte saving this paragraph used to quote (the optimized 32-byte
encode matching the v0.6 constants at 40 bytes), which §8.3 withdrew: it was read off a
ladder of the *pre-adoption* signal path, and no current sweep produces one.
all improving. `spec/EXPERIMENTS.md` §8 records the measurements and what was rejected.
The equal-quality byte saving this paragraph used to quote (the optimized 32-byte encode
matching the v0.6 constants at 40 bytes) was withdrawn for want of a pre-adoption ladder on
the current corpus, and has since been re-measured in `EXPERIMENTS.md` §7.12: on the
holdout split the two are indistinguishable (a match, not an improvement), and on the tune
split the 40-byte v0.6 encode is measurably better, so the saving is at most that figure.

The DCT, OKLAB color pipeline, ℓ2-ball candidate set, µ-law quantizer, decode-aware DC
search, and gamut handling are **inherited from the v0.6 algorithm** (now parameterized by
Expand Down Expand Up @@ -1639,11 +1641,14 @@ SSIMULACRA2 off size-matched WebP. Below 48 bytes no general codec can produce o
all, which is the compact and default tiers' whole argument.

Codes 3 and 4 are **not** justified that way, and this specification does not claim they
are. Measured at equal bytes on the same corpus, WebP overtakes ChromaHash somewhere
around 300 bytes, and by ~1.6 kB even uncoded RGB565 pixels score better than code 4
(`EXPERIMENTS.md` §2, §3). Entropy coding would recover roughly 4% — nowhere near the
20–40% gap, and it would cost the O(1) length check that *is* this format's validity
check (§2.6).
are. Measured at equal bytes on the same corpus (`EXPERIMENTS.md` §11.14, holdout split),
WebP overtakes ChromaHash on ΔE00 between 193 and 411 bytes. At ~1.6 kB AVIF scores
5.471 mean ΔE00 against code 4's 6.768, and WebP and JPEG also beat code 4 on every
metric. Even uncoded RGB565 pixels edge code 4 on ΔE00 there, 6.570, though not on
SSIMULACRA2 or Butteraugli, and not on the tune split, where code 4 wins (§2).
Entropy coding would recover a few percent (§7.13: 1.6% at 32 B, 4.8% at 108 B) — well
short of that gap, and it would cost the O(1) length check that *is* this format's
validity check (§2.6).

They are kept for the operational properties they share with the rest of the format, and
those are real: no codec dependency, no decoder CVE surface, no container or metadata
Expand Down Expand Up @@ -1672,12 +1677,12 @@ is a few hundred bytes and the target is intentionally low-pass.
| Gaborish | Small post-decode smoothing convolution | **Re-evaluated, still off** — the decode-side synthesis window (`window_weights`, a Hann taper, disabled by default). `EXPERIMENTS.md` §12.2–§12.3 measured it at codes 1 and 2 with artifact metrics that did not exist when v0.6 rejected it: it removes up to 73% of the invented structure and costs ΔE00 and SSIMULACRA2 monotonically, failing the guards at every strength. At code 2 the lightest taper is statistically free on ΔE00 and fails on SSIMULACRA2 alone. |
| Edge-preserving filter (EPF) | Adaptive deringing loop filter | **Reject** — a blurred placeholder has few edges to preserve. |
| DC image + DC predictors | Separate DC plane with spatial predictors | **N/A** — chromahash has a single average-color DC per channel, already chosen by the decode-aware DC search (§10.3). |
| **Quantization weighting matrices (HVS/CSF)** | Frequency-dependent quant step | **Evaluate / adopt** — a frequency-shaped bit allocation generalizes the existing two-tier `AcLayout` L split; cheap and on-trend with HVS sensitivity. |
| **Quantization weighting matrices (HVS/CSF)** | Frequency-dependent quant step | **Built as scalefactor bands; below threshold, not adopted** — `EXPERIMENTS.md` §11.9 measured a frequency-shaped scale split on the current corpus: the best arm is −0.13% ΔE00, far under the ≥3% retune rule and unable to pay for signalling it. This row read "evaluate / adopt" until that sweep ran, and is kept as the prediction it scored against. |
| **Entropy coding (rANS + context modeling + clustering, HybridUint tokens)** | Adaptive entropy coding of quantized coefficients | **Highest-impact v0.8+** — fixed-width µ-law leaves the most on the table; many high-frequency coefficients quantize to zero and would cost almost nothing under an entropy coder, raising the quality ceiling per byte. Heaviest to make bit-exact across all language bindings (incl. the hand-written TS decoder) and it trades away the fixed-per-tier length, so it is deferred deliberately. |
| Coefficient ordering / scan + nonzero context | Frequency-ordered scan, run/EOB modeling | **Already frequency-ordered** — the top-K isotropic selection is exactly this; pairs naturally with entropy coding when added. |
| Patches / splines / dots | Reference repeated elements / smooth gradients / point sources | **Reject** — no repeated elements or point sources in a placeholder; the DCT already models smooth gradients compactly. |
| Noise synthesis | Add a per-image perceptual noise model at decode | **Low-priority option** — a few bits of noise amplitude could add cheap perceptual texture; minor. |
| **Progressive / responsive passes** | DC-first, then refinement passes (embedded scalability) | **Compelling v0.8 direction** — make higher tiers *embedded* (the default-tier bytes are a prefix of code 2, etc.) so one stored hash serves both an instant preview and an on-demand detailed render. Constrains the layout but is very LQIP-appropriate. |
| **Progressive / responsive passes** | DC-first, then refinement passes (embedded scalability) | **Built; an operational feature, not a quality one** — `EXPERIMENTS.md` §7.11 implemented embedded tiers (interleaved AC codes, any prefix decodable). The first 32 bytes of a 108-byte hash score 4.21% worse ΔE00 than a native 32-byte encode, so one stored hash can serve both sizes at that price. This row read "compelling v0.8 direction" until that sweep ran; whether the operational gain is worth ~4% is a v0.8 decision, not a measurement. |
| Upsampling (2×/4×/8×) | Store small, upsample at decode with a fixed kernel | **Already covered** — the DCT renders at any target size and `decodeCapped` band-limits (§6.4, §11.3). |

**Summary of the roadmap, as written for v1 — and how it scored.** The four directions
Expand Down
2 changes: 1 addition & 1 deletion tools/comparison/results/adopted-defaults-holdout.json
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@
}
},
"provenance": {
"rev": "c070e1794d867bc69f45b79ece0cb3c9bf5abc65",
"rev": "e3104757d19ea3b0a00f932ab3ec3e7ff937895d",
"dirty": false,
"dirtyPaths": [],
"iqaCli": "iqa-cli 1.2.1",
Expand Down
2 changes: 1 addition & 1 deletion tools/comparison/results/adopted-defaults.json
Original file line number Diff line number Diff line change
Expand Up @@ -19,7 +19,7 @@
}
},
"provenance": {
"rev": "c070e1794d867bc69f45b79ece0cb3c9bf5abc65",
"rev": "e3104757d19ea3b0a00f932ab3ec3e7ff937895d",
"dirty": false,
"dirtyPaths": [],
"iqaCli": "iqa-cli 1.2.1",
Expand Down
Loading
Loading