Skip to content

perf(rust): a decode stage breakdown, bound like §1 - #89

Merged
justin13888 merged 8 commits into
masterfrom
perf/80-decode-stage-breakdown
Sep 25, 2026
Merged

justin13888 merged 8 commits into
masterfrom
perf/80-decode-stage-breakdown

Conversation

@justin13888

@justin13888 justin13888 commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

spec/PERFORMANCE.md §1 breaks encode down by stage, but decode had no breakdown, so §12.1 item 4 was sized by reading decode.rs. This PR measures decode stage by stage, records the result in a committed baseline, binds a new §1.1 table to it through verify:benchmark, and restates item 4 from the measurement.

What the measurement found. Shares of one decode of a 100×100 gradient (see §1.1):

stage t1 natural t4 natural t4 capped 32×32
gamma_lut 66% 0% 4%
render 29% 99% 55%
selection 4% 1% 38%
  • Tier 1 (the default): two thirds of a decode goes to build_gamma_lut. It rebuilds a 4096-entry table (one portable_pow per entry for sRGB/P3) on every call, to render 1024 pixels. The table depends only on the output gamut, so building it once would give byte-identical output. This PR measures that lever; it does not implement it (§12.1 item 4(a)).
  • Tier 4: the render loop takes 99% of the decode.
  • Tier 4 capped: sorting the candidate list takes 38% of the decode.

Changes, by path:

  • rust/src/decode.rs: stage! marks cover render_at_size from start to end (header, selection, ac_dequant, window_filter, cos_tables, gamma_lut, render). Without bench-internals they compile to nothing. The render loop gets one mark on purpose, and its doc comment says why.
  • rust/src/encode.rs: pub(crate) use stage; lets the decoder use the encoder's macro and recorder. The recorder's doc says the two stay separable because neither path calls the other. Also moves analyze()'s doc comment back onto analyze(); it had ended up on the stage_timing module.
  • rust/src/lib.rs: the stage_timing re-export's doc now covers decode as well.
  • rust/Cargo.toml: adds the bench_decode_stages example with required-features = ["bench-internals", "spec-vectors"], and updates the bench-internals feature comment.
  • rust/examples/bench_decode_stages.rs (new): before timing anything, it decodes all 19 shared decode vectors (natural and capped) through the instrumented build and requires the spec's exact bytes. This checks against bytes the build did not produce itself, which is stronger than a repeatability check. It then times the natural or capped decode of a gradient at a given tier, stage by stage, and prints meta.* lines (raster, hash length, vectors checked).
  • mise-tasks/benchmark/decode-stages (new): mise run benchmark:decode-stages SRC_W SRC_H TIER ITERS [CAP]. It runs the example and records one cell.
  • tools/benchmark/record_stages.py: adds a --decode mode that writes perf-decode-stages.json (schema chromahash-perf-decode-stages/1, keys WxH-tTIER-natural|capCWxCH). A decode cell is refused if it is missing its totals, raster, hash length or a vector check of at least 1. The dirty probe now excludes both recorder outputs, so §1 and §1.1 can be recorded back to back. Any other change still records dirty: true.
  • tools/benchmark/test_record_stages.py: six new tests cover the decode mode:
    • the cell's fields and shares
    • refusal without the vector check or totals
    • a malformed CAP
    • a schema mismatch
    • the sibling artifact not dirtying either table
    • any other change still dirtying the record
  • tools/comparison/baselines/perf-decode-stages.json (new): the three §1.1 cells, recorded at 8f5f4d1 from a clean tree. The file is exactly the recorder's output.
  • tools/comparison/src/verify-benchmark-core.ts:
    • DECODE_STAGES_BASELINE, parseDecodeStages (shares the refusal logic with parseStages), checkDecodeStagesProvenance and checkStageRowCoverage.
    • checkDecodeStagesProvenance requires one clean commit, vectorsChecked ≥ 1, a key that matches the cell's fields, iters ≥ 1, whole_decode ≥ stage_sum, and shares equal to their own ns over whole_decode. A non-number share also fails.
    • checkStageRowCoverage fails a table that omits a recorded stage or names one that was never recorded.
    • checkProseClaims now fails a claim whose cell or stage is missing. It used to skip it silently, which also affected §1's claims.
  • tools/comparison/src/verify-benchmark.ts: binds §1.1:
    • A resolver that finds nothing throws MissingCell instead of counting as unbound.
    • An unusable baseline exits 1.
    • A missing §1.1 table fails; before, it was only listed as SKIP.
    • A missing column fails.
    • The number of checked cells must equal rows × columns, so an n/a cell cannot pass silently.
    • Ten prose figures are bound, in §1.1, §12.1 item 4 and §12.3.
  • tools/comparison/src/metric-selftest.ts: four new blocks (B14–B17) cover the functions above.
  • spec/PERFORMANCE.md:
    • New §1.1 table and prose, with provenance.
    • §12.1 item 4 restated in two parts sized from §1.1: (a) the LUT, (b) flattening the cos tables plus SIMD. Line references updated.
    • §12.3's decode bullet rewritten.
    • §10 item 4 gains one sentence pointing at the LUT.
    • §0's reproducibility row and re-measure procedure now include benchmark:decode-stages.
  • .mise.toml: the test:benchmark description and comments now mention §1.1 and the decode recorder.

Validation

All at 2ca3423 unless noted, on this host:

  • cargo fmt --check (rust/): pass
  • cargo clippy --all-targets --features full -- -D warnings: pass
  • cargo clippy --all-targets --features full,bench-internals -- -D warnings: pass. CI does not run this.
  • cargo test --features full: pass (144 + 14 + 5 + 2)
  • cargo build --no-default-features / cargo test --no-default-features: pass
  • cargo +1.85 build --all-targets: pass (MSRV). cargo +1.85 build --all-targets --features bench-internals also passed at 500e78e; rust/ has not changed since.
  • ./tools/ci/check-versions.sh: pass. python3 spec/validate.py: pass
  • uvx ruff format --check tools/benchmark, uvx ruff check tools/benchmark: pass
  • python3 -m unittest discover -s tools/benchmark -p 'test_*.py': pass (13 tests)
  • pnpm --prefix tools/comparison run format:check / lint / build: pass
  • node tools/comparison/dist/metric-selftest.js: pass
  • node tools/comparison/dist/verify-claims.js: pass. verify-experiments.js --strict: pass. verify-sweep-labels.js: pass. corpus-licenses.js --check: pass
  • node tools/comparison/dist/metric-reference-test.js: pass. determinism-check.js: pass. rd-gate.js: PASS (0.00% drift)
  • node tools/comparison/dist/verify-benchmark.js: exit 1, pre-existing. The six disagreements are the §2 tier-3/4 encode and §3 128/1024 TBD placeholders that §0 lists and master also carries. ci-comparison.yml runs this step with continue-on-error. Every §1.1 value, and all 10 of its prose figures, checked clean. verify-benchmark.js --section 1.1 exits 0.
  • A negative probe: with an n/a cell and an extra idct row temporarily in §1.1, the gate reported the missing cell, the count shortfall (27 expected, 23 checked) and the unrecorded row. The edit was then reverted.
  • convco check origin/master..HEAD: no errors.

Replicate runs (not committed)

On the same loaded host (load average 9–19), I re-ran bench_decode_stages five times per column without recording. Ranges of the share per stage:

  • t1 natural: gamma_lut 62.3–65.6%, render 29.0–32.0%. One earlier 200-iteration run gave 55% / 38%. In absolute terms gamma_lut held at 213–219 µs while render ranged 95–158 µs.
  • t4 natural: render 98.8–99.0%.
  • t4 capped: selection 37.8–38.3%, render 54.9–55.5%.

§1.1 is published at whole-percent precision and says in place that the t1 split between gamma_lut and render is its least stable figure. The finding that the LUT outweighs the render loop at tier 1 held in every run.

Coverage gaps

  • CI never builds with bench-internals, so neither the decode stage! marks nor bench_decode_stages is compiled or run there. They ran locally: clippy, MSRV, and the three recorded runs.
  • mise-tasks/benchmark/decode-stages has no automated test. The recorder it calls does.
  • verify-benchmark.ts's process-level exits have no test: the missing-baseline exit 1, the §1.1 cell-count check, the missing-table and missing-column failures. The core functions they call are covered by B14–B17. The count, row and cell failures were exercised once by hand (above).
  • No absolute decode time is published. As perf: a decode stage breakdown, bound like §1 #80 says, that needs a quiet host.

Risks and rollout

  • There is no behaviour change in any shipped build. The marks compile to nothing without bench-internals, and the golden decode tests, the spec vectors and rd-gate all pass.
  • checkProseClaims is stricter for §1 as well. All 13 §1 claims still resolve and pass.

Issue

Closes #80

Decisions taken

  1. How finely to mark the decode render loop.
    Taken: one render mark around the whole O(w·h·K) loop. §1.1 and item 4(b) say its split between the inverse DCT and the colour conversion is not measured.
    Rejected: a timer per pixel, or splitting the loop in two. Either would time a decoder the crate does not ship: the timer overhead is comparable to per-pixel work at tier 1, and a second loop changes memory traffic.
    Reverses: add stage! marks inside the loop in rust/src/decode.rs, then add the rows to §1.1 and to the recorded cells.
  2. Where the decoder gets stage! from.
    Taken: pub(crate) use stage; in encode.rs, reusing the one thread-local recorder. This matches the manifest's "stage!/stage_timing visibility" and is the smallest change.
    Rejected: moving stage_timing into its own module, a larger move with no behavioural gain.
    Reverses: move the module and macro to rust/src/stage_timing.rs and update the two use sites and lib.rs.
  3. What "produces the shipped bytes" means for the decode example.
    Taken: before timing, decode all 19 shared decode vectors and require the spec's bytes (spec-vectors is a required feature). The artifact records the count and the gate requires at least 1.
    Rejected: comparing each timed decode only with its own warm-up, as bench_stages.rs does. That shows repeatability, not agreement with the shipped decoder. The timed decode is still also checked for repeatability.
    Reverses: drop check_spec_vectors and vectorsChecked, and the spec-vectors requirement.
  4. Which decodes form §1.1's columns.
    Taken: a 100×100 gradient (§2's fixture) at tier 1 natural, tier 4 natural and tier 4 capped 32×32. These are the default, the costliest (§2), and §2's mitigation. Iterations are 2000, 20 and 200, so each column times a comparable interval.
    Rejected: §1's 512×512 sources, because decode cost depends on tier and raster, not source size. Also rejected: tiers 2 and 3, to keep three columns like §1.
    Reverses: record more cells with benchmark:decode-stages and add entries to DECODE_STAGE_COLUMNS.
  5. Precision, and whether to publish absolute times.
    Taken: whole percent, shares only.
    Rejected: §1's 0.1% precision, because replicate runs on this loaded host moved the tier-1 shares by several points. Also rejected: a total row in ms, because perf: a decode stage breakdown, bound like §1 #80 reserves absolute figures for a quiet host.
    Reverses: re-record on a quiet host, add decimals or a total row, and bind the total to ns.whole_decode.
  6. One recorder or two.
    Taken: a --decode mode in record_stages.py, which is the one file in the manifest. The dirty probe excludes both recorder outputs.
    Rejected: excluding only the file being written. Then §1 and §1.1 could not be recorded back to back as §0's procedure does, and a measurement artifact is not a source input anyway.
    Reverses: narrow RECORDER_OUTPUTS to the output being written.
  7. How to store the cap.
    Taken: {"width", "height"} or null, so biome format leaves the recorder's output byte-identical.
    Rejected: a [w, h] pair, which biome rewraps, so the committed file would fail format:check or differ from the recorder's output.
    Reverses: change the cap line in record_stages.py and decodeCellKey, and add baselines/** to biome's ignore list.
  8. Where the table lives.
    Taken: ### 1.1 under §1.
    Rejected: a new top-level section, which would renumber §2–§12, and those numbers are cross-referenced throughout.
    Reverses: move the heading and change the binding's section.
  9. How to record the LUT lever.
    Taken: fold it into §12.1 item 4 as part (a).
    Rejected: a new numbered row, which would renumber items 8 and 9, and "Finalize API surface for v1 #8" is referenced in §12.3.
    Reverses: split item 4(a) into its own row and renumber.
  10. What §1.1's resolvers do when nothing matches.
    Taken: throw MissingCell, so every §1.1 cell is required.
    Rejected: return null, which checkTable counts as a deliberately unbound pass. That is how §1 behaves, and it is the defect class this series has been removing.
    Reverses: return null in the §1.1 resolver.
  11. How checkProseClaims treats a claim whose cell or stage is missing.
    Taken: it fails, for §1's claims too.
    Rejected: skipping silently, which meant a claim bound to the wrong cell was never checked.
    Reverses: restore the two continues in checkProseClaims.
  12. Whether the gate checks that the recorded rev's rust/ tree equals HEAD's.
    Taken: no. §1.1 matches §1, which checks one clean commit across columns but not tree equality to HEAD.
    Rejected: recording git rev-parse HEAD:rust and comparing. Every rust/ change on master would then fail verify:benchmark until both stage tables were re-recorded, a policy §1 has not adopted.
    Reverses: record rustTree in record_stages.py and compare it in checkOneCleanCommit.
  13. Paths touched beyond the planned manifest.
    Taken: rust/Cargo.toml (the example needs an [[example]] entry), tools/comparison/src/metric-selftest.ts (the only harness that tests verify-benchmark-core.ts), and PERFORMANCE.md §0 and §10. §0's procedure would otherwise leave §1.1 unrecorded, and §10 item 4 would otherwise contradict §1.1.
    Rejected: leaving them. The example would not build, the new gate functions would be untested, and the document would contradict itself.
    Reverses: revert those hunks.

…ecode_stages example

The decoder's render_at_size now carries stage! marks covering it end to
end (header, selection, ac_dequant, window_filter, cos_tables, gamma_lut,
render), through the same recorder the encoder uses. The macro is made
crate-visible from encode.rs, and expands to nothing without the feature.

bench_decode_stages encodes a gradient at a tier, then times its natural or
capped decode stage by stage. Before timing anything it decodes all 19
shared decode vectors through the instrumented build and requires the
spec's bytes, so the shipped-bytes claim is checked against bytes this
build did not produce, not only against its own warm-up.

Also moves analyze()'s doc comment back onto analyze(); it had been
attached to the stage_timing module.
record_stages.py gains a --decode mode writing perf-decode-stages.json
(schema chromahash-perf-decode-stages/1), one cell per invocation keyed
WxH-tTIER-natural or WxH-tTIER-capCWxCH. A decode cell carries its render
raster, hash length and the number of spec vectors the instrumented build
reproduced, and is refused outright when any of those or its totals are
missing.

The dirty probe now excludes both recorder outputs rather than only the one
being written, so §1 and §1.1 can be recorded back to back before either is
committed; any other change still records dirty: true.
Recorded with benchmark:decode-stages from a clean tree: tier 1 natural
(2000 iterations), tier 4 natural (20) and tier 4 capped 32x32 (200), each
after reproducing all 19 shared decode vectors.
biome format rewraps a short JSON array onto one line, so a [w, h] cap made
the committed baseline either fail format:check or differ from what the
recorder writes. An object is laid out the same by both.
verify:benchmark reads perf-decode-stages.json through parseDecodeStages,
which refuses a missing, unparseable, wrong-schema or empty file the way
parseStages does, and exits 1 when §1.1 is bound but the file is unusable.

checkDecodeStagesProvenance holds every cell to one clean commit, a
recorded spec-vector check, a key that matches its fields, and shares
equal to their own ns over whole_decode. checkStageRowCoverage fails a
table that omits a recorded stage or names one nothing recorded, and a
missing §1.1 table fails rather than being listed as skipped. Ten prose
figures in §1.1, §12.1 item 4 and §12.3 are bound as claims.

checkProseClaims now fails a claim whose cell or stage the baseline does
not hold; it used to skip it silently, for §1's claims as well.
…from it

§1.1 is the decode equivalent of §1: shares of a tier-1 natural, tier-4
natural and tier-4 capped 32x32 decode, bound to perf-decode-stages.json.
It finds the gamma LUT rebuilt on every call at two thirds of a
default-tier decode, the render loop at 99% of a tier-4 one, and the
candidate sort at over a third of a capped tier-4 one.

§12.1 item 4 is restated in two parts sized by that table, §12.3 no longer
says decode is sized by reading decode.rs, §10 item 4 points at the LUT,
and §0's procedure records §1.1 alongside §1.
Re-recorded with the cap written as an object, from a clean tree at the
commit that carries §1.1, and §1.1 quotes that revision. The capped tier-4
render share rounds to 55% in this run, against 56% in the last.
checkTable passes over a cell whose text is not a number, so a §1.1 cell
written as n/a would have been a silent pass. Count what it checked for
§1.1 and fail unless it is every row times every bound column.
@justin13888
justin13888 merged commit ad32d4d into master Sep 25, 2026
25 checks passed
@justin13888
justin13888 deleted the perf/80-decode-stage-breakdown branch September 25, 2026 08:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

perf: a decode stage breakdown, bound like §1

1 participant