Issues #583 and #584 were filed from the first revision of the codec benchmark harness in PR #558
(issue #163). Three later passes on that PR corrected the harness and withdrew several of the
figures those issues quote. Both issues remain valid findings — the defects they name are real
and reproduce — but the numbers, one localisation argument and one verification command in them are
now wrong, and #583's verification command matches zero benchmarks.
The run that filed #583 and #584 is unattended and its contract forbids editing, commenting on,
labelling or closing an existing issue under any authority; it may only file new ones. So the
correction is filed here rather than applied in place. A human with write access should fold this
into #583 and #584 and close this issue.
The replacement text lives in crates/gamut-dng/STATUS.md, section "Codec benchmark harness
(#163)", and in PR #558.
#583 — gamut-dng: lossless-JPEG decode is ~60x slower than the reference implementation
The finding stands. lossless_jpeg::decode_symbol scans the whole 256-entry code table once
per candidate bit length, the BitReader refills one byte at a time, and the effect is one and a
half to two orders of magnitude. The Mechanism and Likely fix sections are read from the source and
are unaffected.
Withdrawn: the two Deflate ratios (1.24x, 1.25x). They are not properties of gamut against the
SDK. On the */deflate rows the SDK's inflate is the system libz, which the oracle links
dynamically (-lz) because the SDK includes <zlib.h> unconditionally, and which libz the loader
resolves is not even a property of the machine: cargo exports every build script's native search
path on the runner's library path, and gamut-dng dev-depends on libtiff-oracle, which builds a
libz.so of its own under target/. Measured both ways on the same box, same fixture bytes:
| arm |
stock zlib 1.3.1 (from target/) |
zlib-ng 2.3.3 (platform) |
decode_dng cfa/deflate gamut |
1.044, 0.925, 1.052 ms |
1.053, 1.051, 1.038 ms |
decode_dng cfa/deflate adobe-sdk |
0.974, 0.991, 1.113 ms |
0.830, 0.829, 0.837 ms |
gamut's arm does not move; the reference arm moves by 1.2-1.3x, which carries the ratio from 0.94
to 1.23-1.26 — across 1.0, with no defect in either implementation. Any Deflate figure has to
travel with the library it was taken against.
Withdrawn, separately: those rows are not a gamut-versus-SDK comparison at all.
gamut-deflate is deliberately encoder-only, so on the Deflate rows gamut inflates with
miniz_oxide. Neither arm's inflate is gamut-authored. What distinguishes them is that
miniz_oxide is pinned by Cargo.lock to one version and one checksum and the system libz is
pinned by nothing.
Withdrawn: the localisation argument. #583 argues "that every other scheme lands within 1.25x
and only this one is 60x out localises the cost to the entropy decoder". It rests on the two
withdrawn ratios, and the rows that falsify it are not in the issue at all — decode_dng's
uncompressed rows, uncorrected, are 2.1-2.5x (cfa/uncompressed) and 1.7-1.8x
(linear-raw/uncompressed).
Replacement, which does not depend on any other row: the isolating evidence is the codestream
pair, decode_lossless_jpeg. It carries no container asymmetry — same bare SOF3 stream in, same
samples out, one counter for all three arms — and its one residual bias, the FFI export path, is
bounded by the third arm adobe-sdk-no-export well below the effect. Across eight runs, on both
photometries and in both arm orders, no fastest-sample ratio is below 34x. The SDK's own
lossless-JPEG decoder is built here from the committed SDK source, so unlike the Deflate rows this
one does not depend on what the machine has installed.
Withdrawn: the precision of "~60x" / "57-59x". That was a two-run figure. Reproduced across
eight runs the medians span 42-92 and the fastest-sample ratios 34-63. Nothing brings the effect
near parity; the defensible form is the 34x floor.
Corrected: the Verification command matches no benchmark.
cargo bench -p gamut-dng --bench codec, comparing decode_lossless_jpeg_gamut against
decode_lossless_jpeg_adobe_sdk in the same run
There are no benchmarks by those names. A later pass merged each pair into one benchmark taking
the implementation as a divan argument, so that the two arms run back to back under the same
instantaneous load rather than minutes apart. The working command is:
cargo bench -p gamut-dng --bench codec -- decode_lossless_jpeg
and the ratio is the cfa gamut row against the cfa adobe-sdk row inside that one benchmark
(likewise linear-raw). -- --sortr name runs the gamut arm first; a published ratio is the mean
of one run each way. The adobe-sdk-no-export row in the same output is the export-path bound.
Correctness fencing is unchanged: the existing SOF3 conformance tests and the Adobe differential
must stay green.
#584 — gamut-dng: lossless-JPEG CFA encode expands the payload past uncompressed
The finding stands, and it is a byte quantity that reproduces exactly. cfa/uncompressed
writes 541 440 bytes and cfa/lossless-jpeg 618 800, so turning compression on grows the file by
14.3 %. The mechanism (one full-width component, so predictor 1 differences a red photosite against
its green neighbour), the spec citation (DNG 1.7.1.0 p. 20) and the reshape remedy are all
unaffected.
Stale: the quoted fixture table. #584 quotes columns IFD0 preview and of raw at 147 456 /
37.5 % / 12.5 %. The harness no longer prints those columns, and the model behind them was doubled:
the preview is stored at 8 bits, but the decoder surfaces every sub-image as
SubImageData::Decoded(Vec<u16>), so the volume it materialises is 294 912 bytes — 75 % of a
16-bit CFA frame and 25 % of a LinearRaw one. The table now prints raw samples, encoded DNG,
of raw, preview, decode vol., / raw and correction. The encoded DNG column and every
percentage #584's argument rests on are unchanged.
Restated: the framing invites a wrong number. #584's Observation quotes whole-file
percentages (137.7 %, 157.4 %) and its remedy table quotes codestream bytes against the same
raw-sample denominator (119.7 % -> 91.5 %, "23.5 % smaller"). Chaining the two gives a reader a
~33 % file saving that does not exist. Against one denominator throughout: both files carry an
identical 148 224 bytes of preview and directory, so the reshaped file is 359 888 + 148 224 =
508 112 bytes, 129.2 % of raw — 6.2 % smaller than the uncompressed file, not 33 %.
Rejected: the proposed Verification. #584 proposes pinning the encoder to the sizes the fixture
table prints. It should not: the margin is a property of frame-uniform synthetic gains — one gain
per CFA colour across the whole frame, which is exactly what makes the interleaved-component
reshape win so cleanly — and pinning an encoder requirement to a single synthetic fixture is the
failure a benchmark harness exists to avoid. Re-measure on the real-camera corpus (mise run fetch-dng-samples, then mise run test-dng-real) and gate there if anything is to be gated.
Refs #163, #558, #583, #584.
Issues #583 and #584 were filed from the first revision of the codec benchmark harness in PR #558
(issue #163). Three later passes on that PR corrected the harness and withdrew several of the
figures those issues quote. Both issues remain valid findings — the defects they name are real
and reproduce — but the numbers, one localisation argument and one verification command in them are
now wrong, and #583's verification command matches zero benchmarks.
The run that filed #583 and #584 is unattended and its contract forbids editing, commenting on,
labelling or closing an existing issue under any authority; it may only file new ones. So the
correction is filed here rather than applied in place. A human with write access should fold this
into #583 and #584 and close this issue.
The replacement text lives in
crates/gamut-dng/STATUS.md, section "Codec benchmark harness(#163)", and in PR #558.
#583 — gamut-dng: lossless-JPEG decode is ~60x slower than the reference implementation
The finding stands.
lossless_jpeg::decode_symbolscans the whole 256-entry code table onceper candidate bit length, the
BitReaderrefills one byte at a time, and the effect is one and ahalf to two orders of magnitude. The Mechanism and Likely fix sections are read from the source and
are unaffected.
Withdrawn: the two Deflate ratios (1.24x, 1.25x). They are not properties of gamut against the
SDK. On the
*/deflaterows the SDK's inflate is the system libz, which the oracle linksdynamically (
-lz) because the SDK includes<zlib.h>unconditionally, and which libz the loaderresolves is not even a property of the machine:
cargoexports every build script's native searchpath on the runner's library path, and
gamut-dngdev-depends onlibtiff-oracle, which builds alibz.soof its own undertarget/. Measured both ways on the same box, same fixture bytes:target/)decode_dng cfa/deflate gamutdecode_dng cfa/deflate adobe-sdkgamut's arm does not move; the reference arm moves by 1.2-1.3x, which carries the ratio from 0.94
to 1.23-1.26 — across 1.0, with no defect in either implementation. Any Deflate figure has to
travel with the library it was taken against.
Withdrawn, separately: those rows are not a gamut-versus-SDK comparison at all.
gamut-deflateis deliberately encoder-only, so on the Deflate rows gamut inflates withminiz_oxide. Neither arm's inflate is gamut-authored. What distinguishes them is thatminiz_oxideis pinned byCargo.lockto one version and one checksum and the system libz ispinned by nothing.
Withdrawn: the localisation argument. #583 argues "that every other scheme lands within 1.25x
and only this one is 60x out localises the cost to the entropy decoder". It rests on the two
withdrawn ratios, and the rows that falsify it are not in the issue at all —
decode_dng'suncompressed rows, uncorrected, are 2.1-2.5x (
cfa/uncompressed) and 1.7-1.8x(
linear-raw/uncompressed).Replacement, which does not depend on any other row: the isolating evidence is the codestream
pair,
decode_lossless_jpeg. It carries no container asymmetry — same bare SOF3 stream in, samesamples out, one counter for all three arms — and its one residual bias, the FFI export path, is
bounded by the third arm
adobe-sdk-no-exportwell below the effect. Across eight runs, on bothphotometries and in both arm orders, no fastest-sample ratio is below 34x. The SDK's own
lossless-JPEG decoder is built here from the committed SDK source, so unlike the Deflate rows this
one does not depend on what the machine has installed.
Withdrawn: the precision of "~60x" / "57-59x". That was a two-run figure. Reproduced across
eight runs the medians span 42-92 and the fastest-sample ratios 34-63. Nothing brings the effect
near parity; the defensible form is the 34x floor.
Corrected: the Verification command matches no benchmark.
There are no benchmarks by those names. A later pass merged each pair into one benchmark taking
the implementation as a divan argument, so that the two arms run back to back under the same
instantaneous load rather than minutes apart. The working command is:
and the ratio is the
cfa gamutrow against thecfa adobe-sdkrow inside that one benchmark(likewise
linear-raw).-- --sortr nameruns the gamut arm first; a published ratio is the meanof one run each way. The
adobe-sdk-no-exportrow in the same output is the export-path bound.Correctness fencing is unchanged: the existing SOF3 conformance tests and the Adobe differential
must stay green.
#584 — gamut-dng: lossless-JPEG CFA encode expands the payload past uncompressed
The finding stands, and it is a byte quantity that reproduces exactly.
cfa/uncompressedwrites 541 440 bytes and
cfa/lossless-jpeg618 800, so turning compression on grows the file by14.3 %. The mechanism (one full-width component, so predictor 1 differences a red photosite against
its green neighbour), the spec citation (DNG 1.7.1.0 p. 20) and the reshape remedy are all
unaffected.
Stale: the quoted fixture table. #584 quotes columns
IFD0 previewandof rawat 147 456 /37.5 % / 12.5 %. The harness no longer prints those columns, and the model behind them was doubled:
the preview is stored at 8 bits, but the decoder surfaces every sub-image as
SubImageData::Decoded(Vec<u16>), so the volume it materialises is 294 912 bytes — 75 % of a16-bit CFA frame and 25 % of a
LinearRawone. The table now printsraw samples,encoded DNG,of raw,preview,decode vol.,/ rawandcorrection. Theencoded DNGcolumn and everypercentage #584's argument rests on are unchanged.
Restated: the framing invites a wrong number. #584's Observation quotes whole-file
percentages (137.7 %, 157.4 %) and its remedy table quotes codestream bytes against the same
raw-sample denominator (119.7 % -> 91.5 %, "23.5 % smaller"). Chaining the two gives a reader a
~33 % file saving that does not exist. Against one denominator throughout: both files carry an
identical 148 224 bytes of preview and directory, so the reshaped file is 359 888 + 148 224 =
508 112 bytes, 129.2 % of raw — 6.2 % smaller than the uncompressed file, not 33 %.
Rejected: the proposed Verification. #584 proposes pinning the encoder to the sizes the fixture
table prints. It should not: the margin is a property of frame-uniform synthetic gains — one gain
per CFA colour across the whole frame, which is exactly what makes the interleaved-component
reshape win so cleanly — and pinning an encoder requirement to a single synthetic fixture is the
failure a benchmark harness exists to avoid. Re-measure on the real-camera corpus (
mise run fetch-dng-samples, thenmise run test-dng-real) and gate there if anything is to be gated.Refs #163, #558, #583, #584.