Summary
libpng_oracle::encode takes each TextChunk's payload as a Rust &str and hands its bytes
straight to libpng:
pub struct TextChunk<'a> {
pub keyword: &'a str,
pub text: &'a str,
pub kind: TextKind<'a>,
}
A Rust str is UTF-8. libpng writes a tEXt/zTXt payload verbatim, and PNG §11.3.3.2 says a
tEXt text string "is interpreted according to the Latin-1 character set [ISO_8859-1]" (§11.3.3.3
makes an inflated zTXt identical to it). So any code point above U+007F is written as its UTF-8
pair — é (U+00E9) becomes C3 A9, which a conforming reader renders é.
The consequence is that the oracle cannot generate a fixture for the one thing gamut-png's text
writer has to get right. A differential test would agree with a gamut encoder that made the same
mistake, which is precisely the defect class an oracle exists to catch.
Proposal
In tooling/libpng-oracle/src/lib.rs, encode the payload for the character set the chunk carries
before handing it to libpng:
TextKind::Text / TextKind::ZTxt: map each char to its Latin-1 byte (u8::try_from on the
code point), and return an error — or panic, as the dev oracle does elsewhere — for a character
Latin-1 cannot represent, rather than silently writing UTF-8;
TextKind::ITxt: keep UTF-8, which is what §11.3.3.4 specifies.
The keyword needs the same treatment in all three: §11.3.3.1 binds it to Latin-1 everywhere.
Why not now
tooling/ was outside PR #550's manifest. That PR implemented the §11.3.3.1/§11.3.3.2 repertoire
rules on the encoder side and pinned them with exact-byte assertions on the chunk payload; this is
the differential test those assertions stand in for.
Related: #502 (a warning count, so a chunk libpng dropped is observable).
Summary
libpng_oracle::encodetakes eachTextChunk's payload as a Rust&strand hands its bytesstraight to libpng:
A Rust
stris UTF-8. libpng writes atEXt/zTXtpayload verbatim, and PNG §11.3.3.2 says atEXttext string "is interpreted according to the Latin-1 character set [ISO_8859-1]" (§11.3.3.3makes an inflated
zTXtidentical to it). So any code point above U+007F is written as its UTF-8pair —
é(U+00E9) becomesC3 A9, which a conforming reader rendersé.The consequence is that the oracle cannot generate a fixture for the one thing gamut-png's text
writer has to get right. A differential test would agree with a gamut encoder that made the same
mistake, which is precisely the defect class an oracle exists to catch.
Proposal
In
tooling/libpng-oracle/src/lib.rs, encode the payload for the character set the chunk carriesbefore handing it to libpng:
TextKind::Text/TextKind::ZTxt: map eachcharto its Latin-1 byte (u8::try_fromon thecode point), and return an error — or panic, as the dev oracle does elsewhere — for a character
Latin-1 cannot represent, rather than silently writing UTF-8;
TextKind::ITxt: keep UTF-8, which is what §11.3.3.4 specifies.The keyword needs the same treatment in all three: §11.3.3.1 binds it to Latin-1 everywhere.
Why not now
tooling/was outside PR #550's manifest. That PR implemented the §11.3.3.1/§11.3.3.2 repertoirerules on the encoder side and pinned them with exact-byte assertions on the chunk payload; this is
the differential test those assertions stand in for.
Related: #502 (a warning count, so a chunk libpng dropped is observable).