Skip to content

A stencil needs decoded bytes, not compressed ones - #20

Merged
tannevaled merged 1 commit into
mainfrom
fix/stencil-needs-decoded-bytes
Aug 27, 2026
Merged

A stencil needs decoded bytes, not compressed ones#20
tannevaled merged 1 commit into
mainfrom
fix/stencil-needs-decoded-bytes

Conversation

@tannevaled

Copy link
Copy Markdown
Contributor

decodeImage asks whether an image is a mask before it asks whether the filter
chain finished, so a mask whose filter stopped at an image format nothing here
decodes was handed to the stencil path still compressed — and drawn. A stencil
is one bit a pixel; compressed bytes painted through it are noise in the shape
of nothing at all.

This is not a rare shape. 273 of the image masks in the 1 633 real forms carry
an encoded filter: 236 CCITT and 9 JBIG2. Fifty-one first pages of real forms
were showing that noise, and it took decoding faxes in the reader
(go-pdfkit/reader#15) to make it visible: those pages came out with less ink
afterwards, because most of a form is white and the noise was not.

The remaining nine are JBIG2, which nothing here decodes. Until something does,
the honest answer is the one the rest of decodeImage already gives: "the image
is not drawn rather than drawn wrong."

The test draws a mask of eight zero bytes — as samples, a solid black stencil,
so anything drawn at all is visible — once with /Filter /JBIG2Decode and once
with no filter, and requires the first to draw nothing and the second to draw.
Against the parent commit the first fails with "drawn = true, want false".

100% statement coverage, go vet clean, nine cross-compile targets.

decodeImage asks whether an image is a mask before it asks whether the filter
chain finished, so a mask whose filter stopped at an image format nothing here
decodes was handed to the stencil path still compressed — and drawn. A stencil
is one bit a pixel; compressed bytes painted through it are noise in the shape
of nothing at all.

This is not a rare shape. 273 of the image masks in the 1 633 real forms carry
an encoded filter: 236 CCITT and 9 JBIG2. Fifty-one first pages of real forms
were showing that noise, and it took decoding faxes in the reader
(go-pdfkit/reader#15) to make it visible: those pages came out with *less* ink
afterwards, because most of a form is white and the noise was not.

The remaining nine are JBIG2, which nothing here decodes. Until something does,
the honest answer is the one the rest of decodeImage already gives: "the image
is not drawn rather than drawn wrong."

The test draws a mask of eight zero bytes — as samples, a solid black stencil,
so anything drawn at all is visible — once with /Filter /JBIG2Decode and once
with no filter, and requires the first to draw nothing and the second to draw.
Against the parent commit the first fails with "drawn = true, want false".

100% statement coverage, go vet clean, nine cross-compile targets.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@tannevaled
tannevaled merged commit 475f700 into main Aug 27, 2026
1 check passed
@tannevaled tannevaled mentioned this pull request Aug 27, 2026
tannevaled added a commit that referenced this pull request Aug 27, 2026
reader v0.5.0 decodes /CCITTFaxDecode. 67 of the 1 633 real forms in the corpus
carry a fax and 6 of their pages have nothing else on them, so those pages drew
blank; requiring an older reader here would leave them blank for anyone who
asked only for this package, since minimum version selection gives them the
version this go.mod names.

fax_test.go guards it with a fax rather than a version string: one Group 4 row,
three white pixels then five black, coded in horizontal mode and built in the
test so nobody else's scan enters the repository. Against reader v0.4.2 it fails
with "the black part of the fax is not black: (6,4) = 255,255,255,255".

Measured through this package over the first page of all 1 633 real forms, 69
pages change: one from blank to drawn (fr-cerfa/cerfa_11818, 0 inked pixels to
28 989), fourteen with more ink, and fifty-one with less — those were showing
their own compressed bytes as a stencil, which #20 stopped and this finishes.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
tannevaled added a commit that referenced this pull request Aug 27, 2026
reader v0.6.0 stopped flateDecode from lying: a damaged stream's inflated
prefix used to come back with a nil error, so a caller could not tell a clean
decode from a fragment. Fixing that makes the strict Decode this package used
return nothing where it used to return the fragment — so taking the new reader
without changing anything here is a REGRESSION, and the measurement says so:

  plots/P12positive01.pdf     188 770 inked pixels -> 0     (a blank page)
  uk-govuk/apply-for-help-...  48 455 -> 38 695
  uk-govuk/form-n1 (x3)        42 347 -> 40 771, and two like it

Every stream is now read through DecodeStreamRecovering, and every one of those
pages comes back exactly as it was: 188 770 -> 0 -> 188 770. The recovery was
already happening; it was happening by accident, inside a function that claimed
success. Now it is asked for.

WHY THIS IS THE RIGHT BEHAVIOUR AND NOT MERELY THE OLD ONE

263 streams in 212 of the 1 633 real forms — 13.0% of them — cannot be decoded
cleanly. This package already draws a page as far as it got when its time runs
out; a damaged stream is that same situation arriving another way, and half a
figure is worth more to somebody reading a form than none of it.

WHAT MAKES IT SAFE

Bytes that no filter decoded are never painted. reader.Decoded keeps them in
Undecoded, set INSTEAD of Data rather than beside it, so reading Data cannot
reach them — a type doing the work rather than a flag anyone has to remember.
That guarantee is why this change is possible at all: the same reader release
found that a fax with a bad /Columns had been putting its encoded bytes where a
caller would paint them, which is precisely the defect #20 had just fixed here
from the other side.

Two cases are still refused rather than salvaged, and the tests say so: a page
whose /Contents is filtered as an image, which would run a JPEG as operators;
and a stream whose filter nothing can apply, which yields no decoded bytes at
all.

MEASURED, THE THREE STATES SEPARATED

reader v0.5.0, then v0.6.0 alone, then v0.6.0 with this change, over the first
page of all 1 633 real forms and 1 334 arXiv files, each configuration run
twice and identical to itself both times:

  real forms   v0.5.0 -> v0.6.0: 6 pages change, 5 of them with less ink
               v0.6.0 -> +this:  4 change, all 4 back to what v0.5.0 drew
               v0.5.0 -> +this:  2 change, both by under thirty pixels

  arXiv        v0.5.0 -> v0.6.0: 4 change, one of them to a blank page
               v0.6.0 -> +this:  3 change, the blank page back to 188 770
               v0.5.0 -> +this:  2 change

One arXiv figure disagrees in the other direction — 2512.03312/Fig3.pdf draws
24 913 pixels under v0.5.0, 38 525 under v0.6.0 alone, and 24 913 again with
this. Which of the two is right has not been established, and it is written
down here rather than left out.

100% statement coverage, go vet and -race clean, nine cross-compile targets.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant