Skip to content

Tag a scan even without Tesseract - #11

Merged
bbertucc merged 4 commits into
mainfrom
approx-scan
Oct 4, 2026
Merged

bbertucc merged 4 commits into
mainfrom
approx-scan

Conversation

@bbertucc

@bbertucc bbertucc commented Oct 4, 2026

Copy link
Copy Markdown
Member

When Tesseract is missing or fails, a scanned page was left untagged. Now it is tagged from Iris's HTML as usual, with the invisible text flowing down the page in reading order instead of over each word's image. Screen readers get the full structure; only selection, search and magnifier highlights are off. The report says textSource: "approximate" and warns no_text_positions.

🤖 Generated with Claude Code

A screen reader reads tags in order, not by position, so the page
need not be left untagged. Each block starts a new line.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All checks pass. No blocking issue: the approximate layer does not change a pixel, survives the text check (I pushed it to 3500 words a page and to 40×40 pt pages, textPreserved: true throughout), and the structure tree a screen reader reads is correct. Three latent findings.

Non-blocking notes

1. breaks only fires on a top-level block change, so nested blocks still run together. src/tag.ts:186

const breaks = new Set(words.length ? [] : ordered.filter((o, k) => k && o.block !== ordered[k - 1].block).map((o) => o.word));

wordsInOrder (src/html/build.ts:238) numbers block by the top-level child of the page plan, so every word under one grouping element shares a block index and never gets a break. A block's last word has space === false, so the next block's first word lands at zero gap and extraction fuses the two. Tagging scan-300dpi.pdf with PATH="":

  • <main><h1>Parking Permit</h1><p>Residents may apply.</p><h2>Fees</h2><p>Twenty dollars a year.</p></main> → "Parking PermitResidents may apply.FeesTwenty dollars a year." (same for div, section, article)
  • <h1>Rates</h1><table><tr><th>Year</th><th>Fee</th></tr><tr><td>2024</td><td>Twenty</td></tr></table> → "Rates\nYearFee2024Twenty"
  • <ul><li>One</li><li>Two</li></ul> → "• One• Two"

So on the common scan — a table, a list, or anything wrapped in a container — Ctrl+F for Fees or Twenty finds nothing: the tokens are FeesTwenty and 2024Twenty. The comment at src/tag.ts:185, "each block starts a new line, so blocks do not run together", holds only for HTML that is a flat list of blocks, which is exactly the shape of the mixed fixture the new test uses (<h2>Office Address</h2><p>…</p>), so the test passes over this. Breaking on the first word of each Run rather than each top-level block would cover the nested cases.

This is a note, not blocking: before the PR such a page carried no tagged text at all, and readingOrder() over the structure tree is correct in all three cases above, so assistive tech is unaffected.

2. The new PDF/UA-1 claim on an approximate page is never checked by veraPDF. A no-Tesseract scan page is now tagged, so it stops counting as untagged and src/tag.ts:109 claims PDF/UA-1 — confirmed, tagging scan-300dpi.pdf with PATH="" writes <pdfuaid:part>1</pdfuaid:part>. But test/pdfua.test.ts skips the scan fixtures without Tesseract (skip: noVera || (SCANS.has(name) && noOcr)) and CI installs Tesseract (.github/workflows/test.yml:21), so no veraPDF run ever sees this output, while the README says "The tests check each such claim with veraPDF". A corpus case that forces the approximate path would close the gap.

3. Dead default. src/tag.ts:255: breaks = new Set<Word>() — fillPositions has one caller and it always passes breaks.

Accessibility impact: a scan without Tesseract gains the full structure tree, alt text and reading order it previously lost entirely, and screen-reader output is correct; only extracted-text fidelity stays imperfect, and more so than the diff's comment claims wherever the HTML nests its blocks.

… claim

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All five checks pass, and the test job passes on this head, so the new veraPDF case really runs.

My three earlier notes on 7dc5816 are addressed:

  1. Blocks running together — fixed. The three cases I gave now break: <main>…</main> → Parking Permit / Residents may apply. / Fees / Twenty dollars a year., the table → Rates / Year / Fee / 2024 / Twenty, the list → • One / • Two. The comment at src/tag.ts:185 is now accurate.
  2. The unchecked PDF/UA-1 claim — fixed. test/pdfua.test.ts:24 forces the approximate path with env: { PATH: "" }, and .github/workflows/test.yml installs veraPDF with IRIS_REQUIRE_VERAPDF: 1, so it is not skipped in CI.
  3. Dead default — fixed. breaks is a required parameter (src/tag.ts:255).

Non-blocking notes

Inline boundaries still fuse two tokens. src/tag.ts:263

let x = prev ? from.box[2] + (prev.space ?? true ? width(" ") * size : 0) : first.box[0];

splitWords (src/align/words.ts:30) leaves a run's last word at space: false, and addText starts a fresh run after an inline child, so the word before an inline element's close gets a zero gap. Tagging scan-300dpi.pdf with PATH="":

  • <p>See <a href='https://example.com'>the form</a> for details.</p> → "See the formfor details."
  • <p>Call <code>tag</code> now.</p> → "Call tagnow."
  • <dl><dt>Permit</dt><dd>A paper you display.</dd></dl> → "PermitA paper you display." (Lbl/LBody are in INLINE, which the bullet case needs)

Ctrl+F for Permit or for fails on such a page. Not blocking, and I will not raise it again: the structure tree keeps those as separate elements, so a screen reader reads them apart, and the new test's • One link case is correct because the link ends the item. Treating "first word of a new Run" as worth a space, not just a new line, would close it.

I also pushed the new line-breaking at the wrap-around: a 93-cell table and 200 identical one-word paragraphs on the 306×396 pt scan both keep textPreserved: true with every word extractable, so the extra lines overflowing back to the top do not cost text.

Accessibility impact: none negative — a no-Tesseract scan now gets a correct structure tree and reading order, and the only residue is two fused tokens in copy-and-search text at inline boundaries.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

My inline-boundary note on 31b3e4d is fixed: <p>See <a>the form</a> for <code>tag</code>.</p> and <dl><dt>Permit</dt><dd>A paper.</dd></dl> now extract as See the form for tag. and Permit A paper., asserted by test/cli.test.ts:62. All five checks pass. One blocking issue, found in the break logic the earlier commit added.

Blocking

A word in breaks is placed off the right edge, so the text check fails and nothing is written. src/tag.ts:265

if (breaks.has(w) || x + wide > page[2] - 1) {
  x = breaks.has(w) ? first.box[0] : page[0] + 1;

The ordinary wrap sets x = page[0] + 1, and wide is clamped to page[2] - page[0] - 2 (src/tag.ts:262), so a wrapped word always ends at or before page[2] - 1. The break branch instead sets x = first.box[0] — in the approximate path no word is placed, so first is the fallback at page[0] + 10 (src/tag.ts:256) — and then never re-checks the right edge. Any break word wider than page[2] - page[0] - 11 lands partly off the page, which is exactly the loss the comment at src/tag.ts:251 names ("extractors drop text off the page").

Input that reaches it — a no-Tesseract scan whose HTML starts a block with a long token:

{"lang":"en","title":"Fees","pages":[{"sourcePage":1,"html":"<h1>Fees</h1><p>https://example.gov/parking/permits/residential/zone-14/application-form-2024-revised-final.pdf</p><p>After.</p>"}]}
$ env -i PATH="" node src/cli.ts tag --pdf test/fixtures/scan-300dpi.pdf --pages p.json --out o.pdf --report r.json
iris-pdf: text_lost: Page 1 lost text.
exit=2
"verification": { "pixels": "identical-outside-fields", "differingPixels": 0, "textPreserved": false },
"warnings": [ … { "code": "text_lost", "page": 1, "detail": "2 characters of the tagged text are not readable" } ]

No PDF is written. On this PR's base the same input exits 0 with the page left untagged and warned, so the PR turns a warned page into a hard exit 2 with no output at all. The same token placed mid-block — <p>Fees https://example.gov/…-final.pdf After.</p> — takes the ordinary wrap and gives textPreserved: true, exit 0, which isolates the break branch as the cause.

scan-300dpi.pdf is 306×396 pt, so ~57 characters is enough; on a 612 pt page the threshold is ~115, which a printed URL or an Iris-transcribed run-together token reaches. Clamping the break the way the wrap already does — x = page[0] + 1 for breaks too, or recomputing wide against page[2] - x — closes it. A test with an over-wide first word in a block would hold it; the dense cases do not catch this (60 <dt>/<dd> pairs on the same scan keep textPreserved: true, 126 mcids).

Non-blocking notes

wordsInOrder now mutates the tree it is asked to read. src/html/build.ts:242: out.at(-1)!.word.space = true. space is not only layout — src/tag.ts:300 passes it to overlay.word, which draws an extra space glyph (src/pdf/content.ts:141), so this query changes the written content stream on the pdf-text and ocr paths as well, not just the approximate one. It is idempotent and the corpus pixel/text tests pass, so nothing is broken today; setting the flag in listItem/build, where the tree is constructed, would keep the query pure.

Accessibility impact: the structure tree, reading order and alt text an approximate scan gains are correct, but a long token at the start of a block now loses overlay text and aborts the run, so that page ends up with no tagged output at all instead of the warned, untagged page it got before.

…where the tree is built

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All five checks pass. My blocking finding on 73838d9 is fixed, and I verified the fix is really guarded by a test.

Over-wide first word of a block — fixed. src/tag.ts:266

x = breaks.has(w) && first.box[0] + wide <= page[2] - 1 ? first.box[0] : page[0] + 1;

wide is clamped to page[2] - page[0] - 2, so the page[0] + 1 branch always ends at or before page[2] - 1. The input I gave now exits 0 with textSource: "approximate", textPreserved: true, no text_lost. The new case in test/cli.test.ts:62 (<p>${"x".repeat(150)}</p>) holds it: reverting only that line in a copy of the tree makes it fail —

✖ approximately placed blocks do not run together, nested or not
  AssertionError: iris-pdf: text_lost: Page 1 lost text.
  2 !== 0

Tree mutation during a read — fixed. The space flags are now set where the tree is built (src/html/build.ts:146 for <dd>, :218 for leading whitespace), and wordsInOrder only reads.

Non-blocking notes

A <pre> block does not start a new line, so its text can fuse with its neighbours. src/html/build.ts:236

const INLINE = new Set(["Link", "Reference", "Code", "Lbl", "LBody"]);

pre maps to Code (src/html/map.ts:13) alongside code, so a <pre> — a block element — is treated as inline and gets no newLine. With no whitespace in the HTML on either side there is also no space flag, so the tokens join. Tagging scan-300dpi.pdf with PATH="" and <div>Intro<pre>npm ci</pre>After</div><p>Tail.</p>:

["Intronpm ciAfter", "Tail."]

Latent and small: the structure tree keeps the Code element separate, so a screen reader reads it apart, and reachable only through a <pre> written with no surrounding whitespace. Keying the break on the HTML tag, or leaving pre out of the inline set, would close it.

Accessibility impact: a scan without Tesseract now gets the full structure tree, alt text and reading order it previously lost, and no page loses overlay text — the earlier abort on a long first token is gone.

@bbertucc
bbertucc merged commit ca2fad5 into main Oct 4, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant