fix(pdf): honor ActualText and skip ToUnicode FFFD - #90
Conversation
ToUnicode U+FFFD no longer overwrites Encoding Differences. Marked-content ActualText (inline, named Properties, and StructTreeRoot MCID) replaces the shown glyphs. closes #81
Deploying with
|
| Status | Name | Latest Commit | Updated (UTC) |
|---|---|---|---|
| ✅ Deployment successful! View logs |
mdgate-demo | 41d8720 | Aug 27 2026, 12:04 PM |
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using high effort and found 1 potential issue.
Autofix Details
Bugbot Autofix prepared a fix for the issue found in the latest run.
- ✅ Fixed: Parent ActualText repeats on child MCIDs
- Parent /ActualText is now bound to the first descendant MCID and later child sequences are suppressed, so the replacement emits once.
You can send follow-ups to the cloud agent here.
There was a problem hiding this comment.
Stale comment
Left a non-blocking comment. Cursor Bugbot finished as skipped and reported one unresolved Medium finding (parent ActualText repeating on child MCIDs), so this run did not approve. Human review is needed; no reviewer was assigned because the only assignable user is the PR author.
Sent by Cursor Approval Agent: Pull Request Router and Approver
Parent /ActualText replaces the whole structure element. Bind it to the first descendant MCID and suppress the rest so child sequences do not repeat the replacement.
There was a problem hiding this comment.
Stale comment
Left a non-blocking comment. Cursor Bugbot finished as skipped and left one unresolved Medium finding on the current head (parent ActualText can drop text), so this run did not approve. Human review is needed; no reviewer was assigned because the only assignable user is the PR author.
Sent by Cursor Approval Agent: Pull Request Router and Approver
Share one ActualText binding across descendant MCIDs and emit it once. An empty or nested first MCID no longer drops the replacement or later glyphs.
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
Bugbot Autofix prepared a fix for the issue found in the latest run.
- ✅ Fixed: Nested MCIDs emit ActualText too early
- Shared glyph capture now lives on StructActual and ActualText emits only after the last open frame that uses that bind, so nested MCIDs keep the outer start position.
You can send follow-ups to the cloud agent here.
Reviewed by Cursor Bugbot for commit bcb4577. Configure here.
There was a problem hiding this comment.
Stale comment
Left a non-blocking comment. Cursor Bugbot finished as skipped and left one unresolved Medium finding on the current head (nested MCIDs emit ActualText too early), so this run did not approve. Human review is needed; no reviewer was assigned because no assignable reviewer other than the PR author is available.
Sent by Cursor Approval Agent: Pull Request Router and Approver
Share glyph capture on the parent StructActual and emit only after the last open frame that uses it. Nested child MCIDs no longer replace text at the inner span.



Summary
#80 is already merged. It closed #71 by keeping
/Artifacttext. The leftover is #81: after that merge,01030000000001.pdfstill starts with3�4 Yarrowinstead of gold314/YARROW.Confirmed. The
314run is Type1GKDCHH+Brill-Roman. ToUnicode maps0x13toU+FFFDand overwrites/Differences/one.SP. The producer also wraps that glyph in/Span << /ActualText (UTF-16BE "1") >> BDC.@mdgate/pdfignored both.YARROWvsYarrowis the small-cap half of the same font. Glyph-name suffix mapping from #88 already turns/Y.c2scand/a.smcpinto uppercase. This PR keeps that path when ToUnicode is missing or FFFD.Changes
U+FFFD, so Encoding Differences names (one.SP,Y.c2sc) remain./ActualTextfrom inline BDC dicts, named/Properties, andStructTreeRootMCID entries. Replace the shown glyphs; keep artifact graphics skipped.Test plan
bunx vitest run packages/pdf/test/encoding.test.ts packages/pdf/test/layout.test.tsbunx vitest run packages/pdf/test(169 pass; images needs dist)bun run lintbun test(565 pass)Note
Medium Risk
Changes core PDF text decoding and marked-content extraction paths; regressions could affect markdown output for many PDFs, but scope is limited to font mapping and accessibility metadata, not security or I/O.
Overview
Improves PDF-to-markdown text extraction when producers lie in the glyph stream or ToUnicode maps codes to U+FFFD.
ToUnicode / encoding: Entries that are empty or only replacement characters are no longer applied to the font cmap or treated as successful decode results, so Encoding
/Differencesglyph names (e.g.one.SP, small-cap names) and other fallbacks can still produce the intended Unicode.Marked content
/ActualText: While a marked-content span carries actual text (inline BDC dict, named/Properties, orStructTreeRootvia MCID—including shared parentActualTextacross sibling/nested MCIDs), shown glyphs are not emitted; layout is tracked andActualTextis output once when the span closes./Artifacthandling is refactored to a frame stack but behavior stays the same for skipping decorative graphics.Tests cover FFFD + Differences, all ActualText sources, and small-cap Differences mapping.
Reviewed by Cursor Bugbot for commit 41d8720. Bugbot is set up for automated code reviews on this repo. Configure here.