fix(pdf): keep footnotes out of body reading order - #84
Merged
Merged
Conversation
Raised markers after punctuation stay with the word. Numbered note blocks are peeled from the page before XY-cut and table detection so they do not interleave with columns or glue onto the last paragraph. fixes #74
Deploying with
|
| Status | Name | Latest Commit | Updated (UTC) |
|---|---|---|---|
| ✅ Deployment successful! View logs |
mdgate-demo | 90928eb | Aug 26 2026, 05:27 AM |
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 4 potential issues.
Bugbot Autofix prepared fixes for all 4 issues found in the latest run.
- ✅ Fixed: Notes glue onto trailing list items
- roleChange and noteStart now close an open list before note emission, so a following note starts a new block instead of continuing the last item.
- ✅ Fixed: Matching page numbers get dropped
- Footer-only lines are skipped as note starts, and bottom page numbers are classified as footers before orphan-marker drop.
- ✅ Fixed: Figure caption exclusion never matches
- The figure/table guard now tests the text after the leading number, so numbered captions are no longer treated as notes.
- ✅ Fixed: Trailing marker scan is overbroad
- Trailing marker collection now requires two or more letters before 1-3 digits at a token end, so Fig.1, p.45, v2, and decimals no longer peel body lists.
You can send follow-ups to the cloud agent here.
Reviewed by Cursor Bugbot for commit 25cac30. Configure here.
Close lists before a note so footnotes do not glue onto the last item. Classify bottom page numbers as footers before dropping orphan markers. Test figure/table captions against the text after the leading number. Limit trailing marker scans to end-of-word digits.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



Summary
Confirmed #74 against the cited opendataloader-bench pages. The PDF converter had no footnote region detector, so raised markers became ordinary digits in the wrong sentence, and numbered note blocks mixed into body columns or glued onto the last paragraph.
This change:
FY2019.⁴,equilibrium?¹²).Verified on the issue examples (
01030000000008,57,86,97,135). Notes stay out of the body. Remaining gaps (a page number that still sits mid-page on 08, a wrapped note line ordered before its start on 135) are outside the mix-into-body failure.fixes #74
Note
Medium Risk
Changes core PDF reading-order and paragraph assembly heuristics; misclassification could still affect some layouts, but scope is limited to the PDF package with broad new test coverage.
Overview
Improves PDF-to-markdown so footnote markers and bottom-of-page note blocks no longer pollute main body text or column reading order.
Superscript merge in
pdf.tsis reworked: small digit runs attach to the nearest preceding text item using position and punctuation rules (e.g.FY2019.⁴,equilibrium?¹²), not only when the parent ends in a letter.Footnote peeling adds exported
peelFootnoteLinesinlayout.ts, which per page infers note regions from superscript markers, smaller type, and numbered note starts (with guards for page numbers, figure captions, and false positives when too much of the page would be classified as notes). Peeled lines are excluded from table detection; markdown flow emits body → notes → footer per page withnote/footerroles so two-column notes do not interleave with columns. Note lines get explicit paragraph breaks and skip heading heuristics so footer notes are not space-joined onto the last body sentence or list item.Unit and integration tests cover peeling heuristics and end-to-end markdown behavior.
Reviewed by Cursor Bugbot for commit 90928eb. Bugbot is set up for automated code reviews on this repo. Configure here.