Skip to content

fix(xlsx): read a worksheet serialized with a namespace prefix - #1927

Merged
developer0hye merged 4 commits into
mainfrom
fix/xlsx-worksheet-namespace
Sep 27, 2026
Merged

developer0hye merged 4 commits into
mainfrom
fix/xlsx-worksheet-namespace

Conversation

@developer0hye

Copy link
Copy Markdown
Owner

File submission policy

Before attaching or committing files, read the submission policy.

  • Any submitted sample files or attachments satisfy the submission policy, or none are submitted. The synthetic fixture pair is already tracked from the defect report; only derived evidence is added.

Summary

A worksheet may bind the SpreadsheetML URI to a prefix instead of making it the default namespace. Excel writes and accepts both spellings, and the two packages are equivalent XML — identical expanded element names, attributes and text. The dependency's worksheet reader matched literal e.name() values such as row and headerFooter, so s:row and s:headerFooter matched nothing: every sheet element was skipped, the CLI reported success, and the page came out blank with body cells, header and footer all gone.

The fix is a streaming NsReader adapter in the dependency (developer0hye/umya-spreadsheet#14, upstream MathNya/umya-spreadsheet#371, still open) that resolves each event's namespace and hands the legacy sub-readers the unqualified local name when, and only when, the URI is SpreadsheetML. An element in a foreign namespace takes an internal prefix so the legacy parser cannot mistake its local name for a worksheet element. The adapter buffers one event, so worksheet streaming survives, and the stored source XML is untouched.

Only the lock's source line moves, from 10efd229 to fa350180 on the same fix/panic-safety-v2 fork branch. git diff 10efd229 fa350180 in the fork is exactly that adapter, its wiring in src/reader/xlsx.rs and src/reader/xlsx/worksheet.rs, and tests/worksheet_namespaces.rs — 4 files, no other dependency. It is our own fork, not a registry release, so the 7-day crates.io age check does not apply; every command here ran with --locked.

The measured outcome is stronger than "the content comes back": the prefixed and default-namespace packages now convert to a byte-identical PDF (8ec36306…), and the default-namespace output is unchanged from main. Namespace serialization no longer reaches the output at all.

Related issue

Related: #1803

Testing

  • TDD, red first. On the old pin worksheet_namespace_prefixes_preserve_body_and_header_footer fails with left: [] against the five expected body values for the s: package, while the default-namespace control in the same test passes. After the pin it passes.
  • The test pins the rule, not the reported input: besides the committed s: fixture it generates four more serializations from the default-namespace control — x:, ns0:, spreadsheetml: and row:. The last would still fail a substring-matching reader. Each is checked through both parse and parse_streaming, for all five body values, the header, and the footer's three sections including its page-number and total-pages fields.
  • Pre-fix output of the reported source: 1 byte of extracted text against the ground truth's 114. Post-fix: 121 bytes, and shasum of the prefixed and control outputs is identical.
  • cargo test --locked --workspace --profile ci: 3335 lib tests and every integration suite pass. (A first run reported 8 failures in ooxml-package; they were a stale test binary in the shared target directory, baked with a removed worktree's CARGO_MANIFEST_DIR. Forcing a rebuild turned them green — no source change.)
  • cargo fmt --all -- --check clean; cargo clippy --locked --workspace --all-targets clean on stable and on 1.97 (CI's newer toolchain).
  • compare_text_layer.py GT vs after: no codepoint-class delta, normalized content identical. GT vs before: -6 spaces and 114 characters against 0.
  • Dependency side: the fork carries 157 passing tests including three new namespace integration tests; upstream Resolve worksheet namespace prefixes in streaming readers MathNya/umya-spreadsheet#371 passed all five checks and is open.
  • Documentation freshness reviewer: PASS before each of the two commits.

Visual impact

  • No rendered PDF change
  • Rendered PDF change or visual evidence added
  • Reason: the prefixed source rendered a blank page and now renders its full content.

Visual audit

  • Issue: XLSX worksheet namespace prefixes silently produce a blank page #1803
  • Fixture: tests/visual_audits/issue-1803/prefixed.xlsx (unmodified)
  • Page(s): 1
  • Renderer and DPI: pdftoppm at 150 DPI for the stored evidence and the cluster census, ImageMagick 5% fuzz for the diff mask, mutool draw -F trace for every position quoted below
  • Evidence mode: fix
  • Layout audit report: assets/bugfixes/issue-1803/layout-audit.json
  • Render cluster reports: assets/bugfixes/issue-1803/render-clusters-page-1.json
  • Reference exporter differences: None
  • Fine-detail threshold: 0.5pt
  • Layout audit page count: Pass
  • Layout audit text flow: Pass
  • Layout audit visible fills: Pass
  • Layout audit rectangle geometry: Pass
  • Layout audit large shifts: Pass
  • Layout audit fine shifts: XLSX: a sheet header baseline sits 1.6pt above Excel's #1731, XLSX: a 14-15pt bottom-aligned cell in a 19.5pt fixed row is centred as a tight row where Excel bottom-seats it #1721, XLSX: a uniform-run header/footer section starts one point right of our anchor #1728, Tracker: sub-materiality visual deviations (revisit in bulk) #1874
  • New follow-up issues found in this audit: None
  • Model vision findings: The full GT and output pages, the 150 DPI 5% diff and all five matched-shift crops were opened and inspected, on the before render as well as the after. The before page is completely white — not one glyph, rule or fill anywhere on it, which is why its only visible "element" is the page box. In the after page every native element is back and in the same reading order: the centred header Quarterly report, the Quarterly summary title, the Research/120 and Development/240 rows with the values right-aligned in the same column, and the three footer sections Internal, Generated by Example Reporting System and 1 / 1. The crops show matching Arial faces at matching weight with no italic, no underline, no clipping and no overflow on either side; nothing on this synthetic page is bold, coloured, rotated, dashed or filled, and neither side draws a single rule or rectangle (the layout audit counts 0 rectangles on both sides), so the hairline and emphasis inventories are empty by construction rather than by assumption. In the diff mask every lit pixel is a thin outline hugging a glyph edge; no region is solidly filled, which is what a sub-2pt displacement of otherwise identical text looks like. Measuring the traces rather than the raster: the body is uniformly 1.0pt high (baselines 84/113/128 against 85/114/129) with x matching exactly at 53.000 and 246.000; the centred header is 1.6pt high (50.400 against 52.000) and 0.487pt right (267.487 against 267.000); the footer baseline matches exactly at 767.000 on all three sections, with Internal 1.000pt left (50.000 against 51.000), the centred section 0.495pt left (204.505 against 205.000) and 1 / 1 0.596pt right (540.596 against 540.000) on a 1.39% wider run. Every one of these is also present in the default-namespace control, whose output this change leaves byte-identical, so none of them is introduced here.
  • GT: assets/bugfixes/issue-1803/gt.jpg
  • Before: assets/bugfixes/issue-1803/before.jpg
  • After: assets/bugfixes/issue-1803/after.jpg
  • Native: None
  • Compare: assets/bugfixes/issue-1803/compare.jpg

Visual comparison

GT Before After
GT Before After

Required inspection

  • Rendered all evidence at 150 DPI or higher
  • Stored progressive JPEG quality 86 assets with metadata stripped
  • Used Codex/Claude vision to inspect the full GT/output pages, diff, and matched crops
  • Inspected matched region crops at full resolution
  • Ran compare_layout.py --audit --fine-shift PT and dispositioned every fine/large text-instance shift, rectangle geometry deviation, painted-text visibility mismatch, and visible-fill occlusion
  • Ran compare_render.py --cluster-report PATH --strict-clusters and dispositioned every material 5% fuzz diff cluster by explicit ID
  • Inventoried hairlines and border dash styles
  • Inventoried font weight, italic, and underline emphasis

Deviation audit

Check Result
Page count/order Matches GT: one page on both sides, before and after.
Element presence Fixed: the before page is blank; the after page matches GT with 5/5 lines matched, 0 missing, 0 extra.
Position/size Remaining: #1731 (centred header baseline 1.6pt high), #1721 (all five body instances 1.0pt high in fixed 15pt rows), #1728 (Internal, a plain uniform-run left footer section, starts at 50.000 against native 51.000 on a 0.7in margin), #1874 (centred header 0.487pt right, centred footer 0.495pt left, 1 / 1 0.596pt right on a 1.39% wider run). 0 large shifts at 5pt; 8 fine shifts at 0.5pt, all listed here.
Rotation/flip Matches GT: no rotated or flipped content on either side.
Fill Matches GT: neither side paints a fill; 0 visible-fill occlusions.
Stroke/border Matches GT: neither side emits a stroke, rule or dash pattern; the layout audit counts 0 rectangles on both sides.
Shape outline geometry Matches GT: the page carries no shape geometry.
Text content Fixed: compare_text_layer.py reported 114 GT characters against 0 before; after, no codepoint-class delta and identical normalized content.
Font family/weight/style Matches GT: ArialMT on both sides for all nine runs, regular weight throughout, no italic and no underline on either side.
Text color Matches GT: every run is black on both sides.
Alignment Matches GT: Quarterly summary, Research and Development at x 53.000 on both sides, 120 and 240 at 246.000 on both, header and centre footer centred, 1 / 1 right-aligned. Residual horizontal offsets are the #1728 and #1874 rows above.
Line/paragraph spacing Matches GT: row pitch worst 1.00pt, which is the uniform #1721 body offset against the unshifted footer rather than a pitch difference — the three body baselines are 29pt and 15pt apart on both sides. Width worst 1.4%, the 1 / 1 run of #1874.
Clipping/overflow Matches GT: nothing clipped or overflowing on either side.

Checklist

  • Commits include a Signed-off-by line
  • PR scope contains one root cause
  • Remaining converter or harness deviations each reference an open issue

🤖 Generated with Claude Code

developer0hye and others added 4 commits September 21, 2026 01:16
Signed-off-by: Yonghye Kwon <developer.0hye@gmail.com>
A worksheet may bind the SpreadsheetML URI to a prefix instead of making
it the default namespace. Excel writes and accepts both spellings, and
the two packages are equivalent XML: identical expanded element names,
attributes and text.

The dependency's worksheet reader matched literal `e.name()` values such
as `row` and `headerFooter`, so `s:row` and `s:headerFooter` matched
nothing. Every sheet element was skipped, the converter reported success,
and the page came out blank — body cells, header and footer all gone.

The fix is a streaming `NsReader` adapter in the dependency
(developer0hye/umya-spreadsheet#14, upstream MathNya/umya-spreadsheet#371)
that resolves each event's namespace and hands the legacy sub-readers the
unqualified local name when, and only when, the URI is SpreadsheetML.
Elements in a foreign namespace take an internal prefix so the legacy
parser cannot mistake their local names for worksheet elements. It buffers
one event, so worksheet streaming is preserved, and the stored source XML
is untouched. Only the lock's `source` line moves; the fork delta over the
previous pin is that adapter, its wiring and its tests.

The regression test pins the rule rather than the reported input: it reads
the committed `s:`-prefixed fixture and four more prefixes generated from
the default-namespace control, including `row:`, which a substring-matching
reader would still get wrong. Both the whole-document and the streaming
parse paths are checked for every body value, the header, and the footer's
three sections with their page-number fields.

The prefixed and default-namespace packages now convert to a
byte-identical PDF, and the default-namespace output is unchanged.

Related: #1803

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Yonghye Kwon <developer.0hye@gmail.com>
Page 1 of the native Excel export, the pre-fix blank output and the
post-fix output for the prefixed worksheet, all at 150 DPI, with the
layout audit and the strict render-cluster report against the same pair.

Every one of the 16 diff clusters is a deviation the default-namespace
path already had; none is introduced here, and the dispositions route
them to #1731, #1721, #1728 and the #1874 tracker.

Related: #1803

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: Yonghye Kwon <developer.0hye@gmail.com>
@developer0hye
developer0hye merged commit b3f0f5c into main Sep 27, 2026
21 checks passed
@developer0hye
developer0hye deleted the fix/xlsx-worksheet-namespace branch September 27, 2026 23:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant