Resolve worksheet namespace prefixes in streaming readers - #371
Merged
MathNya merged 1 commit intoOct 1, 2026
Merged
Conversation
Signed-off-by: Yonghye Kwon <developer.0hye@gmail.com>
Owner
|
@developer0hye |
developer0hye
added a commit
to developer0hye/office2pdf
that referenced
this pull request
Sep 27, 2026
A worksheet may bind the SpreadsheetML URI to a prefix instead of making it the default namespace. Excel writes and accepts both spellings, and the two packages are equivalent XML: identical expanded element names, attributes and text. The dependency's worksheet reader matched literal `e.name()` values such as `row` and `headerFooter`, so `s:row` and `s:headerFooter` matched nothing. Every sheet element was skipped, the converter reported success, and the page came out blank — body cells, header and footer all gone. The fix is a streaming `NsReader` adapter in the dependency (developer0hye/umya-spreadsheet#14, upstream MathNya/umya-spreadsheet#371) that resolves each event's namespace and hands the legacy sub-readers the unqualified local name when, and only when, the URI is SpreadsheetML. Elements in a foreign namespace take an internal prefix so the legacy parser cannot mistake their local names for worksheet elements. It buffers one event, so worksheet streaming is preserved, and the stored source XML is untouched. Only the lock's `source` line moves; the fork delta over the previous pin is that adapter, its wiring and its tests. The regression test pins the rule rather than the reported input: it reads the committed `s:`-prefixed fixture and four more prefixes generated from the default-namespace control, including `row:`, which a substring-matching reader would still get wrong. Both the whole-document and the streaming parse paths are checked for every body value, the header, and the footer's three sections with their page-number fields. The prefixed and default-namespace packages now convert to a byte-identical PDF, and the default-namespace output is unchanged. Related: #1803 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Signed-off-by: Yonghye Kwon <developer.0hye@gmail.com>
This was referenced Sep 27, 2026
Owner
|
@developer0hye |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Resolve SpreadsheetML namespace prefixes before dispatching worksheet elements to the existing sub-readers. A worksheet using
s:row,s:cand prefixed header/footer elements currently reads successfully but loses its content. The equivalent default-namespace worksheet works.Related: #370, developer0hye/office2pdf#1803.
A streaming
NsReaderadapter presents the existing sub-readers with their expected unqualified SpreadsheetML names. It resolves each element in scope, retains qualified foreign names, and gives foreign-default-namespace elements an internal prefix so their local names cannot be confused with worksheet elements. The original stored XML is unchanged. Normal, lazy/lite and callback-streaming reads use the adapter. No dependencies or public APIs change.Tests
35bbffa4; boths:andspreadsheet:variants lose A1.cargo test --locked: 347 passed; 23 existing ignored tests.cargo clippy --locked --test worksheet_namespaces -- -D warnings: passed, including the library.cargo +nightly-2026-09-11 fmt --all -- --checkandgit diff --check: passed.Performance tradeoff
The adapter preserves streaming memory behavior but adds a namespace-resolution/serialization pass before the existing parser. An optimized public-API probe with 20,000 rows / 40,000 cells measured these median reader times across six fresh-process runs per variant and mode, interleaved in alternating order:
That is approximately 1.6% and 34.1% slower respectively on this fixture. Measurements were on an Apple M2 with 16 GB RAM under variable host load; every repetition, load observation and process resource report was retained. These are host observations, not a performance guarantee. Baseline and patched executables had distinct SHA-256 hashes; the baseline loses all cells on the prefixed benchmark while the patch preserves all 40,000.
Using an adapter keeps namespace handling consistent through nested worksheet readers without changing their shared
Readersignatures. The additional streaming-read cost is a tradeoff of that bounded approach.