Skip to content

:strict character-range scan: bytes, not characters - #148

Merged
mathieu17g merged 4 commits into
mainfrom
strict-char-scan
Sep 11, 2026
Merged

mathieu17g merged 4 commits into
mainfrom
strict-char-scan

Conversation

@mathieu17g

Copy link
Copy Markdown
Collaborator

Closes #143.

The :strict character-range scan decoded every character of every text, attribute value, comment, CDATA section and processing-instruction body, about 1 ns per byte. It now reads the bytes. An ASCII byte is legal or not by itself, so only a byte at or above 0x80 is decoded, by the string's own Char iteration. The verdict, the message and the error on malformed UTF-8 are those of the loop it replaces.

Two paths, by span length. A span of 16 bytes or less, an attribute value or the white space between two tags, is checked as two 8-byte words without a loop. A longer span is checked 64 bytes at a time, in a loop the compiler vectorises.

Section (8) of benchmarks/profile.jl on Julia 1.13.0, medians, GC share in parentheses:

parse(…, Node; wellformed = …) before after
text-only :structural 0.49 ms 0.50 ms
text-only :strict 7.54 ms 0.81 ms
plain :structural 48.44 ms 49.07 ms
plain :strict 62.98 ms (GC 7.5) 54.32 ms (GC 6.9)
escaped :structural 57.74 ms (GC 3.4) 58.0 ms (GC 3.3)
escaped :strict 78.38 ms (GC 13.4) 64.22 ms (GC 7.8)

Allocations are unchanged: :strict allocates what :structural does, and a guard in test/test_allocations.jl holds the scan to zero on both paths, for a String, a SubString and a StringView.

The issue asked for plain :strict within 2 ms of plain :structural; measured: 5.3 ms between the medians, all of it within the 6.9 ms GC share of the :strict run, so net of GC the two do not separate. What remains is a cost per span, not per byte: the document has 623,553 spans, half of them white space of a few bytes between tags, and each one costs the call and the two-word test.

Tests: the loop the scan replaces is recopied as the specification and compared with the scan on every code point, on 18,420 byte sequences of an alphabet covering lead, continuation, control and boundary bytes, and on 28,908 placed cases. Each case runs on a String, on a SubString through both paths and on a StringView.

Docs: Table 7 and the well-formedness paragraphs of PERFORMANCE-v0.4.md re-measured on Julia 1.13.0; CHANGELOG entry.

An ASCII byte is legal or not by itself, so the scan decodes only a
byte at or above 0x80. A span of 16 bytes or less is checked as two
8-byte machine words, without a loop; a longer one 64 bytes at a
time, with vector instructions. Verdict, message and the error on
malformed UTF-8 are unchanged. Text-only `:strict` 7.5 → 0.8 ms; the
plain document's 11 ms of `:strict` fall within the GC spread.

Assisted-by: Claude (Anthropic)
Table 7 re-measured on Julia 1.13.0: `:strict` on the character data
alone 0.81 ms from 7.0, and on the document within the GC spread. The
well-formedness section says how the scan reads a span, the libxml2
comparison is redone, and the CHANGELOG gets its entry. The section (8)
note of profile.jl no longer claims a cost proportional to the text share.

Assisted-by: Claude (Anthropic)
@codecov-commenter

codecov-commenter commented Sep 11, 2026 •

Copy link
Copy Markdown

⚠️ Please install the 'codecov app svg image' to ensure uploads and comments are reliably processed by Codecov.

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 96.67%. Comparing base (b9c005c) to head (dd0b05b).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files

Impacted file tree graph

@@            Coverage Diff             @@
##             main     #148      +/-   ##
==========================================
+ Coverage   96.61%   96.67%   +0.06%     
==========================================
  Files          15       15              
  Lines        2597     2649      +52     
==========================================
+ Hits         2509     2561      +52     
  Misses         88       88              
Files with missing lines Coverage Δ
src/parse.jl 99.54% <100.00%> (+0.13%) ⬆️
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

The comment said the character-range scan scales with the document's
text share. It costs by the span it reads, sixteen bytes at a time on
a long text and as two 8-byte words on a short one.

Assisted-by: Claude (Anthropic)
"Sixteen bytes at a time" and "two 8-byte words" read as the same
thing. The long path is a loop that checks 64 bytes at a time with
vector instructions; the short path, 16 bytes or less, is two 8-byte
words and no loop at all. PERFORMANCE, CHANGELOG and the profile.jl
header now say so.

Assisted-by: Claude (Anthropic)
@mathieu17g
mathieu17g marked this pull request as ready for review September 11, 2026 21:07
@mathieu17g
mathieu17g merged commit ad977cd into main Sep 11, 2026
13 checks passed
@mathieu17g
mathieu17g deleted the strict-char-scan branch September 11, 2026 21:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

The :strict character-range scan alone takes longer than libxml2's full parse of the same text

2 participants