:strict character-range scan: bytes, not characters - #148
Merged
Merged
Conversation
An ASCII byte is legal or not by itself, so the scan decodes only a byte at or above 0x80. A span of 16 bytes or less is checked as two 8-byte machine words, without a loop; a longer one 64 bytes at a time, with vector instructions. Verdict, message and the error on malformed UTF-8 are unchanged. Text-only `:strict` 7.5 → 0.8 ms; the plain document's 11 ms of `:strict` fall within the GC spread. Assisted-by: Claude (Anthropic)
Table 7 re-measured on Julia 1.13.0: `:strict` on the character data alone 0.81 ms from 7.0, and on the document within the GC spread. The well-formedness section says how the scan reads a span, the libxml2 comparison is redone, and the CHANGELOG gets its entry. The section (8) note of profile.jl no longer claims a cost proportional to the text share. Assisted-by: Claude (Anthropic)
|
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #148 +/- ##
==========================================
+ Coverage 96.61% 96.67% +0.06%
==========================================
Files 15 15
Lines 2597 2649 +52
==========================================
+ Hits 2509 2561 +52
Misses 88 88
🚀 New features to boost your workflow:
|
The comment said the character-range scan scales with the document's text share. It costs by the span it reads, sixteen bytes at a time on a long text and as two 8-byte words on a short one. Assisted-by: Claude (Anthropic)
"Sixteen bytes at a time" and "two 8-byte words" read as the same thing. The long path is a loop that checks 64 bytes at a time with vector instructions; the short path, 16 bytes or less, is two 8-byte words and no loop at all. PERFORMANCE, CHANGELOG and the profile.jl header now say so. Assisted-by: Claude (Anthropic)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #143.
The
:strictcharacter-range scan decoded every character of every text, attribute value, comment, CDATA section and processing-instruction body, about 1 ns per byte. It now reads the bytes. An ASCII byte is legal or not by itself, so only a byte at or above 0x80 is decoded, by the string's ownChariteration. The verdict, the message and the error on malformed UTF-8 are those of the loop it replaces.Two paths, by span length. A span of 16 bytes or less, an attribute value or the white space between two tags, is checked as two 8-byte words without a loop. A longer span is checked 64 bytes at a time, in a loop the compiler vectorises.
Section (8) of
benchmarks/profile.jlon Julia 1.13.0, medians, GC share in parentheses:parse(…, Node; wellformed = …):structural:strict:structural:strict:structural:strictAllocations are unchanged:
:strictallocates what:structuraldoes, and a guard intest/test_allocations.jlholds the scan to zero on both paths, for aString, aSubStringand aStringView.The issue asked for
plain :strictwithin 2 ms ofplain :structural; measured: 5.3 ms between the medians, all of it within the 6.9 ms GC share of the:strictrun, so net of GC the two do not separate. What remains is a cost per span, not per byte: the document has 623,553 spans, half of them white space of a few bytes between tags, and each one costs the call and the two-word test.Tests: the loop the scan replaces is recopied as the specification and compared with the scan on every code point, on 18,420 byte sequences of an alphabet covering lead, continuation, control and boundary bytes, and on 28,908 placed cases. Each case runs on a
String, on aSubStringthrough both paths and on aStringView.Docs: Table 7 and the well-formedness paragraphs of
PERFORMANCE-v0.4.mdre-measured on Julia 1.13.0; CHANGELOG entry.