Preserve letters outside the BMP in every case conversion - #71
Merged
Merged
Conversation
The three regexes tested Unicode categories per UTF-16 code unit. A letter
outside the Basic Multilingual Plane is a surrogate pair whose halves are both
categorised as Surrogate rather than as a letter, so `[^\p{L}0-9]` matched both
halves and the letter was replaced with spaces and then collapsed away. Every
conversion silently dropped it: "set 𐐀abc" became "SetAbc".
Two further faults shared that root cause. `IsAllCaps` stripped both surrogate
halves before testing, so an astral lowercase letter was skipped entirely and
"HELLO 𐐨" reported true. The splitter's letter-to-non-letter alternative read
the high surrogate as a non-letter and inserted a spurious word boundary, so
"abc𝒳def" split into two tokens.
Replace all three regexes with walks by code point. The conversions now treat
an astral letter exactly as they treat its BMP analogue, and astral characters
that are not letters stay separators as before.
Behaviour for BMP input is unchanged: the new code was compared against the
regex implementation over 12,191 inputs across all six public methods with no
differences. Four of the five new tests fail without this change.
Fixes #70
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RUsrYfSAuSQ7VWws7vzsr
|
This was referenced Sep 24, 2026
This was referenced Sep 28, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



Fixes #70.
The bug
All three regexes tested Unicode categories per UTF-16 code unit. A letter outside the Basic
Multilingual Plane is a surrogate pair, and each half is
UnicodeCategory.Surrogaterather than aletter — so
[^\p{L}0-9]matched both halves,ReplaceNonAlphaNumericWithSpace's predecessorturned the letter into two spaces, and
CollapseSpacesremoved them. The letter was gone.Measured at
main(5a20678), .NET SDK 10.0.401, Linux:"set 𐐀abc".ToPascalCase()SetAbcSet𐐀abc"𐐀abc".ToSnakeCase()abc𐐨abc"abc𝒳def".ToMacroCase()ABC_DEFABC_𝒳DEF"HELLO 𐐨".IsAllCaps()truefalseTwo further faults shared the root cause, and both are fixed here:
IsAllCapsskipped astral letters. It stripped non-letters with[^\p{L}]— which removedboth surrogate halves — then tested the remainder, so an astral lowercase letter was never seen.
ToTitleCaseuses that result to decide whether to lowercase before title-casing.(?<=[\p{L}])(?=[^\p{L}])read the highsurrogate as a non-letter, so
abc𝒳defsplit into two tokens.Changes
All three regexes are replaced by walks over code points:
ReplaceNonAlphaNumericWithSpace— keeps letters and ASCII digits, replaces everything elsewith a space.
SplitOnCaseChange/IsWordBoundary— the regex's three alternatives, expressed as explicitconditions over code points. The comments name the case each one covers.
IsAllCaps— walks code points directly rather than stripping and re-testing.ToTitleCasegainsEnsure.NotNull(input), which preserves theArgumentNullExceptionthatRegex.Replaceused to throw and satisfies CA1062 now that the regex call is gone. The#if NET7_0_OR_GREATERsplit disappears with the regexes, so all six target frameworks now runidentical code.
Astral characters that are not letters — emoji, symbols — remain separators, exactly as before.
Verification
original regex implementation over 12,191 inputs (all three-part combinations of 23 atoms
covering ASCII case, digits, separators, punctuation, and accented/Greek/eszett letters, plus 24
realistic identifiers) across all six public methods — no differences. The same harness was
then re-run with an astral input added and correctly reported the four expected differences, so
it is not vacuously passing.
fail, 25 pass. The fifth pins that non-letter astral characters are still dropped, and passes
either way by design.
net10.0,net9.0,net8.0,netstandard2.0,netstandard2.1,0 warnings, 0 errors.
Not in this PR
"MAX_SIZE".ToPascalCase()isMaxSizebut"set MAX_SIZE".ToPascalCase()isSetMAXSIZE— thesame token converting two ways depending on what precedes it. That falls out of
ToTitleCasetesting
IsAllCapson the whole string to decide whether words are acronyms, which is adocumented heuristic rather than an oversight. Changing it is a design decision about acronym
policy, so it is left alone and noted on #70's sibling discussion in
ktsu-dev/Coder#71.🤖 Generated with Claude Code
https://claude.ai/code/session_014RUsrYfSAuSQ7VWws7vzsr
Generated by Claude Code