Skip to content

Preserve letters outside the BMP in every case conversion - #71

Merged
matt-edmondson merged 1 commit into
mainfrom
claude/issue-70-astral-letters
Sep 24, 2026
Merged

matt-edmondson merged 1 commit into
mainfrom
claude/issue-70-astral-letters

Conversation

@matt-edmondson

Copy link
Copy Markdown
Contributor

Fixes #70.

The bug

All three regexes tested Unicode categories per UTF-16 code unit. A letter outside the Basic
Multilingual Plane is a surrogate pair, and each half is UnicodeCategory.Surrogate rather than a
letter — so [^\p{L}0-9] matched both halves, ReplaceNonAlphaNumericWithSpace's predecessor
turned the letter into two spaces, and CollapseSpaces removed them. The letter was gone.

Measured at main (5a20678), .NET SDK 10.0.401, Linux:

input before after
"set 𐐀abc".ToPascalCase() SetAbc Set𐐀abc
"𐐀abc".ToSnakeCase() abc 𐐨abc
"abc𝒳def".ToMacroCase() ABC_DEF ABC_𝒳DEF
"HELLO 𐐨".IsAllCaps() true false

Two further faults shared the root cause, and both are fixed here:

  • IsAllCaps skipped astral letters. It stripped non-letters with [^\p{L}] — which removed
    both surrogate halves — then tested the remainder, so an astral lowercase letter was never seen.
    ToTitleCase uses that result to decide whether to lowercase before title-casing.
  • The splitter inserted a spurious word boundary. (?<=[\p{L}])(?=[^\p{L}]) read the high
    surrogate as a non-letter, so abc𝒳def split into two tokens.

Changes

All three regexes are replaced by walks over code points:

  • ReplaceNonAlphaNumericWithSpace — keeps letters and ASCII digits, replaces everything else
    with a space.
  • SplitOnCaseChange / IsWordBoundary — the regex's three alternatives, expressed as explicit
    conditions over code points. The comments name the case each one covers.
  • IsAllCaps — walks code points directly rather than stripping and re-testing.

ToTitleCase gains Ensure.NotNull(input), which preserves the ArgumentNullException that
Regex.Replace used to throw and satisfies CA1062 now that the regex call is gone. The
#if NET7_0_OR_GREATER split disappears with the regexes, so all six target frameworks now run
identical code.

Astral characters that are not letters — emoji, symbols — remain separators, exactly as before.

Verification

  • Behaviour for BMP input is unchanged. The new implementation was compared against the
    original regex implementation over 12,191 inputs (all three-part combinations of 23 atoms
    covering ASCII case, digits, separators, punctuation, and accented/Greek/eszett letters, plus 24
    realistic identifiers) across all six public methods — no differences. The same harness was
    then re-run with an astral input added and correctly reported the four expected differences, so
    it is not vacuously passing.
  • New tests fail without the fix. Reverted the source change and re-ran: 4 of the 5 new tests
    fail, 25 pass. The fifth pins that non-letter astral characters are still dropped, and passes
    either way by design.
  • Full suite: 29/29 pass. The 24 pre-existing tests are unchanged.
  • Build: clean across net10.0, net9.0, net8.0, netstandard2.0, netstandard2.1,
    0 warnings, 0 errors.

Not in this PR

"MAX_SIZE".ToPascalCase() is MaxSize but "set MAX_SIZE".ToPascalCase() is SetMAXSIZE — the
same token converting two ways depending on what precedes it. That falls out of ToTitleCase
testing IsAllCaps on the whole string to decide whether words are acronyms, which is a
documented heuristic rather than an oversight. Changing it is a design decision about acronym
policy, so it is left alone and noted on #70's sibling discussion in ktsu-dev/Coder#71.

🤖 Generated with Claude Code

https://claude.ai/code/session_014RUsrYfSAuSQ7VWws7vzsr


Generated by Claude Code

The three regexes tested Unicode categories per UTF-16 code unit. A letter
outside the Basic Multilingual Plane is a surrogate pair whose halves are both
categorised as Surrogate rather than as a letter, so `[^\p{L}0-9]` matched both
halves and the letter was replaced with spaces and then collapsed away. Every
conversion silently dropped it: "set 𐐀abc" became "SetAbc".

Two further faults shared that root cause. `IsAllCaps` stripped both surrogate
halves before testing, so an astral lowercase letter was skipped entirely and
"HELLO 𐐨" reported true. The splitter's letter-to-non-letter alternative read
the high surrogate as a non-letter and inserted a spurious word boundary, so
"abc𝒳def" split into two tokens.

Replace all three regexes with walks by code point. The conversions now treat
an astral letter exactly as they treat its BMP analogue, and astral characters
that are not letters stay separators as before.

Behaviour for BMP input is unchanged: the new code was compared against the
regex implementation over 12,191 inputs across all six public methods with no
differences. Four of the five new tests fail without this change.

Fixes #70

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014RUsrYfSAuSQ7VWws7vzsr
@sonarqubecloud

Copy link
Copy Markdown

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Letters outside the BMP are silently deleted by every case conversion

2 participants