Skip to content

Pascal/Camel/Snake/Kebab/Macro delete combining marks and non-ASCII digits: "हिन्दी भाषा" → "ह_न_द_भ_ष", NFD "éclair" → "EClair" #80

Description

@matt-edmondson

What's wrong

ReplaceNonAlphaNumericWithSpace (CaseConverter/CaseConverter.cs:43) keeps a code point only if it passes char.IsLetter or is an ASCII digit '0'..'9'. Two kinds of word characters fail that test and get replaced with a space:

  1. Combining marks (UnicodeCategory.NonSpacingMark, SpacingCombiningMark, EnclosingMark). These carry Indic vowel signs and viramas, Thai and Arabic vowel marks, and the accents in NFD-decomposed Latin text. IsWordBoundary (around :139-144) has a related problem: a mark that follows a letter counts as "a non-letter", so a word boundary is placed before it.
  2. Non-ASCII decimal digits (Arabic-Indic ١٢٣, Devanagari १२३, and so on). The word splitter already uses the Unicode-aware char.IsDigit, so these fail only in this one helper.

ToTitleCase does not go through this path, so it keeps both kinds of character. The sibling APIs therefore disagree on the same input.

Repro

Input Call Actual Expected
"हिन्दी भाषा" ToSnakeCase() "ह_न_द_भ_ष" "हिन्दी_भाषा"
"हिन्दी भाषा" ToPascalCase() "हनदभष" "हिन्दीभाषा"
"éclair" (NFD: e + U+0301) ToPascalCase() "EClair" "Éclair" (or "Éclair")
"éclair" (NFD) ToSnakeCase() "e_clair" "éclair"
"café au lait" (NFD) ToPascalCase() "CafeAuLait" "CaféAuLait"
"x١٢٣y" ToPascalCase() / ToSnakeCase() "XY" / "x_y" "X١٢٣Y" / "x_١٢٣y"
"x١٢٣y" ToTitleCase() "X ١٢٣Y" (keeps the digits) —

Why it matters

These calls corrupt text without any error. Whole scripts come out unreadable (Devanagari, Bengali, Thai, Tamil, vowelled Arabic). NFD input also breaks, and NFD is common: macOS filesystem names arrive decomposed, as does text from some editors and normalisers. This is the same class of bug as the astral-letter fix (#70/#71), which already established that letters from every script are kept, just for a different set of Unicode categories.

Suggested fix

  • In ReplaceNonAlphaNumericWithSpace, also keep code points whose category is NonSpacingMark, SpacingCombiningMark or EnclosingMark. Replace the ASCII digit check with char.IsDigit(input, i).
  • In IsWordBoundary, never break before a combining mark, and treat a mark as continuing the letter before it. For example, when deciding "previous is a letter", skip back over marks.
  • In the first-character upper/lower mapping, keep any marks that follow the base letter attached to it.

Acceptance: the table above produces the expected column, and there are regression tests for Devanagari, NFD Latin and Arabic-Indic digits across Pascal, Camel, Snake, Kebab and Macro.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    readyFully specified; implement as written

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions