Skip to content

İ and ı are never case-mapped: "İzmir city".ToSnakeCase() returns "İzmir_city", "kırmızı".ToMacroCase() returns "KıRMıZı" #97

Description

@matt-edmondson

What's wrong

CaseConverter changes case with ToLowerInvariant() / ToUpperInvariant():

  • ToSnakeCase, around line 428 of CaseConverter/CaseConverter.cs
  • ToMacroCase, around line 461
  • ToLowercaseFirstChar / ToUppercaseFirstChar, around lines 189 and 206

ToCamelCase, ToPascalCase and ToKebabCase go through these same calls.

.NET's invariant casing deliberately leaves two code points unmapped:

  • U+0130 İ (capital I with dot above)
  • U+0131 ı (dotless small i)

For example, "İ".ToLowerInvariant() returns "İ". Unicode's simple mappings in UnicodeData.txt are İ → i and ı → I. The result is that an uppercase letter stays in snake, kebab and camel output, and a lowercase letter stays in MACRO and Pascal output. Each of those breaks the case style's own rule.

Reproduction

Observed by running the current source on net10.0:

Call Actual Expected
"İzmir city".ToSnakeCase() İzmir_city izmir_city
"İzmir city".ToKebabCase() İzmir-city izmir-city
"İzmir city".ToCamelCase() İzmirCity izmirCity
"İstanbul".ToLowercaseFirstChar() İstanbul istanbul
"kırmızı".ToMacroCase() KıRMıZı KIRMIZI
"ılık su".ToPascalCase() ılıkSu IlıkSu
"ılık su".ToUppercaseFirstChar() ılık su Ilık su

There is a likely knock-on effect, which I traced from the code but did not run. SplitOnCaseChange treats each ı followed by a capital as a word boundary. So ToMacroCase applied to its own output KıRMıZı should yield Kı_RMı_Zı, which means the conversion is not idempotent for Turkish or Azerbaijani text.

Why it matters

These conversions are typically used to derive identifiers, file names or config keys. Turkish place names and words are ordinary input, and the output currently contains characters that the target convention forbids. Consumers that validate the result, for example with ^[a-z0-9_]+$-style checks after an ASCII fold, or that compare it with a case-sensitive match, will reject it or mismatch.

This is not covered by #84 (lowercase letters rewritten by an upper→lower round trip) or #87 (ß). This issue is about two letters that are never mapped at all.

Suggested fix / acceptance criteria

  • Route the four casing calls through a small helper. It applies the invariant mapping, then patches the two excluded code points to their Unicode simple mappings: İ → i when lowercasing, ı → I when uppercasing. Do not switch to a culture's TextInfo: its result depends on the ICU data on the machine.
  • Add tests:
    • "İzmir city".ToSnakeCase() == "izmir_city"
    • "kırmızı".ToMacroCase() == "KIRMIZI"
    • ToMacroCase is idempotent on "kırmızı"
    • "ılık su".ToPascalCase() starts with I

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't workingreadyFully specified; implement as written

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions