What's wrong
CaseConverter changes case with ToLowerInvariant() / ToUpperInvariant():
ToSnakeCase, around line 428 of CaseConverter/CaseConverter.cs
ToMacroCase, around line 461
ToLowercaseFirstChar / ToUppercaseFirstChar, around lines 189 and 206
ToCamelCase, ToPascalCase and ToKebabCase go through these same calls.
.NET's invariant casing deliberately leaves two code points unmapped:
- U+0130
İ (capital I with dot above)
- U+0131
ı (dotless small i)
For example, "İ".ToLowerInvariant() returns "İ". Unicode's simple mappings in UnicodeData.txt are İ → i and ı → I. The result is that an uppercase letter stays in snake, kebab and camel output, and a lowercase letter stays in MACRO and Pascal output. Each of those breaks the case style's own rule.
Reproduction
Observed by running the current source on net10.0:
| Call |
Actual |
Expected |
"İzmir city".ToSnakeCase() |
İzmir_city |
izmir_city |
"İzmir city".ToKebabCase() |
İzmir-city |
izmir-city |
"İzmir city".ToCamelCase() |
İzmirCity |
izmirCity |
"İstanbul".ToLowercaseFirstChar() |
İstanbul |
istanbul |
"kırmızı".ToMacroCase() |
KıRMıZı |
KIRMIZI |
"ılık su".ToPascalCase() |
ılıkSu |
IlıkSu |
"ılık su".ToUppercaseFirstChar() |
ılık su |
Ilık su |
There is a likely knock-on effect, which I traced from the code but did not run. SplitOnCaseChange treats each ı followed by a capital as a word boundary. So ToMacroCase applied to its own output KıRMıZı should yield Kı_RMı_Zı, which means the conversion is not idempotent for Turkish or Azerbaijani text.
Why it matters
These conversions are typically used to derive identifiers, file names or config keys. Turkish place names and words are ordinary input, and the output currently contains characters that the target convention forbids. Consumers that validate the result, for example with ^[a-z0-9_]+$-style checks after an ASCII fold, or that compare it with a case-sensitive match, will reject it or mismatch.
This is not covered by #84 (lowercase letters rewritten by an upper→lower round trip) or #87 (ß). This issue is about two letters that are never mapped at all.
Suggested fix / acceptance criteria
- Route the four casing calls through a small helper. It applies the invariant mapping, then patches the two excluded code points to their Unicode simple mappings:
İ → i when lowercasing, ı → I when uppercasing. Do not switch to a culture's TextInfo: its result depends on the ICU data on the machine.
- Add tests:
"İzmir city".ToSnakeCase() == "izmir_city"
"kırmızı".ToMacroCase() == "KIRMIZI"
ToMacroCase is idempotent on "kırmızı"
"ılık su".ToPascalCase() starts with I
What's wrong
CaseConverter changes case with
ToLowerInvariant()/ToUpperInvariant():ToSnakeCase, around line 428 ofCaseConverter/CaseConverter.csToMacroCase, around line 461ToLowercaseFirstChar/ToUppercaseFirstChar, around lines 189 and 206ToCamelCase,ToPascalCaseandToKebabCasego through these same calls..NET's invariant casing deliberately leaves two code points unmapped:
İ(capital I with dot above)ı(dotless small i)For example,
"İ".ToLowerInvariant()returns"İ". Unicode's simple mappings in UnicodeData.txt areİ→iandı→I. The result is that an uppercase letter stays in snake, kebab and camel output, and a lowercase letter stays in MACRO and Pascal output. Each of those breaks the case style's own rule.Reproduction
Observed by running the current source on net10.0:
"İzmir city".ToSnakeCase()İzmir_cityizmir_city"İzmir city".ToKebabCase()İzmir-cityizmir-city"İzmir city".ToCamelCase()İzmirCityizmirCity"İstanbul".ToLowercaseFirstChar()İstanbulistanbul"kırmızı".ToMacroCase()KıRMıZıKIRMIZI"ılık su".ToPascalCase()ılıkSuIlıkSu"ılık su".ToUppercaseFirstChar()ılık suIlık suThere is a likely knock-on effect, which I traced from the code but did not run.
SplitOnCaseChangetreats eachıfollowed by a capital as a word boundary. SoToMacroCaseapplied to its own outputKıRMıZıshould yieldKı_RMı_Zı, which means the conversion is not idempotent for Turkish or Azerbaijani text.Why it matters
These conversions are typically used to derive identifiers, file names or config keys. Turkish place names and words are ordinary input, and the output currently contains characters that the target convention forbids. Consumers that validate the result, for example with
^[a-z0-9_]+$-style checks after an ASCII fold, or that compare it with a case-sensitive match, will reject it or mismatch.This is not covered by #84 (lowercase letters rewritten by an upper→lower round trip) or #87 (
ß). This issue is about two letters that are never mapped at all.Suggested fix / acceptance criteria
İ→iwhen lowercasing,ı→Iwhen uppercasing. Do not switch to a culture'sTextInfo: its result depends on the ICU data on the machine."İzmir city".ToSnakeCase() == "izmir_city""kırmızı".ToMacroCase() == "KIRMIZI"ToMacroCaseis idempotent on"kırmızı""ılık su".ToPascalCase()starts withI