Skip to content

ToPascalCase/ToCamelCase change their own output when run again on adjacent one-letter words: "vector x y" → "VectorXY" → "VectorXy" #95

Description

@matt-edmondson

What's wrong

Adjacent one-letter words turn into a run of capitals in PascalCase: "vector x y" → "VectorXY", "a b c" → "ABC", "x_y" → "XY". When that output is converted again, IsWordBoundary reads the run as a single all-caps word, and LowercaseAllCapsWords (the per-word all-caps normalisation added for #72) lowercases it. So the converters don't return their own output unchanged, and round trips silently lose word boundaries.

Relevant code, all in CaseConverter/CaseConverter.cs:

  • ToPascalCase (~L370)
  • LowercaseAllCapsWords (~L240)
  • the acronym-tail rule in IsWordBoundary (~L150)

Reproduction

Against the net10.0 build of HEAD (1c1c555):

"vector x y".ToPascalCase()               = "VectorXY"
"VectorXY".ToPascalCase()                 = "VectorXy"   // expected "VectorXY"
"vector x y".ToCamelCase()                = "vectorXY"
"vectorXY".ToCamelCase()                  = "vectorXy"   // expected "vectorXY"
"a b c".ToPascalCase().ToPascalCase()     = "Abc"        // first pass "ABC"
"x_y".ToPascalCase().ToPascalCase()       = "Xy"         // first pass "XY"

"vector_x_y".ToPascalCase().ToSnakeCase() = "vector_xy"  // boundary between x and y lost
"a_b".ToPascalCase().ToSnakeCase()        = "ab"

A single one-letter word does survive, because the next letter is lowercase and the acronym-tail rule applies: "get_a_value" → "GetAValue" → "get_a_value".

Why it matters

Identifiers like coord_x_y, matrix_m_n and get_x_y are common. A code generator or rename tool that normalises names more than once, or goes snake → Pascal → snake, gets different identifiers on each pass: VectorXY becomes VectorXy, and coord_x_y becomes coord_xy.

Suggested fix / acceptance criteria

ToPascalCase and ToCamelCase should not produce output they would convert differently on a second pass. One option: when a one-letter word follows another one-letter word, lowercase it, so "vector x y" → "VectorXy" on the first pass. That output is stable, although vector_x_y → Pascal → snake would still come back as vector_xy. If that fold is acceptable, document it.

Tests:

  • s.ToPascalCase().ToPascalCase() == s.ToPascalCase(), and the same for ToCamelCase, for "vector x y", "a b c" and "x_y"
  • a tested expectation for "vector_x_y".ToPascalCase().ToSnakeCase()

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions