Skip to content

Escape control characters and Unicode line terminators in string literals [patch] - #144

Merged
matt-edmondson merged 1 commit into
mainfrom
fix/escape-control-chars-98
Sep 30, 2026
Merged

matt-edmondson merged 1 commit into
mainfrom
fix/escape-control-chars-98

Conversation

@matt-edmondson

Copy link
Copy Markdown
Contributor

Fixes #98

What was wrong

LanguageGeneratorBase.EscapeString escaped only \, ", \n, \r and \t. Every other control character, and the Unicode line terminators, was copied raw into the quoted literal. That caused three failures:

  • Go: a NUL is a compile error ("invalid NUL character").
  • Python: a NUL is a SyntaxError.
  • C#: U+0085, U+2028 and U+2029 give CS1010: Newline in constant.

Change

  • EscapeString is now an instance method. After the named escapes, it writes every remaining C0 control, DEL, U+0085, U+2028 and U+2029 through a new virtual EscapeCodeUnit hook:
Generator Escape
C#, JavaScript, Python, Go (default) \uXXXX
Rust \u{XXXX}
C, C++ (CFamilyGenerator) octal, one \NNN per UTF-8 byte
  • C and C++ use octal for two reasons: their \x is greedy and would swallow a following hex digit, and a universal character name may not name a control character. The UTF-8 bytes are the same bytes the character would have been written as raw, so the string's contents don't change.
  • A string with nothing to escape numerically takes the old path unchanged.

Tests

  • New StringLiteralControlCharacterTests checks one string in all seven generators. The string holds \0, ESC followed by a hex digit (b), DEL, \n, NEL, U+2028 and U+2029. For each generator the test checks two things: the exact escaped literal, and that no raw control character is left in the output. A second test checks that printable text, including non-ASCII, is still written as-is.
  • With the generator change reverted, all 7 per-language cases fail.
  • I also checked each escaped spelling by hand in the real toolchain: python3, node, gcc -pedantic, go run and rustc. Each one decodes back to the original characters.
  • Full suite: 995 passed, 0 failed.

🤖 Generated with Claude Code

https://claude.ai/code/session_01HsaGezbczZSL9fXfQafs7x


Generated by Claude Code

…ing literals

EscapeString only escaped backslash, quote, \n, \r and \t, so a NUL,
U+0085, U+2028 or U+2029 was copied raw into the literal: Go and Python
reject a NUL in source, and C# reads the line terminators as ending the
line. Every remaining C0 control, DEL, U+0085, U+2028 and U+2029 is now
written as a numeric escape through a per-generator EscapeCodeUnit hook:
\uXXXX by default (C#, JavaScript, Python, Go), \u{XXXX} in Rust, and
octal UTF-8 bytes in C and C++, whose \x is greedy.

Fixes #98

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HsaGezbczZSL9fXfQafs7x
@sonarqubecloud

Copy link
Copy Markdown

@matt-edmondson
matt-edmondson merged commit 04fcd25 into main Sep 30, 2026
14 checks passed
@matt-edmondson
matt-edmondson deleted the fix/escape-control-chars-98 branch September 30, 2026 09:52
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

String literals containing NUL or U+2028/U+2029/U+0085 generate code that fails to compile in Go, Python and C#

2 participants