Skip to content

stop truncating out-of-bmp identifier escapes in the scanner - #4337

Open
dxbjavid wants to merge 1 commit into
google:masterfrom
dxbjavid:identifier-escape-bmp
Open

stop truncating out-of-bmp identifier escapes in the scanner#4337
dxbjavid wants to merge 1 commit into
google:masterfrom
dxbjavid:identifier-escape-bmp

Conversation

@dxbjavid

Copy link
Copy Markdown

Out-of-range identifier escapes get silently truncated to a char

In processUnicodeEscapes a braced unicode escape inside an identifier has its parsed code point cast straight to a char, so only the low 16 bits survive. That means an escape like \u{10041} is quietly accepted as the identifier A, and even values above the U+10FFFF maximum such as \u{110041} slip through the same way, so two identifiers that ought to be distinct (or rejected outright) can end up collapsing onto one. The scanner already rejects \u{10000} here, and the string-literal path range-checks the code point, so this brings identifiers into line by rejecting anything that does not fit in a char rather than aliasing it to a different character.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant