sfnt: read the format 4 glyph id array from its real start - #54
Merged
Merged
Conversation
The glyph id array of a format 4 cmap begins after the idRangeOffset array, at 16+4*segCountX2. The reader sliced it at 16+3*segCountX2 - the start of the offset array - while lookup measured from its end, so every code point in a range-offset segment read the id segCountX2 bytes early, a neighbour's glyph or none. Delta segments were unaffected, which is all the test font had. In Noto Sans that drew "•" as "," in every embedded PDF and left 753 BMP code points (Romanian Ș/ț, much of Cyrillic Extended, the combining marks) without a glyph. The PDF's ToUnicode map came from the rune, so copied text still read correctly and hid it. Against x/image/font/sfnt the lookup now agrees on every BMP code point of Noto Sans Regular, Bold and Italic. sfnttest.Cmap4Ranged is a format 4 map with two range-offset segments, so the offset arithmetic is checked at two different segment indices. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The glyph id array of a format 4 cmap begins after the idRangeOffset array (
16+4*segCountX2).read4sliced it at16+3*segCountX2— the start of the offset array — whilelookupmeasured from its end, so every code point in a range-offset segment read the idsegCountX2bytes early: a neighbour's glyph, or none. Delta segments were unaffected, and the test font only had those.Impact with Noto Sans (the Fyne UI font):
•drew as,in every embedded PDF, and 753 BMP code points (Romanian Ș/ț, much of Cyrillic Extended, combining marks) had no glyph. ToUnicode is built from the rune, so copied text still read correctly and hid it.c.glyphIDs = t[16+4*c.segX2:]sfnttest.Cmap4Ranged: a format 4 map with two range-offset segments, so the offset arithmetic is checked at two segment indices;TestTheCmapFormats/format 4 with range offsetsis red before, green after.golang.org/x/image/font/sfnt: every BMP code point of Noto Sans Regular, Bold and Italic now maps identically (was 753 mismatches).🤖 Generated with Claude Code