Repository navigation
fix(social): send Bluesky rich-text facets, so links and tags are live - #131
Merged
Merged
Conversation
Bluesky parses nothing out of post text. A URL posted as plain text
stays plain text and a hashtag is just a word starting with '#'.
Anything clickable has to be described by a facet giving its byte range
and what it points at, and there is no auto-parse flag to turn on. We
were posting neither, so every link and tag we published was inert.
The part that makes this easy to get wrong is that facet offsets are
counted in UTF-8 bytes while JavaScript string indices are UTF-16 code
units. One emoji, accented character or CJK word before a link shifts
the two apart, and the facet then highlights the wrong span — mid-word,
or past the end of the string. Every offset here goes through
Buffer.byteLength, and the tests assert by slicing the UTF-8 buffer at
the offsets we emit rather than by trusting the numbers.
Two related length bugs came out of the same confusion. Truncation used
text.slice(0, 300): that spends two of the 300 on every emoji, and can
cut between the halves of a surrogate pair, producing a lone surrogate
that is not valid UTF-8. The pre-flight check compared text.length
against the limit, so a post of 200 emoji measured 400 and was rejected
though the API would have taken it. Both now count graphemes, which is
how Bluesky counts and how a reader reads.
URLs keep sentence punctuation outside the link, tags drop trailing
punctuation, purely numeric tags are ignored as prose ("ranked #1"), and
a '#' inside a URL fragment is not mistaken for a tag.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
vu1nz Security Review0 finding(s) in PR #? No security issues found. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Bluesky parses nothing out of post text. A URL posted as plain text stays plain text; a hashtag is just a word starting with
#. Anything clickable has to be described by a facet giving its byte range and target, and there is no auto-parse flag.lib/sp/platforms/bluesky.tsbuilt its record as:No
facets. So every link and tag we've published has been inert.The offsets are the hard part
Facet offsets are UTF-8 bytes; JavaScript string indices are UTF-16 code units. One emoji, accented character, or CJK word before a link shifts the two apart, and the facet highlights the wrong span — mid-word, or past the end of the string.
Every offset goes through
Buffer.byteLength. The tests assert by slicing the UTF-8 buffer at the emitted offsets rather than trusting the numbers:Covered for emoji (
🚀 https://…→byteStart5, not 2), CJK, and accented text.Two more bugs from the same confusion
Truncation.
text.slice(0, 300)spends two of the 300 on every emoji, and can cut between the halves of a surrogate pair — producing a lone surrogate that isn't valid UTF-8. There's a test that demonstrates the old behaviour failing:The pre-flight check.
lib/sp/post.tscomparedtext.lengthagainst the limit, so a post of 200 emoji measured 400 and was rejected though the API would have accepted it.Both now count graphemes — how Bluesky counts, and how a reader reads.
Parsing details
https://example.com.))survives#inside a URL fragment isn't mistaken for a tagNot included
Mentions (
@handle) still post as plain text — linking those needs aresolveHandlecall per mention, which is a network round-trip inside post composition. Happy to add it if you want it.Checks
tsc --noEmitclean🤖 Generated with Claude Code