Skip to content

fix(social): send Bluesky rich-text facets, so links and tags are live - #131

Merged
ralyodio merged 1 commit into
masterfrom
fix/bluesky-facets
Jul 28, 2026
Merged

ralyodio merged 1 commit into
masterfrom
fix/bluesky-facets

Conversation

@ralyodio

Copy link
Copy Markdown
Contributor

Bluesky parses nothing out of post text. A URL posted as plain text stays plain text; a hashtag is just a word starting with #. Anything clickable has to be described by a facet giving its byte range and target, and there is no auto-parse flag.

lib/sp/platforms/bluesky.ts built its record as:

const record = { $type: "app.bsky.feed.post", text, createdAt };

No facets. So every link and tag we've published has been inert.

The offsets are the hard part

Facet offsets are UTF-8 bytes; JavaScript string indices are UTF-16 code units. One emoji, accented character, or CJK word before a link shifts the two apart, and the facet highlights the wrong span — mid-word, or past the end of the string.

Every offset goes through Buffer.byteLength. The tests assert by slicing the UTF-8 buffer at the emitted offsets rather than trusting the numbers:

sliceByBytes(text, f.index.byteStart, f.index.byteEnd) === "https://example.com"

Covered for emoji (🚀 https://… → byteStart 5, not 2), CJK, and accented text.

Two more bugs from the same confusion

Truncation. text.slice(0, 300) spends two of the 300 on every emoji, and can cut between the halves of a surrogate pair — producing a lone surrogate that isn't valid UTF-8. There's a test that demonstrates the old behaviour failing:

const naive = ("a".repeat(300) + "🚀").slice(0, 301);
expect(naive).toMatch(LONE_SURROGATE);   // the bug

The pre-flight check. lib/sp/post.ts compared text.length against the limit, so a post of 200 emoji measured 400 and was rejected though the API would have accepted it.

Both now count graphemes — how Bluesky counts, and how a reader reads.

Parsing details

  • URLs keep sentence punctuation outside the link (https://example.com.)
  • unbalanced closing brackets are trimmed, but a URL that legitimately ends in ) survives
  • tags drop trailing punctuation, cap at 64 chars
  • purely numeric tags are ignored as prose ("ranked feat(aeo-score): time-series rollup + sparkline on project page #1")
  • a # inside a URL fragment isn't mistaken for a tag
  • facets come back sorted and non-overlapping, which is what the API expects

Not included

Mentions (@handle) still post as plain text — linking those needs a resolveHandle call per mention, which is a network round-trip inside post composition. Happy to add it if you want it.

Checks

  • tsc --noEmit clean
  • 718/718 tests pass, 19 new
  • production build compiles

🤖 Generated with Claude Code

Bluesky parses nothing out of post text. A URL posted as plain text
stays plain text and a hashtag is just a word starting with '#'.
Anything clickable has to be described by a facet giving its byte range
and what it points at, and there is no auto-parse flag to turn on. We
were posting neither, so every link and tag we published was inert.

The part that makes this easy to get wrong is that facet offsets are
counted in UTF-8 bytes while JavaScript string indices are UTF-16 code
units. One emoji, accented character or CJK word before a link shifts
the two apart, and the facet then highlights the wrong span — mid-word,
or past the end of the string. Every offset here goes through
Buffer.byteLength, and the tests assert by slicing the UTF-8 buffer at
the offsets we emit rather than by trusting the numbers.

Two related length bugs came out of the same confusion. Truncation used
text.slice(0, 300): that spends two of the 300 on every emoji, and can
cut between the halves of a surrogate pair, producing a lone surrogate
that is not valid UTF-8. The pre-flight check compared text.length
against the limit, so a post of 200 emoji measured 400 and was rejected
though the API would have taken it. Both now count graphemes, which is
how Bluesky counts and how a reader reads.

URLs keep sentence punctuation outside the link, tags drop trailing
punctuation, purely numeric tags are ignored as prose ("ranked #1"), and
a '#' inside a URL fragment is not mistaken for a tag.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

vu1nz Security Review

0 finding(s) in PR #?

No security issues found.

@ralyodio
ralyodio merged commit 19910c5 into master Jul 28, 2026
8 checks passed
@ralyodio
ralyodio deleted the fix/bluesky-facets branch July 28, 2026 02:21
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant