Skip to content

Carry sponsored items in the syndicated feeds - #103

Merged
ralyodio merged 1 commit into
mainfrom
feed-ads
Aug 18, 2026
Merged

ralyodio merged 1 commit into
mainfrom
feed-ads

Conversation

@ralyodio

Copy link
Copy Markdown
Contributor

The site sells ads on its pages but gave them away on its feeds — which is where most of the reading actually happens. A topic river is a subscription somebody keeps in a reader for months, and until now it was the one surface distributing our content with no way to earn from it.

Consumes the new feed_item format from profullstack/crawlproof.com#200.

Two decisions worth stating

We take as=fields, not the ready-made <item>. Splicing CrawlProof’s XML into ours would be less code, and would put two different pieces of software in charge of escaping inside one document. The day their idea of escaping differs from ours is the day every subscriber’s reader reports a parse error on the whole feed, not on the ad. Taking raw fields and rendering them through buildRss / buildAtom / buildJsonFeed keeps that decision in the one place it is already made for the other fifty items.

Failure is silent and total. Every path out of fetchFeedAds returns [] — no retry, no error raised, a 2s timeout, and the empty answer cached too so an unsold slot does not add a round trip to every response. A feed is the product; an ad is revenue on top of it.

Placement

One in ten, capped at three. The cap matters more than it looks: feeds here run to 200 items, and one in ten with no ceiling would be twenty ads, which is not a river with advertising in it. Ads never trail the feed.

adSlotsFor() exists so the caller fetches exactly as many as will be placed — a fetched-but-dropped ad is an impression metered against an advertiser for reach nobody got. A test pins the two functions together across 13 list lengths.

What carries them

  • Topic feeds (/topics/x.rss|.atom|.json) and the following river.
  • Playlists carry none. An .m3u is an ordered list of things to play, and a sponsored line has nothing a player can open — the same reason video/youtube enclosures are already excluded from them.
  • /api/*, /llms.txt and /opml stay clean. The line is who the document is for: a feed is a subscription a person reads in a reader; those are the machine-readable copy an agent consumes, where an ad is noise in a data structure rather than a placement anybody sees.

Disclosure

The title arrives with its prefix already applied and is used as-is. RSS gets <category>Sponsored</category>, Atom a <category term="sponsored"> plus <rights>, JSON Feed both a human tags entry and a _crawlproof.sponsored extension field an agent can branch on.

Kill switch

FEED_ADS=0 turns every sponsored item off without a deploy — which matters because what it switches off is written into documents other people keep.

Verification

All 11 workspaces pass (21 new tests), next build clean. Tests cover XML well-formedness with a real validator, hostile advertiser copy (markup in a title, ]]> in a body), and that a real post still renders byte-identically.

🤖 Generated with Claude Code

The site sells ads on its pages but gave them away on its feeds, which is
where most of the reading actually happens: a topic river is a subscription
somebody keeps in a reader for months, and until now it was the one surface
distributing our content with no way to earn from it.

An ad in a feed is not the web unit in a different place. There is no DOM for
ad.js to fill and no second request a reader will make, so the ad has to be
*in* the document, fetched while it is built.

Two decisions are worth stating, because both are easy to undo later.

We take crawlproof's `as=fields` rather than its ready-made `<item>`. Splicing
their XML into ours would be less code and would put two different pieces of
software in charge of escaping inside one document -- and the day their idea
of escaping differs from ours is the day every subscriber's reader reports a
parse error on the whole feed, not on the ad. Taking raw fields and rendering
them through buildRss/buildAtom/buildJsonFeed keeps that decision in the one
place it is already made for the other fifty items.

Failure is silent and total. Every path out of the fetcher returns []: no
retry, no error raised, a 2s timeout, and the empty answer cached too so an
unsold slot does not add a round trip to every response. A feed is the
product; an ad is revenue on top of it.

Placement is one in ten, capped at three. The cap matters more than it looks
-- feeds here run to 200 items, and one in ten with no ceiling would be twenty
ads, which is not a river with advertising in it. Ads never trail the feed,
and `adSlotsFor` exists so the caller fetches exactly as many as will be
placed: a fetched-but-dropped ad is an impression metered against an
advertiser for reach nobody got.

Playlists carry none. An .m3u is an ordered list of things to play and a
sponsored line has nothing a player can open -- the same reason
`video/youtube` enclosures are already excluded from them.

FEED_ADS=0 turns the whole thing off without a deploy, which matters because
what it switches off is written into documents other people keep.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@ralyodio
ralyodio marked this pull request as ready for review August 18, 2026 14:39
@ralyodio
ralyodio merged commit 5880c01 into main Aug 18, 2026
2 checks passed
@ralyodio
ralyodio deleted the feed-ads branch August 18, 2026 14:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant