Take guest-post subjects from what the small web is actually publishing - #209
Merged
Merged
Conversation
The crossed-seed topics the matcher produces are combinations of two sites' own keyword lists — reliable, and finite. A partner written for a few times exhausts them, `planGuestPost` returns null, and that site's guest-post slot falls back to the author's own blog for ever after. This adds the other source before giving up: a real, recently published post from an RSS Amplifier topic feed, picked at random, which the generator writes a full article about. Nothing is copied. What the feed contributes is a subject somebody in the niche genuinely cared about this week, rather than one assembled from two keyword lists. **Random rather than ranked, deliberately.** Ranking headlines would mean deciding one publisher's is a better subject than another's on evidence we do not have, and the failure it introduces is worse than the one it prevents: a stable ranking over a slow-moving feed writes about the same thing repeatedly, which is the exact problem this source exists to solve. **The subject is a post title, not a topic keyword.** Pulling at random from the directory's topic list would have been the obvious reading and produces nonsense: its largest topics are "one" (11,502 feeds) and "time" (8,544) — stopwords, not subjects. The items inside a topic feed are real editorial. The filter that matters most: **our own ads are excluded.** Those feeds now carry CrawlProof fills as syndication items, so without this the cron could pick one of our advertisements, commission a guest post about it, and publish that on a partner's blog under our name — an ad laundered into editorial. Both the `<category>Sponsored</category>` element and the "(Sponsored)" title suffix are checked, because either alone is a single point of failure for something that must never happen once. Parsed with a regex rather than an XML dependency: one known document shape from one known publisher, one field wanted, and a malformed feed has to degrade to "no subject today" rather than throw inside the publishing cron. Every failure path — a missing topic, a slow directory, an unparseable document — returns the same answer, and the caller publishes an ordinary post instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ThreatCrush Security Scan35 finding(s) HIGH/CRITICAL: 3 | MEDIUM: 23 | LOW: 9
Snippets are redacted; ThreatCrush never prints matched credential material. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The crossed-seed topics the matcher produces are combinations of two sites' own keyword lists — reliable, and finite. A partner written for a few times exhausts them,
planGuestPostreturns null, and that site's guest-post slot falls back to its own blog forever. This adds the other source before giving up: a real, recently published post from an RSS Amplifier topic feed, picked at random, which the generator writes a full article about.Nothing is copied. What the feed contributes is a subject somebody in the niche genuinely cared about this week, rather than one assembled from two keyword lists.
Two calls worth reviewing
Random rather than ranked, deliberately. Ranking headlines means deciding one publisher's is a better subject than another's on evidence we don't have — and the failure it introduces is worse than the one it prevents: a stable ranking over a slow-moving feed writes about the same thing repeatedly, which is the exact problem this source exists to solve.
The subject is a post title, not a topic keyword. Pulling at random from the directory's topic list is the obvious reading and produces nonsense — its largest topics are "one" (11,502 feeds) and "time" (8,544). Stopwords, not subjects. The items inside a topic feed are real editorial.
The filter that matters most
Our own ads are excluded. Those feeds now carry CrawlProof fills as syndication items (from the RSS ad work), so without this the cron could pick one of our advertisements, commission a guest post about it, and publish that on a partner's blog under our name — an ad laundered into editorial. Both
<category>Sponsored</category>and the(Sponsored)title suffix are checked, because either alone is a single point of failure for something that must never happen once.Smaller things
'and—, and a subject reading "Don't" would go into the article verbatim.1,688 tests pass (13 new);
tsc --noEmitclean.🤖 Generated with Claude Code