Skip to content

Take guest-post subjects from what the small web is actually publishing - #209

Merged
ralyodio merged 1 commit into
masterfrom
feature/guest-topics-from-feeds
Aug 19, 2026
Merged

ralyodio merged 1 commit into
masterfrom
feature/guest-topics-from-feeds

Conversation

@ralyodio

Copy link
Copy Markdown
Contributor

The crossed-seed topics the matcher produces are combinations of two sites' own keyword lists — reliable, and finite. A partner written for a few times exhausts them, planGuestPost returns null, and that site's guest-post slot falls back to its own blog forever. This adds the other source before giving up: a real, recently published post from an RSS Amplifier topic feed, picked at random, which the generator writes a full article about.

Nothing is copied. What the feed contributes is a subject somebody in the niche genuinely cared about this week, rather than one assembled from two keyword lists.

Two calls worth reviewing

Random rather than ranked, deliberately. Ranking headlines means deciding one publisher's is a better subject than another's on evidence we don't have — and the failure it introduces is worse than the one it prevents: a stable ranking over a slow-moving feed writes about the same thing repeatedly, which is the exact problem this source exists to solve.

The subject is a post title, not a topic keyword. Pulling at random from the directory's topic list is the obvious reading and produces nonsense — its largest topics are "one" (11,502 feeds) and "time" (8,544). Stopwords, not subjects. The items inside a topic feed are real editorial.

The filter that matters most

Our own ads are excluded. Those feeds now carry CrawlProof fills as syndication items (from the RSS ad work), so without this the cron could pick one of our advertisements, commission a guest post about it, and publish that on a partner's blog under our name — an ad laundered into editorial. Both <category>Sponsored</category> and the (Sponsored) title suffix are checked, because either alone is a single point of failure for something that must never happen once.

Smaller things

  • Regex rather than an XML dependency: one known document shape, one field wanted, and a malformed feed must degrade to "no subject today" rather than throw inside the publishing cron.
  • Entities are decoded — titles arrive from 50k publishers carrying &apos; and &#8212;, and a subject reading "Don't" would go into the article verbatim.
  • Titles under 24 or over 160 characters are dropped: too short is a label ("Weeknotes"), too long is a paragraph that won't survive a prompt.
  • Every failure path returns null and the caller publishes an ordinary post.

1,688 tests pass (13 new); tsc --noEmit clean.

🤖 Generated with Claude Code

The crossed-seed topics the matcher produces are combinations of two sites' own
keyword lists — reliable, and finite. A partner written for a few times
exhausts them, `planGuestPost` returns null, and that site's guest-post slot
falls back to the author's own blog for ever after. This adds the other source
before giving up: a real, recently published post from an RSS Amplifier topic
feed, picked at random, which the generator writes a full article about.

Nothing is copied. What the feed contributes is a subject somebody in the niche
genuinely cared about this week, rather than one assembled from two keyword
lists.

**Random rather than ranked, deliberately.** Ranking headlines would mean
deciding one publisher's is a better subject than another's on evidence we do
not have, and the failure it introduces is worse than the one it prevents: a
stable ranking over a slow-moving feed writes about the same thing repeatedly,
which is the exact problem this source exists to solve.

**The subject is a post title, not a topic keyword.** Pulling at random from the
directory's topic list would have been the obvious reading and produces
nonsense: its largest topics are "one" (11,502 feeds) and "time" (8,544) —
stopwords, not subjects. The items inside a topic feed are real editorial.

The filter that matters most: **our own ads are excluded.** Those feeds now
carry CrawlProof fills as syndication items, so without this the cron could pick
one of our advertisements, commission a guest post about it, and publish that on
a partner's blog under our name — an ad laundered into editorial. Both the
`<category>Sponsored</category>` element and the "(Sponsored)" title suffix are
checked, because either alone is a single point of failure for something that
must never happen once.

Parsed with a regex rather than an XML dependency: one known document shape from
one known publisher, one field wanted, and a malformed feed has to degrade to
"no subject today" rather than throw inside the publishing cron. Every failure
path — a missing topic, a slow directory, an unparseable document — returns the
same answer, and the caller publishes an ordinary post instead.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

ThreatCrush Security Scan

35 finding(s)

HIGH/CRITICAL: 3 | MEDIUM: 23 | LOW: 9

Severity Rule Location
HIGH tls-verification-disabled lib/onion.ts:47
HIGH secret-generic-credential lib/sp/platforms/facebook.ts:32
HIGH sh-remote-script-execution prober/deploy/provision.sh:30
MEDIUM js-unescaped-html-sink app/(app)/dashboard/admin/email-broadcast/EmailBroadcastForm.tsx:125
MEDIUM js-unescaped-html-sink app/(app)/dashboard/projects/[id]/autoblog/articles/[articleId]/page.tsx:214
MEDIUM js-unescaped-html-sink app/(marketing)/blog/[slug]/page.tsx:67
MEDIUM js-unescaped-html-sink app/(marketing)/blog/[slug]/page.tsx:97
MEDIUM js-unescaped-html-sink app/(marketing)/blog/[slug]/page.tsx:104
MEDIUM js-unescaped-html-sink app/(marketing)/blog/[slug]/page.tsx:110
MEDIUM js-unescaped-html-sink app/(marketing)/recent/page.tsx:186
MEDIUM js-unescaped-html-sink app/(marketing)/recent/page.tsx:190
MEDIUM js-unescaped-html-sink app/c/[project]/[slug]/page.tsx:77
MEDIUM js-unescaped-html-sink app/c/[project]/page.tsx:57
MEDIUM js-unescaped-html-sink app/careers.js/route.ts:228
MEDIUM js-unescaped-html-sink app/careers.js/route.ts:285
MEDIUM js-unescaped-html-sink app/layout.tsx:129
MEDIUM js-open-redirect app/login/form.tsx:39
MEDIUM js-unescaped-html-sink app/r/[token]/page.tsx:176
MEDIUM js-open-redirect app/signup/form.tsx:43
MEDIUM js-open-redirect components/billing/buy-credits-modal.tsx:98
MEDIUM js-unescaped-html-sink components/json-ld.tsx:8
MEDIUM js-unescaped-html-sink components/report/markdown-view.tsx:15
MEDIUM redos-nested-quantifier lib/careers/jobs.ts:139
MEDIUM js-unescaped-html-sink lib/careers/page-templates.ts:198
MEDIUM redos-nested-quantifier lib/emailMarkdown.ts:130
MEDIUM redos-nested-quantifier lib/lx/articleGen.ts:98
LOW secret-generic-credential app/(marketing)/docs/autoblog-webhook/page.tsx:145
LOW secret-generic-credential lib/sp/platforms/linkedin.ts:25
LOW js-dynamic-code-execution tests/careers-page-templates.test.ts:21
LOW js-dynamic-code-execution tests/careers-widget-script.test.ts:19
LOW js-dynamic-code-execution tests/careers-widget-script.test.ts:69
LOW js-dynamic-code-execution tests/contract/ad-visitor-id.test.ts:51
LOW js-dynamic-code-execution tests/contract/ad-visitor-id.test.ts:52
LOW secret-generic-credential tests/contract/posthog-integration.test.ts:13
LOW secret-generic-credential tests/lead-campaign.test.ts:16

Snippets are redacted; ThreatCrush never prints matched credential material.

@ralyodio
ralyodio merged commit 04c804b into master Aug 19, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant