Skip to content

Give Promote campaigns content sources, so they feed themselves - #204

Merged
ralyodio merged 1 commit into
masterfrom
worktree-promote-sources
Aug 18, 2026
Merged

ralyodio merged 1 commit into
masterfrom
worktree-promote-sources

Conversation

@ralyodio

Copy link
Copy Markdown
Contributor

A Promote list was a hand-pasted set of links. This adds content sources — standing subscriptions that keep a campaign supplied with fresh links — plus the blend that decides which content posts next.

Implements §3, §4 and §11 of the new promote-engine-architecture.md, which this PR also lands.

What's in it

  • Keyword sources. bitcoin becomes https://rssamplifier.com/topics/bitcoin.rss. One keyword is one source. Slugs are hyphenated (artificial-intelligence), verified against the live directory; unknown topics 404, so a bad keyword is caught in the form.
  • Custom RSS/Atom feeds, validated by fetching and parsing before saving, with the same private-address guard the audit engine uses.
  • Shared fetch registry. promo_feed is keyed on feed URL with no user_id: 200 users tracking "bitcoin" poll RSS Amplifier once between them. Conditional on ETag/Last-Modified; geometric backoff on failure.
  • Deficit-based blend. A 70/30 campaign posts exactly 70/30 over 100 ticks and never runs one class more than 3× consecutively — weighted-random gives streaks, and a streak of shared content reads as a content farm.
  • Fallback policy so a user with no original content can still run a campaign, capped per rolling 24h.
  • Attribution. Shared content is written about, not as, and credits the publisher by name.
  • Sources + content-mix UI on /dashboard/promote/[id], four new MCP tools, five new server actions.

Incidental fix

promo_link.times_promoted was re-stamped to 1 every tick — the old query never selected the column it was incrementing. Selection reads it back now, so it counts.

Deploy order

Apply supabase/migrations/20260818190000_promote_sources.sql before deploying. The sweep selects the new promo_list columns; without them the select errors and lists stop posting. Not yet applied — house rule is one migration at a time via the Supabase MCP.

Verification

  • tsc --noEmit clean
  • vitest run: 1629 passed / 0 failed (was 1524 — 105 new tests)
  • next build succeeds
  • Migration parses under the real Postgres grammar (pglast), 27 statements
  • Feed parser tested against a captured live RSS Amplifier feed, namespaces and all

Draft: the migration has not been applied, and nothing here is exercised against prod data yet.

🤖 Generated with Claude Code

A Promote list was a hand-pasted set of links: the user typed 20 URLs and the
drip engine rotated through them forever. That works for a fixed set of product
pages and nothing else. It cannot promote what the user published this morning,
and it gives a user with no back catalogue nothing to post at all.

A source is a standing subscription that keeps supplying links. A keyword
becomes an RSS Amplifier topic feed — "bitcoin" is /topics/bitcoin.rss, and one
keyword is one source, so a list of keywords never collapses into a single
ambiguous URL. Any RSS or Atom feed can be added directly, classified as our
content, a partner's, or the industry's.

Feeds are fetched once and shared. promo_feed is keyed on the feed URL and
carries no user_id, so two hundred users tracking "bitcoin" poll RSS Amplifier
once between them; each subscribing campaign then gets promo_link rows by
reference. Polls are conditional on ETag/Last-Modified, so an unchanged feed
costs a 304 rather than a parse, and a failing feed backs off geometrically and
surfaces its error on the sources that subscribe to it.

Which content posts next is now a blend decision. Selection is deficit-based
rather than weighted-random: on each tick the ownership class furthest below its
target share is the one that posts. Weighted random gives streaks, and a streak
of shared content is what makes an automated account read as a content farm. A
70/30 campaign now posts exactly 70 owned and 30 shared over 100 ticks and never
runs the same class more than three times consecutively. Within a class the
rotation is unchanged: least recently promoted first. When one side runs dry the
fallback policy decides whether to cover for it, capped per rolling day so a user
with no original content cannot become a shared-content firehose.

Shared content is now written about rather than as: generatePitch is told whose
content it is and credits the publisher by name, instead of announcing somebody
else's blog post as though we shipped it.

Two smaller things fall out of this. promo_link.times_promoted was being
re-stamped to 1 every tick because the old query never selected the column it was
incrementing; selection reads it back now, so it counts. And the URL identity
used for dedupe is deliberately separate from the URL we publish — folding www.
is right for "have I posted this story before?" and wrong for "which URL do I
hand to Reddit?".

docs/promote-engine-architecture.md records the target architecture and, more
importantly, corrects the runtime baseline it arrived with: this is Next.js on
Supabase with one Railway worker, not Bun services on Turso with Redis queues.
Building to the latter would have forked the product.

The migration must be applied before this deploys — the sweep selects the new
promo_list columns.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

ThreatCrush Security Scan

35 finding(s)

HIGH/CRITICAL: 3 | MEDIUM: 23 | LOW: 9

Severity Rule Location
HIGH tls-verification-disabled lib/onion.ts:47
HIGH secret-generic-credential lib/sp/platforms/facebook.ts:32
HIGH sh-remote-script-execution prober/deploy/provision.sh:30
MEDIUM js-unescaped-html-sink app/(app)/dashboard/admin/email-broadcast/EmailBroadcastForm.tsx:125
MEDIUM js-unescaped-html-sink app/(app)/dashboard/projects/[id]/autoblog/articles/[articleId]/page.tsx:214
MEDIUM js-unescaped-html-sink app/(marketing)/blog/[slug]/page.tsx:67
MEDIUM js-unescaped-html-sink app/(marketing)/blog/[slug]/page.tsx:97
MEDIUM js-unescaped-html-sink app/(marketing)/blog/[slug]/page.tsx:104
MEDIUM js-unescaped-html-sink app/(marketing)/blog/[slug]/page.tsx:110
MEDIUM js-unescaped-html-sink app/(marketing)/recent/page.tsx:186
MEDIUM js-unescaped-html-sink app/(marketing)/recent/page.tsx:190
MEDIUM js-unescaped-html-sink app/c/[project]/[slug]/page.tsx:77
MEDIUM js-unescaped-html-sink app/c/[project]/page.tsx:57
MEDIUM js-unescaped-html-sink app/careers.js/route.ts:228
MEDIUM js-unescaped-html-sink app/careers.js/route.ts:285
MEDIUM js-unescaped-html-sink app/layout.tsx:129
MEDIUM js-open-redirect app/login/form.tsx:39
MEDIUM js-unescaped-html-sink app/r/[token]/page.tsx:176
MEDIUM js-open-redirect app/signup/form.tsx:43
MEDIUM js-open-redirect components/billing/buy-credits-modal.tsx:98
MEDIUM js-unescaped-html-sink components/json-ld.tsx:8
MEDIUM js-unescaped-html-sink components/report/markdown-view.tsx:15
MEDIUM redos-nested-quantifier lib/careers/jobs.ts:139
MEDIUM js-unescaped-html-sink lib/careers/page-templates.ts:198
MEDIUM redos-nested-quantifier lib/emailMarkdown.ts:130
MEDIUM redos-nested-quantifier lib/lx/articleGen.ts:98
LOW secret-generic-credential app/(marketing)/docs/autoblog-webhook/page.tsx:145
LOW secret-generic-credential lib/sp/platforms/linkedin.ts:25
LOW js-dynamic-code-execution tests/careers-page-templates.test.ts:21
LOW js-dynamic-code-execution tests/careers-widget-script.test.ts:19
LOW js-dynamic-code-execution tests/careers-widget-script.test.ts:69
LOW js-dynamic-code-execution tests/contract/ad-visitor-id.test.ts:51
LOW js-dynamic-code-execution tests/contract/ad-visitor-id.test.ts:52
LOW secret-generic-credential tests/contract/posthog-integration.test.ts:13
LOW secret-generic-credential tests/lead-campaign.test.ts:16

Snippets are redacted; ThreatCrush never prints matched credential material.

@ralyodio
ralyodio marked this pull request as ready for review August 18, 2026 19:02
@ralyodio
ralyodio merged commit 0c5b34a into master Aug 18, 2026
8 checks passed
@ralyodio
ralyodio deleted the worktree-promote-sources branch August 18, 2026 19:02
ralyodio added a commit that referenced this pull request Aug 19, 2026
* docs(promote): reconcile the architecture doc with #205 and #206

The architecture doc was written in #204 and hasn't moved since, but two PRs
landed under it the same day and both changed things it describes as settled.

#205 added effectiveMix(): the ownership mix is narrowed to the classes a
campaign can actually supply before any deficit is computed. §4.1 still
described the raw configured mix, which is the version that mislabelled six
production posts as via_fallback and was three posts from silencing the
campaign. Someone reading §4 to extend the blend would have rebuilt the bug,
so the rule and the reason it exists are now in the doc rather than only in
the PR that fixed it. §4.2 gets the corresponding scope limit: covering for a
class the campaign never had is not a fallback and is not capped as one.

#206 added lib/sp/accountHealth.ts, which partly satisfies §14's "pause
campaigns after repeated provider rejections" — at the account layer, not the
campaign layer. Noted as half-built, because the Reddit provider will still
need the campaign-level version: a subreddit rejection is a destination fact,
not an account one.

Docs only; no behaviour change.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Make a Promote publication happen at most once (#AC10, #AC14)

The sweep decided and published in one pass, "claiming" a campaign by pushing
next_run_at forward. That UPDATE carried no predicate on next_run_at, so it was
a read-then-write: two sweeps that both read the row both won it. And the
worker runs the sweep on a 60s interval *and* out-of-band whenever someone
clicks "Post now", so overlapping runs are designed in, not rare.

Nothing downstream was idempotent either. If the process died after
postViaAccount() published but before the promo_post insert, there was no
record it had happened — and last_promoted_at is only stamped at the end of the
campaign, so the same link was still least-recently-promoted on the next tick
and went out again. The user sees a duplicate; the logs show nothing.

promo_job makes the intended publication the unit of work, written down before
anything is sent. Two mechanisms:

- Plan before publishing. Jobs are keyed on sha256(list, link, account,
  destination, kind, slot), where the slot is the next_run_at value the sweep
  observed as due. A racing sweep reads the same due row, derives the same
  keys, loses to the unique index, and gets nothing back — so it publishes
  nothing. Keying on the wall clock instead would give each sweep its own key
  and rebuild the bug.
- Claim by compare-and-swap: update ... where id = ? and state = 'queued'. Read
  and write are one statement, so two workers cannot both see 'queued'.

The campaign-level claim keeps its place but now carries its predicate. It is
an optimization — it avoids duplicated work. The guarantee is in the job.

AT MOST ONCE, ON PURPOSE. A job still 'publishing' past a 10 minute lease is
failed, never retried. No provider we publish through accepts an idempotency
key, so an interrupted publish has genuinely unknown outcome — it may be live.
Re-running it is the duplicate this exists to prevent. The reaper closes it with
the outcome recorded as unknown and leaves it in history for a human; the credit
is not refunded, because refunding a post that did land is the other way to be
wrong. Retry stays available for failures that provably happened before the
publish call, which is what 'retrying' and attempt_count are for.

Stray jobs are closed rather than left queued: nothing reclaims a queued job, so
one stranded by a disconnected account or a credit pause would sit there forever
misreporting the campaign as backed up.

MIGRATION FIRST. planJobs() logs loudly and publishes nothing if promo_job is
missing, rather than failing the select and stopping every campaign in silence
the way source_mix did. That is a guard, not the fix — apply
20260819120000_promote_jobs.sql before this deploys.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Claim a campaign on "still due" rather than an exact timestamp match

The claim predicate compared next_run_at to the exact string we read back. That
is correct only if a timestamptz round-trips to a byte-identical value through
PostgREST, and if it ever did not, the update would match nothing, no campaign
would be claimed, and Promote would stop posting entirely with nothing in the
logs that looks like a failure. That is the same silent-stop shape the
source_mix migration cost us a day for, and it is not worth risking to save a
predicate.

Re-asserting the condition the select already used gives the same atomicity
with none of that exposure: the minimum cadence is 300s, so whoever wins the
claim pushes next_run_at well into the future and the loser's `lte` cannot
match. The scheduling slot the idempotency key is built from is still the
observed next_run_at, so racing sweeps still agree on the slot.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* docs(promote): match §9.2 to the claim predicate that shipped

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(promote): key promo_job.user_id to profiles, like every other promote table

promo_list.user_id references public.profiles(id). Pointing promo_job at
auth.users instead would let a job outlive the profile row the rest of the
feature is keyed to, and would cascade differently from its own campaign.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant