Skip to content

Teach the ad generator to write prose, for ads that live inside content - #202

Merged
ralyodio merged 1 commit into
masterfrom
ad-summaries
Aug 18, 2026
Merged

ralyodio merged 1 commit into
masterfrom
ad-summaries

Conversation

@ralyodio

Copy link
Copy Markdown
Contributor

Every format so far renders the same tiny copy set: a headline, one benefit line averaging 76 characters, and a CTA. Right size for a 300×250 box; far too little for a placement that sits inside somebody's writing — a sponsored paragraph in a blog post, or a feed item that is read rather than glanced at.

The generator now also writes two lengths of editorial prose, from the destination site it already fetches for colours and copy:

column shape use
summary_short one or two sentences the inline mention, a card body, a "sponsored by" line
summary_long two or three paragraphs the blog-post form

The register is the point

Stated explicitly in the prompt: third person, factual, no second person, no CTA, no hype adjectives — "the way a journalist would describe the product in one line of a round-up."

Without that instruction the model returns three restatements of the headline, which is useless: ad voice is exactly what makes a sponsored paragraph read as an intrusion and get skipped.

The no-invention rule is restated harder than for display copy, because more room is more room to fabricate: if the page does not say who it is for or what it costs, neither do we.

The domain decides what is stored

summary_domain records which site the prose was written from.

A user can paste a URL, get a preview, then edit the URL before saving — prose confidently describing the first site is worse than no prose at all. So it is stored only when the domain it was written from matches where the campaign points, and serving re-checks the same thing on read: a mismatch is treated as absent and the ad falls back to its short creative body. Regeneration re-derives the domain from the campaign's current URL, since a changed destination is exactly when a summary goes stale.

Rendering

New feed style article — artwork, heading, real paragraphs, CTA, disclosure line — in HTML, Markdown and plain text. card now prefers the short summary over the banner line.

Both fall back to the previous behaviour for the 106 campaigns generated before this existed, because an "article" carrying a single line is just a worse card.

Disclosure is strongest in the article form, deliberately: the whole point of the long form is that it reads like editorial, so the line saying it is paid for is the only thing distinguishing it from the post above it. Tests assert every link stays rel="sponsored nofollow".

Two structural choices

The summary lookup is its own query, not two more columns on the creative join in serveAd. That join is the serving query — if it fails, every unit on every slot goes dark — and these columns sit behind an add column if not exists in a hand-applied migration, so a deploy can run ahead of the schema. One extra round trip on the feed path (one fetch per publisher build, not per reader) is a cheap price for not being able to take banner and terminal serving down with it.

The write path is resilient for the same reason: an unknown column retries without the summary rather than refusing to create the campaign.

A real bug the tests caught

^\s* in the markdown-stripping regex also matches the newline before the anchor under the m flag, so stripping a heading or bullet swallowed the blank line separating it from the previous paragraph and silently merged two paragraphs into one.

Deploying

20260818180000_ad_campaign_summaries.sql — already applied to prod, all four columns verified present. Additive and nullable, so it is safe ahead of the deploy.

Existing campaigns have no prose until regenerated; everything degrades to current behaviour until they are.

Verification

1524 tests pass (14 new), tsc --noEmit clean, next build clean.

🤖 Generated with Claude Code

Every format so far renders the same tiny copy set: a headline, one benefit
line averaging 76 characters, and a call to action. That is the right size for
a 300x250 box and far too little for a placement that sits *inside* somebody's
writing -- a sponsored paragraph in a blog post, or the body of a feed item,
where the ad is read rather than glanced at.

So the generator now also writes two lengths of editorial prose from the
destination site it already fetches for colours and copy:

  summary_short  one or two sentences -- the inline mention
  summary_long   two or three paragraphs -- the blog-post form

The register is the point, and it is stated explicitly in the prompt: third
person, factual, no second person, no call to action, no hype adjectives, "the
way a journalist would describe the product in one line of a round-up".
Without that instruction the model returns three restatements of the headline,
which is useless -- ad voice is exactly what makes a sponsored paragraph read
as an intrusion and get skipped. The no-invention rule is restated harder for
these than for the display copy, because more room is more room to fabricate:
if the page does not say who it is for or what it costs, neither do we.

**The domain decides what is stored.** summary_domain records which site the
prose was written from. A user can paste a URL, get a preview, then edit the
URL before saving, and prose confidently describing the first site is worse
than no prose at all. So it is only stored when the domain it was written from
matches where the campaign points, and serving re-checks the same thing on
read: a mismatch is treated as absent and the ad falls back to its short
creative body. Regeneration re-derives the domain from the campaign's current
URL, since a changed destination is exactly when a summary goes stale.

Rendering: a new feed style `article` -- artwork, heading, the real paragraphs,
the call to action, and a disclosure line -- in HTML, Markdown and plain text.
`card` now prefers the short summary over the banner line. Both fall back to
the old behaviour for the campaigns generated before this existed, because an
"article" carrying a single line is just a worse card.

Disclosure is strongest in the article form, deliberately. The whole point of
the long form is that it reads like editorial, so the line saying it is paid
for is the only thing distinguishing it from the post above it.

Two structural choices worth keeping:

The summary lookup is its own query rather than two more columns on the
creative join in serveAd. That join is *the* serving query -- if it fails,
every unit on every slot goes dark -- and these columns sit behind an `add
column if not exists` in a migration applied by hand, so a deploy can run ahead
of the schema. One extra round trip on the feed path, which is one fetch per
publisher build rather than one per reader, is a cheap price for not being able
to take banner and terminal serving down with it. The write path is resilient
for the same reason: an unknown column retries without the summary rather than
refusing to create the campaign.

A test caught a real bug on the way: `^\s*` in the markdown-stripping regex
also matches the newline *before* the anchor under the m flag, so stripping a
heading or bullet swallowed the blank line separating it from the previous
paragraph and silently merged two paragraphs into one.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

ThreatCrush Security Scan

35 finding(s)

HIGH/CRITICAL: 3 | MEDIUM: 23 | LOW: 9

Severity Rule Location
HIGH tls-verification-disabled lib/onion.ts:47
HIGH secret-generic-credential lib/sp/platforms/facebook.ts:32
HIGH sh-remote-script-execution prober/deploy/provision.sh:30
MEDIUM js-unescaped-html-sink app/(app)/dashboard/admin/email-broadcast/EmailBroadcastForm.tsx:125
MEDIUM js-unescaped-html-sink app/(app)/dashboard/projects/[id]/autoblog/articles/[articleId]/page.tsx:214
MEDIUM js-unescaped-html-sink app/(marketing)/blog/[slug]/page.tsx:67
MEDIUM js-unescaped-html-sink app/(marketing)/blog/[slug]/page.tsx:97
MEDIUM js-unescaped-html-sink app/(marketing)/blog/[slug]/page.tsx:104
MEDIUM js-unescaped-html-sink app/(marketing)/blog/[slug]/page.tsx:110
MEDIUM js-unescaped-html-sink app/(marketing)/recent/page.tsx:186
MEDIUM js-unescaped-html-sink app/(marketing)/recent/page.tsx:190
MEDIUM js-unescaped-html-sink app/c/[project]/[slug]/page.tsx:77
MEDIUM js-unescaped-html-sink app/c/[project]/page.tsx:57
MEDIUM js-unescaped-html-sink app/careers.js/route.ts:228
MEDIUM js-unescaped-html-sink app/careers.js/route.ts:285
MEDIUM js-unescaped-html-sink app/layout.tsx:129
MEDIUM js-open-redirect app/login/form.tsx:39
MEDIUM js-unescaped-html-sink app/r/[token]/page.tsx:176
MEDIUM js-open-redirect app/signup/form.tsx:43
MEDIUM js-open-redirect components/billing/buy-credits-modal.tsx:98
MEDIUM js-unescaped-html-sink components/json-ld.tsx:8
MEDIUM js-unescaped-html-sink components/report/markdown-view.tsx:15
MEDIUM redos-nested-quantifier lib/careers/jobs.ts:139
MEDIUM js-unescaped-html-sink lib/careers/page-templates.ts:198
MEDIUM redos-nested-quantifier lib/emailMarkdown.ts:130
MEDIUM redos-nested-quantifier lib/lx/articleGen.ts:98
LOW secret-generic-credential app/(marketing)/docs/autoblog-webhook/page.tsx:145
LOW secret-generic-credential lib/sp/platforms/linkedin.ts:25
LOW js-dynamic-code-execution tests/careers-page-templates.test.ts:21
LOW js-dynamic-code-execution tests/careers-widget-script.test.ts:19
LOW js-dynamic-code-execution tests/careers-widget-script.test.ts:69
LOW js-dynamic-code-execution tests/contract/ad-visitor-id.test.ts:51
LOW js-dynamic-code-execution tests/contract/ad-visitor-id.test.ts:52
LOW secret-generic-credential tests/contract/posthog-integration.test.ts:13
LOW secret-generic-credential tests/lead-campaign.test.ts:16

Snippets are redacted; ThreatCrush never prints matched credential material.

@ralyodio
ralyodio merged commit 30bd7f6 into master Aug 18, 2026
8 checks passed
@ralyodio
ralyodio deleted the ad-summaries branch August 18, 2026 16:16
ralyodio added a commit that referenced this pull request Aug 18, 2026
#203)

* Backfill the ad summaries, and refuse the ones that describe the fetch

#202 added the two editorial summaries but left the 111 existing campaigns
with none, so nothing rendered the long form. This backfills them.

The script is TypeScript run through tsx rather than the .mjs the other
scripts use, so it can import generateAdSummary from lib/ads/creative and
share one prompt with the live generator. Reimplementing the prompt in plain
JS would mean backfilled prose reading differently from freshly generated
prose with nothing to say why -- so SUMMARY_RULES is now a single exported
constant that both the full creative prompt and the summary-only prompt
compose.

generateAdSummary is deliberately not generateAdCreatives. That regenerates
the whole concept -- colours, four creatives, and a hero image through
gpt-image when a Supabase client is passed -- which for a backfill would be
both far more expensive and actively destructive: it would replace copy an
advertiser may have edited by hand. This reads the page, writes the prose, and
touches nothing else.

**The dry run is why there are guards.** Run against real data before spending
anything, the very first pass produced this for a .onion campaign:

    "A Tor .onion address; the fetched page contains no readable text and
     provides no information about the site's..."

That is the model doing exactly the right thing -- it was told never to invent,
the page was unreadable, so it described what it actually saw. It is also the
last string on earth that should be published inside somebody's blog post as
the advertiser's own description of themselves. So:

  input guard   under 200 readable characters on the page and we do not ask.
                You cannot summarise a page you could not read, and asking
                anyway is what produces prose about the fetch.

  output guard  prose matching the shapes a model reaches for when it has
                nothing -- "could not be accessed", "appears to be a
                placeholder", "no information about" -- is discarded and the
                campaign keeps falling back to its short creative body.

Both live in generateAdSummary rather than in the script, so a new campaign
for an unreadable site is protected too, not just this one backfill. Tested
against five meta-phrasings and four real ones, the last of which legitimately
mentions "page" and "content" so the matcher cannot over-trigger.

The pass is safe to re-run: it selects only campaigns that still lack usable
prose, so an interrupted run continues where it stopped and a summary written
by hand afterwards is left alone. Failures are per-campaign and never stop the
pass -- a dead site in a list this old is ordinary, and that campaign simply
keeps the behaviour it has today.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Loosen the summary schema caps, and truncate on a boundary

The production pass failed one campaign outright with a zod `too_big` error.
The cause is the trap already documented six lines above `body` in this file:
structured-output SDKs strip maxLength from the schema they send and validate
client-side instead, so a cap set to the length you actually want throws away
the *entire generation* the moment the model runs a few characters over. I set
400 and 1600, which are the real limits, and got exactly that.

The caps are now 1200 and 6000 -- generous enough that overrunning them means
something is genuinely wrong -- and the real limits are applied by cleanSummary
afterwards, where going over is a trim rather than an error. Length guidance
stays in the field description, which is advisory rather than enforced.
Retrying the failed campaign with the looser cap wrote it on the first attempt.

While there: cleanSummary sliced at the byte and landed mid-word about as often
as not. This prose is published as an advertiser's own description of
themselves, and "a deployment tool for small te" is worse than a sentence less,
so it now cuts on a sentence break where one exists and a word break otherwise.
The fraction guarding the sentence break started at 0.6, which rejected a
perfectly good break at 48% of the budget and fell through to the mid-sentence
cut it exists to prevent; 0.4 keeps the guard for the pathological
one-huge-sentence case without failing the ordinary one.

Backfill result across all 111 campaigns: 97 have prose, 14 do not, 0 domain
mismatches, 0 summaries describing the fetch rather than the product. The 14
are referral and signup walls -- Kraken, IBKR, a claude.ai referral link, Turso
signup, Vast.ai console -- which render nothing readable to a fetcher, so the
input guard declines them and they keep falling back to their short creative
body. That is the correct outcome, not a gap to close.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant