Teach the ad generator to write prose, for ads that live inside content - #202
Merged
Merged
Conversation
Every format so far renders the same tiny copy set: a headline, one benefit line averaging 76 characters, and a call to action. That is the right size for a 300x250 box and far too little for a placement that sits *inside* somebody's writing -- a sponsored paragraph in a blog post, or the body of a feed item, where the ad is read rather than glanced at. So the generator now also writes two lengths of editorial prose from the destination site it already fetches for colours and copy: summary_short one or two sentences -- the inline mention summary_long two or three paragraphs -- the blog-post form The register is the point, and it is stated explicitly in the prompt: third person, factual, no second person, no call to action, no hype adjectives, "the way a journalist would describe the product in one line of a round-up". Without that instruction the model returns three restatements of the headline, which is useless -- ad voice is exactly what makes a sponsored paragraph read as an intrusion and get skipped. The no-invention rule is restated harder for these than for the display copy, because more room is more room to fabricate: if the page does not say who it is for or what it costs, neither do we. **The domain decides what is stored.** summary_domain records which site the prose was written from. A user can paste a URL, get a preview, then edit the URL before saving, and prose confidently describing the first site is worse than no prose at all. So it is only stored when the domain it was written from matches where the campaign points, and serving re-checks the same thing on read: a mismatch is treated as absent and the ad falls back to its short creative body. Regeneration re-derives the domain from the campaign's current URL, since a changed destination is exactly when a summary goes stale. Rendering: a new feed style `article` -- artwork, heading, the real paragraphs, the call to action, and a disclosure line -- in HTML, Markdown and plain text. `card` now prefers the short summary over the banner line. Both fall back to the old behaviour for the campaigns generated before this existed, because an "article" carrying a single line is just a worse card. Disclosure is strongest in the article form, deliberately. The whole point of the long form is that it reads like editorial, so the line saying it is paid for is the only thing distinguishing it from the post above it. Two structural choices worth keeping: The summary lookup is its own query rather than two more columns on the creative join in serveAd. That join is *the* serving query -- if it fails, every unit on every slot goes dark -- and these columns sit behind an `add column if not exists` in a migration applied by hand, so a deploy can run ahead of the schema. One extra round trip on the feed path, which is one fetch per publisher build rather than one per reader, is a cheap price for not being able to take banner and terminal serving down with it. The write path is resilient for the same reason: an unknown column retries without the summary rather than refusing to create the campaign. A test caught a real bug on the way: `^\s*` in the markdown-stripping regex also matches the newline *before* the anchor under the m flag, so stripping a heading or bullet swallowed the blank line separating it from the previous paragraph and silently merged two paragraphs into one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ThreatCrush Security Scan35 finding(s) HIGH/CRITICAL: 3 | MEDIUM: 23 | LOW: 9
Snippets are redacted; ThreatCrush never prints matched credential material. |
ralyodio
added a commit
that referenced
this pull request
Aug 18, 2026
#203) * Backfill the ad summaries, and refuse the ones that describe the fetch #202 added the two editorial summaries but left the 111 existing campaigns with none, so nothing rendered the long form. This backfills them. The script is TypeScript run through tsx rather than the .mjs the other scripts use, so it can import generateAdSummary from lib/ads/creative and share one prompt with the live generator. Reimplementing the prompt in plain JS would mean backfilled prose reading differently from freshly generated prose with nothing to say why -- so SUMMARY_RULES is now a single exported constant that both the full creative prompt and the summary-only prompt compose. generateAdSummary is deliberately not generateAdCreatives. That regenerates the whole concept -- colours, four creatives, and a hero image through gpt-image when a Supabase client is passed -- which for a backfill would be both far more expensive and actively destructive: it would replace copy an advertiser may have edited by hand. This reads the page, writes the prose, and touches nothing else. **The dry run is why there are guards.** Run against real data before spending anything, the very first pass produced this for a .onion campaign: "A Tor .onion address; the fetched page contains no readable text and provides no information about the site's..." That is the model doing exactly the right thing -- it was told never to invent, the page was unreadable, so it described what it actually saw. It is also the last string on earth that should be published inside somebody's blog post as the advertiser's own description of themselves. So: input guard under 200 readable characters on the page and we do not ask. You cannot summarise a page you could not read, and asking anyway is what produces prose about the fetch. output guard prose matching the shapes a model reaches for when it has nothing -- "could not be accessed", "appears to be a placeholder", "no information about" -- is discarded and the campaign keeps falling back to its short creative body. Both live in generateAdSummary rather than in the script, so a new campaign for an unreadable site is protected too, not just this one backfill. Tested against five meta-phrasings and four real ones, the last of which legitimately mentions "page" and "content" so the matcher cannot over-trigger. The pass is safe to re-run: it selects only campaigns that still lack usable prose, so an interrupted run continues where it stopped and a summary written by hand afterwards is left alone. Failures are per-campaign and never stop the pass -- a dead site in a list this old is ordinary, and that campaign simply keeps the behaviour it has today. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Loosen the summary schema caps, and truncate on a boundary The production pass failed one campaign outright with a zod `too_big` error. The cause is the trap already documented six lines above `body` in this file: structured-output SDKs strip maxLength from the schema they send and validate client-side instead, so a cap set to the length you actually want throws away the *entire generation* the moment the model runs a few characters over. I set 400 and 1600, which are the real limits, and got exactly that. The caps are now 1200 and 6000 -- generous enough that overrunning them means something is genuinely wrong -- and the real limits are applied by cleanSummary afterwards, where going over is a trim rather than an error. Length guidance stays in the field description, which is advisory rather than enforced. Retrying the failed campaign with the looser cap wrote it on the first attempt. While there: cleanSummary sliced at the byte and landed mid-word about as often as not. This prose is published as an advertiser's own description of themselves, and "a deployment tool for small te" is worse than a sentence less, so it now cuts on a sentence break where one exists and a word break otherwise. The fraction guarding the sentence break started at 0.6, which rejected a perfectly good break at 48% of the budget and fell through to the mid-sentence cut it exists to prevent; 0.4 keeps the guard for the pathological one-huge-sentence case without failing the ordinary one. Backfill result across all 111 campaigns: 97 have prose, 14 do not, 0 domain mismatches, 0 summaries describing the fetch rather than the product. The 14 are referral and signup walls -- Kraken, IBKR, a claude.ai referral link, Turso signup, Vast.ai console -- which render nothing readable to a fetcher, so the input guard declines them and they keep falling back to their short creative body. That is the correct outcome, not a gap to close. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Every format so far renders the same tiny copy set: a headline, one benefit line averaging 76 characters, and a CTA. Right size for a 300×250 box; far too little for a placement that sits inside somebody's writing — a sponsored paragraph in a blog post, or a feed item that is read rather than glanced at.
The generator now also writes two lengths of editorial prose, from the destination site it already fetches for colours and copy:
summary_shortsummary_longThe register is the point
Stated explicitly in the prompt: third person, factual, no second person, no CTA, no hype adjectives — "the way a journalist would describe the product in one line of a round-up."
Without that instruction the model returns three restatements of the headline, which is useless: ad voice is exactly what makes a sponsored paragraph read as an intrusion and get skipped.
The no-invention rule is restated harder than for display copy, because more room is more room to fabricate: if the page does not say who it is for or what it costs, neither do we.
The domain decides what is stored
summary_domainrecords which site the prose was written from.A user can paste a URL, get a preview, then edit the URL before saving — prose confidently describing the first site is worse than no prose at all. So it is stored only when the domain it was written from matches where the campaign points, and serving re-checks the same thing on read: a mismatch is treated as absent and the ad falls back to its short creative body. Regeneration re-derives the domain from the campaign's current URL, since a changed destination is exactly when a summary goes stale.
Rendering
New feed style
article— artwork, heading, real paragraphs, CTA, disclosure line — in HTML, Markdown and plain text.cardnow prefers the short summary over the banner line.Both fall back to the previous behaviour for the 106 campaigns generated before this existed, because an "article" carrying a single line is just a worse card.
Disclosure is strongest in the article form, deliberately: the whole point of the long form is that it reads like editorial, so the line saying it is paid for is the only thing distinguishing it from the post above it. Tests assert every link stays
rel="sponsored nofollow".Two structural choices
The summary lookup is its own query, not two more columns on the creative join in
serveAd. That join is the serving query — if it fails, every unit on every slot goes dark — and these columns sit behind anadd column if not existsin a hand-applied migration, so a deploy can run ahead of the schema. One extra round trip on the feed path (one fetch per publisher build, not per reader) is a cheap price for not being able to take banner and terminal serving down with it.The write path is resilient for the same reason: an unknown column retries without the summary rather than refusing to create the campaign.
A real bug the tests caught
^\s*in the markdown-stripping regex also matches the newline before the anchor under themflag, so stripping a heading or bullet swallowed the blank line separating it from the previous paragraph and silently merged two paragraphs into one.Deploying
20260818180000_ad_campaign_summaries.sql— already applied to prod, all four columns verified present. Additive and nullable, so it is safe ahead of the deploy.Existing campaigns have no prose until regenerated; everything degrades to current behaviour until they are.
Verification
1524 tests pass (14 new),
tsc --noEmitclean,next buildclean.🤖 Generated with Claude Code