Skip to content

Backfill the ad summaries, and refuse the ones that describe the fetch - #203

Merged
ralyodio merged 2 commits into
masterfrom
ad-summary-backfill
Aug 18, 2026
Merged

ralyodio merged 2 commits into
masterfrom
ad-summary-backfill

Conversation

@ralyodio

Copy link
Copy Markdown
Contributor

#202 added the two editorial summaries but left the 111 existing campaigns with none, so nothing rendered the long form. This backfills them, and adds the guards the dry run proved were needed.

One prompt, two callers

The script is TypeScript run through tsx rather than the .mjs the other scripts use, so it can import generateAdSummary from lib/ads/creative and share one prompt with the live generator. Reimplementing it in plain JS would mean backfilled prose reading differently from freshly generated prose with nothing to say why — so SUMMARY_RULES is now a single exported constant that both prompts compose.

generateAdSummary is deliberately not generateAdCreatives. That regenerates the whole concept — colours, four creatives, and a hero image through gpt-image when a Supabase client is passed — which for a backfill would be both far more expensive and actively destructive: it would replace copy an advertiser may have edited by hand. This reads the page, writes the prose, and touches nothing else.

The dry run is why there are guards

Run against real data before spending anything, the very first pass produced this for a .onion campaign:

"A Tor .onion address; the fetched page contains no readable text and provides no information about the site's…"

That is the model doing exactly the right thing — told never to invent, handed an unreadable page, it described what it actually saw. It is also the last string on earth that should be published inside somebody's blog post as the advertiser's own description of themselves.

guard rule
input under 200 readable characters on the page → do not ask. You cannot summarise a page you could not read, and asking anyway is what produces prose about the fetch.
output prose matching the shapes a model reaches for when it has nothing — "could not be accessed", "appears to be a placeholder", "no information about" — is discarded, and the campaign keeps falling back to its short creative body.

Both live in generateAdSummary rather than in the script, so a new campaign for an unreadable site is protected too, not just this backfill. Tested against five meta-phrasings and four real ones — the last of which legitimately mentions "page" and "content", so the matcher cannot over-trigger.

Operationally

Safe to re-run: it selects only campaigns that still lack usable prose, so an interrupted run continues where it stopped and a summary written by hand afterwards is left alone. Failures are per-campaign and never stop the pass — a dead site in a list this old is ordinary, and that campaign simply keeps the behaviour it has today.

Flags: --dry-run, --limit N, --concurrency N, --env PATH.

Verification

1526 tests pass (2 new), tsc --noEmit clean, next build clean. The production pass is running as this goes up; final written/skipped counts to follow.

🤖 Generated with Claude Code

#202 added the two editorial summaries but left the 111 existing campaigns
with none, so nothing rendered the long form. This backfills them.

The script is TypeScript run through tsx rather than the .mjs the other
scripts use, so it can import generateAdSummary from lib/ads/creative and
share one prompt with the live generator. Reimplementing the prompt in plain
JS would mean backfilled prose reading differently from freshly generated
prose with nothing to say why -- so SUMMARY_RULES is now a single exported
constant that both the full creative prompt and the summary-only prompt
compose.

generateAdSummary is deliberately not generateAdCreatives. That regenerates
the whole concept -- colours, four creatives, and a hero image through
gpt-image when a Supabase client is passed -- which for a backfill would be
both far more expensive and actively destructive: it would replace copy an
advertiser may have edited by hand. This reads the page, writes the prose, and
touches nothing else.

**The dry run is why there are guards.** Run against real data before spending
anything, the very first pass produced this for a .onion campaign:

    "A Tor .onion address; the fetched page contains no readable text and
     provides no information about the site's..."

That is the model doing exactly the right thing -- it was told never to invent,
the page was unreadable, so it described what it actually saw. It is also the
last string on earth that should be published inside somebody's blog post as
the advertiser's own description of themselves. So:

  input guard   under 200 readable characters on the page and we do not ask.
                You cannot summarise a page you could not read, and asking
                anyway is what produces prose about the fetch.

  output guard  prose matching the shapes a model reaches for when it has
                nothing -- "could not be accessed", "appears to be a
                placeholder", "no information about" -- is discarded and the
                campaign keeps falling back to its short creative body.

Both live in generateAdSummary rather than in the script, so a new campaign
for an unreadable site is protected too, not just this one backfill. Tested
against five meta-phrasings and four real ones, the last of which legitimately
mentions "page" and "content" so the matcher cannot over-trigger.

The pass is safe to re-run: it selects only campaigns that still lack usable
prose, so an interrupted run continues where it stopped and a summary written
by hand afterwards is left alone. Failures are per-campaign and never stop the
pass -- a dead site in a list this old is ordinary, and that campaign simply
keeps the behaviour it has today.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

ThreatCrush Security Scan

35 finding(s)

HIGH/CRITICAL: 3 | MEDIUM: 23 | LOW: 9

Severity Rule Location
HIGH tls-verification-disabled lib/onion.ts:47
HIGH secret-generic-credential lib/sp/platforms/facebook.ts:32
HIGH sh-remote-script-execution prober/deploy/provision.sh:30
MEDIUM js-unescaped-html-sink app/(app)/dashboard/admin/email-broadcast/EmailBroadcastForm.tsx:125
MEDIUM js-unescaped-html-sink app/(app)/dashboard/projects/[id]/autoblog/articles/[articleId]/page.tsx:214
MEDIUM js-unescaped-html-sink app/(marketing)/blog/[slug]/page.tsx:67
MEDIUM js-unescaped-html-sink app/(marketing)/blog/[slug]/page.tsx:97
MEDIUM js-unescaped-html-sink app/(marketing)/blog/[slug]/page.tsx:104
MEDIUM js-unescaped-html-sink app/(marketing)/blog/[slug]/page.tsx:110
MEDIUM js-unescaped-html-sink app/(marketing)/recent/page.tsx:186
MEDIUM js-unescaped-html-sink app/(marketing)/recent/page.tsx:190
MEDIUM js-unescaped-html-sink app/c/[project]/[slug]/page.tsx:77
MEDIUM js-unescaped-html-sink app/c/[project]/page.tsx:57
MEDIUM js-unescaped-html-sink app/careers.js/route.ts:228
MEDIUM js-unescaped-html-sink app/careers.js/route.ts:285
MEDIUM js-unescaped-html-sink app/layout.tsx:129
MEDIUM js-open-redirect app/login/form.tsx:39
MEDIUM js-unescaped-html-sink app/r/[token]/page.tsx:176
MEDIUM js-open-redirect app/signup/form.tsx:43
MEDIUM js-open-redirect components/billing/buy-credits-modal.tsx:98
MEDIUM js-unescaped-html-sink components/json-ld.tsx:8
MEDIUM js-unescaped-html-sink components/report/markdown-view.tsx:15
MEDIUM redos-nested-quantifier lib/careers/jobs.ts:139
MEDIUM js-unescaped-html-sink lib/careers/page-templates.ts:198
MEDIUM redos-nested-quantifier lib/emailMarkdown.ts:130
MEDIUM redos-nested-quantifier lib/lx/articleGen.ts:98
LOW secret-generic-credential app/(marketing)/docs/autoblog-webhook/page.tsx:145
LOW secret-generic-credential lib/sp/platforms/linkedin.ts:25
LOW js-dynamic-code-execution tests/careers-page-templates.test.ts:21
LOW js-dynamic-code-execution tests/careers-widget-script.test.ts:19
LOW js-dynamic-code-execution tests/careers-widget-script.test.ts:69
LOW js-dynamic-code-execution tests/contract/ad-visitor-id.test.ts:51
LOW js-dynamic-code-execution tests/contract/ad-visitor-id.test.ts:52
LOW secret-generic-credential tests/contract/posthog-integration.test.ts:13
LOW secret-generic-credential tests/lead-campaign.test.ts:16

Snippets are redacted; ThreatCrush never prints matched credential material.

The production pass failed one campaign outright with a zod `too_big` error.
The cause is the trap already documented six lines above `body` in this file:
structured-output SDKs strip maxLength from the schema they send and validate
client-side instead, so a cap set to the length you actually want throws away
the *entire generation* the moment the model runs a few characters over. I set
400 and 1600, which are the real limits, and got exactly that.

The caps are now 1200 and 6000 -- generous enough that overrunning them means
something is genuinely wrong -- and the real limits are applied by cleanSummary
afterwards, where going over is a trim rather than an error. Length guidance
stays in the field description, which is advisory rather than enforced.
Retrying the failed campaign with the looser cap wrote it on the first attempt.

While there: cleanSummary sliced at the byte and landed mid-word about as often
as not. This prose is published as an advertiser's own description of
themselves, and "a deployment tool for small te" is worse than a sentence less,
so it now cuts on a sentence break where one exists and a word break otherwise.
The fraction guarding the sentence break started at 0.6, which rejected a
perfectly good break at 48% of the budget and fell through to the mid-sentence
cut it exists to prevent; 0.4 keeps the guard for the pathological
one-huge-sentence case without failing the ordinary one.

Backfill result across all 111 campaigns: 97 have prose, 14 do not, 0 domain
mismatches, 0 summaries describing the fetch rather than the product. The 14
are referral and signup walls -- Kraken, IBKR, a claude.ai referral link, Turso
signup, Vast.ai console -- which render nothing readable to a fetcher, so the
input guard declines them and they keep falling back to their short creative
body. That is the correct outcome, not a gap to close.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@ralyodio
ralyodio merged commit 2258f51 into master Aug 18, 2026
8 checks passed
@ralyodio
ralyodio deleted the ad-summary-backfill branch August 18, 2026 18:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant