Skip to content

feat(leads): render seed pages in a browser, and follow listings one hop - #127

Merged
ralyodio merged 1 commit into
masterfrom
feat/seed-render-two-hop
Jul 28, 2026
Merged

ralyodio merged 1 commit into
masterfrom
feat/seed-render-two-hop

Conversation

@ralyodio

Copy link
Copy Markdown
Contributor

Seeding from a directory only worked when the directory was server-rendered. A growing share of the pages worth seeding — marketplace categories, artist and agency directories — ship an empty shell and load listings over XHR. Fetching one returns HTTP 200 with zero links, which reads as "this directory has no businesses on it" rather than "we couldn't see them."

Rendering

Seeds now fall back to Chromium when a plain fetch is refused or returns nothing usable. Fetch still runs first — it's an order of magnitude cheaper and most directories don't need more. Deliberately generic: no per-directory API clients to write and rewrite as sites move their endpoints.

Verified against a live JS-rendered page:

listings found in HTML
plain fetch 0
rendered 10

Two-hop crawling

The second problem was structural, and rendering alone would not have fixed it. extractOutboundProspects drops same-host links as internal navigation (discover.ts:94) — but a platform directory keeps every listing on its own domain, so the businesses were filtered out by design. The artist's or agency's real site only appears one level deeper, on the profile page.

Seeds can now take a second hop: open the listing entries, take the outbound site from each.

Gated on cost. The second hop only fires when the first comes up short (<3 businesses). Opening a dozen pages per seed would otherwise dominate a campaign tick. Verified both ways:

  • thin first hop (0 found) → second hop taken, businesses found
  • Hacker News (20 outbound hosts on hop 1) → second hop skipped entirely

SSRF guard

Seeds are user-supplied URLs, and a JS-executing browser is a sharper tool than fetch. Seed loading now refuses hosts resolving into private space — including the cloud metadata endpoint (169.254.169.254), which was reachable before this change. Verified blocked.

What this does not do

It does not defeat bot protection. A site behind a Cloudflare managed challenge stays blocked — headless and headed Chromium under xvfb both sit on the interstitial from a datacenter IP, so the page never renders and the XHR never fires. The failure is now reported as a challenge rather than a bare HTTP 403.

Rendering solves "the HTML arrives empty," not "the site doesn't want us."

Infra

Chromium was already in the production image (mcr.microsoft.com/playwright:v1.60.0-jammy), so this costs runtime memory rather than image size — no Dockerfile change.

  • playwright moves to dependencies (it's runtime code now) and is pinned to 1.60.0 to match the image it launches from. It was floating on the next tag, which could have drifted off that pin on any fresh install and broken the browser match.
  • Added to serverExternalPackages — Playwright resolves a real binary through its own package layout and breaks if bundled into a server chunk.
  • Browser is launched once and shared, images/media/fonts are blocked, and it shuts down after 60s idle rather than pinning a Chromium process for the life of the container.

Checks

  • tsc --noEmit clean
  • 675/675 tests pass, 12 new
  • production build compiles
  • renderer, SSRF guard, challenge detection, and both hop paths exercised against live sites

🤖 Generated with Claude Code

Seeding from a directory only worked when the directory was
server-rendered. A growing share of the pages worth seeding ship an
empty shell and load their listings over XHR, and fetching one of those
returns HTTP 200 with no links -- which reads as "this directory has no
businesses on it" rather than "we couldn't see them".

Seeds now fall back to Chromium when a plain fetch is refused or comes
back with nothing usable. Fetch still runs first because it is an order
of magnitude cheaper and most directories don't need more. Rendering is
deliberately generic: no per-directory API clients to write and rewrite
as each site moves its endpoints.

The second problem was structural. extractOutboundProspects drops
same-host links as internal navigation, but a platform directory keeps
every listing on its own domain -- so the businesses were being filtered
out by design, and rendering alone would still have found nothing. Seeds
can now take a second hop: open the listing entries and take the
outbound site from each.

That hop is gated on the first one coming up short. A listicle that
already yielded a page of businesses has nothing to gain from opening
its own internal links, and opening a dozen pages per seed would
dominate a campaign tick.

Seeds are user-supplied URLs and a JS-executing browser is a sharper
tool than fetch, so seed loading now refuses hosts that resolve into
private space -- including the cloud metadata endpoint, which was
reachable before.

Chromium was already in the production image, so this costs runtime
memory rather than image size. Playwright moves to dependencies and is
pinned to 1.60.0 to match the image it is launched from; floating it on
the "next" tag would have drifted off that pin on any fresh install. It
is also marked external so Next doesn't bundle a package that resolves a
real binary through its own layout.

Rendering does not defeat bot protection. A site behind a Cloudflare
managed challenge stays blocked -- headless and headed Chromium both sit
on the interstitial from a datacenter IP -- but the failure is now
reported as a challenge instead of a bare HTTP 403.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

vu1nz Security Review

0 finding(s) in PR #?

No security issues found.

@ralyodio
ralyodio merged commit 936c484 into master Jul 28, 2026
8 checks passed
@ralyodio
ralyodio deleted the feat/seed-render-two-hop branch July 28, 2026 00:59
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant