feat(leads): render seed pages in a browser, and follow listings one hop - #127
Merged
Merged
Conversation
Seeding from a directory only worked when the directory was server-rendered. A growing share of the pages worth seeding ship an empty shell and load their listings over XHR, and fetching one of those returns HTTP 200 with no links -- which reads as "this directory has no businesses on it" rather than "we couldn't see them". Seeds now fall back to Chromium when a plain fetch is refused or comes back with nothing usable. Fetch still runs first because it is an order of magnitude cheaper and most directories don't need more. Rendering is deliberately generic: no per-directory API clients to write and rewrite as each site moves its endpoints. The second problem was structural. extractOutboundProspects drops same-host links as internal navigation, but a platform directory keeps every listing on its own domain -- so the businesses were being filtered out by design, and rendering alone would still have found nothing. Seeds can now take a second hop: open the listing entries and take the outbound site from each. That hop is gated on the first one coming up short. A listicle that already yielded a page of businesses has nothing to gain from opening its own internal links, and opening a dozen pages per seed would dominate a campaign tick. Seeds are user-supplied URLs and a JS-executing browser is a sharper tool than fetch, so seed loading now refuses hosts that resolve into private space -- including the cloud metadata endpoint, which was reachable before. Chromium was already in the production image, so this costs runtime memory rather than image size. Playwright moves to dependencies and is pinned to 1.60.0 to match the image it is launched from; floating it on the "next" tag would have drifted off that pin on any fresh install. It is also marked external so Next doesn't bundle a package that resolves a real binary through its own layout. Rendering does not defeat bot protection. A site behind a Cloudflare managed challenge stays blocked -- headless and headed Chromium both sit on the interstitial from a datacenter IP -- but the failure is now reported as a challenge instead of a bare HTTP 403. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
vu1nz Security Review0 finding(s) in PR #? No security issues found. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Seeding from a directory only worked when the directory was server-rendered. A growing share of the pages worth seeding — marketplace categories, artist and agency directories — ship an empty shell and load listings over XHR. Fetching one returns HTTP 200 with zero links, which reads as "this directory has no businesses on it" rather than "we couldn't see them."
Rendering
Seeds now fall back to Chromium when a plain fetch is refused or returns nothing usable. Fetch still runs first — it's an order of magnitude cheaper and most directories don't need more. Deliberately generic: no per-directory API clients to write and rewrite as sites move their endpoints.
Verified against a live JS-rendered page:
Two-hop crawling
The second problem was structural, and rendering alone would not have fixed it.
extractOutboundProspectsdrops same-host links as internal navigation (discover.ts:94) — but a platform directory keeps every listing on its own domain, so the businesses were filtered out by design. The artist's or agency's real site only appears one level deeper, on the profile page.Seeds can now take a second hop: open the listing entries, take the outbound site from each.
Gated on cost. The second hop only fires when the first comes up short (<3 businesses). Opening a dozen pages per seed would otherwise dominate a campaign tick. Verified both ways:
SSRF guard
Seeds are user-supplied URLs, and a JS-executing browser is a sharper tool than
fetch. Seed loading now refuses hosts resolving into private space — including the cloud metadata endpoint (169.254.169.254), which was reachable before this change. Verified blocked.What this does not do
It does not defeat bot protection. A site behind a Cloudflare managed challenge stays blocked — headless and headed Chromium under xvfb both sit on the interstitial from a datacenter IP, so the page never renders and the XHR never fires. The failure is now reported as a challenge rather than a bare
HTTP 403.Rendering solves "the HTML arrives empty," not "the site doesn't want us."
Infra
Chromium was already in the production image (
mcr.microsoft.com/playwright:v1.60.0-jammy), so this costs runtime memory rather than image size — no Dockerfile change.playwrightmoves todependencies(it's runtime code now) and is pinned to1.60.0to match the image it launches from. It was floating on thenexttag, which could have drifted off that pin on any fresh install and broken the browser match.serverExternalPackages— Playwright resolves a real binary through its own package layout and breaks if bundled into a server chunk.Checks
tsc --noEmitclean🤖 Generated with Claude Code