Repository navigation
feat(leads): read team pages and linked PDFs before paying to search - #150
Merged
Merged
Conversation
The address you want is often published on the company's own domain in a place the crawler never looked. A capability statement or a media kit carries it while the page linking to that file says nothing, and a team page names the individual rather than the shared inbox. Both are tried after the site crawl finds nothing and before the search fallback, because both stay on the prospect's own domain and neither spends a SERP call. Guessing an address remains last, where it belongs. Documents are opened sparingly. A PDF is a slow, large fetch that yields at most a couple of addresses, so only files linked from a page already being read are considered, only the first few, and only up to a size worth waiting for. Which few is decided by filename: a capability statement earns a twelve-megabyte download and a terms-and-conditions does not. A .pdf href that answers with HTML is a login wall or a 404 in disguise and is dropped before parsing. Extracted text is wrapped and handed to discoverContactEmails rather than re-implementing address extraction, so obfuscation handling and same-domain ranking stay in one place instead of drifting into two. Verified against a real data-centre PDF: 4,396 characters parsed, inquiry@5nines.com recovered on-domain, and a phone number the HTML never mentioned. Also schedules the AI spend warning hourly via pg_cron, matching how every other job here is scheduled. Hourly rather than daily so a runaway day is noticed while it is still running; the alert de-duplicates per day and threshold, so the extra runs cost a query and send nothing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
vu1nz Security Review0 finding(s) in PR #? No security issues found. |
|
Review the following changes in direct dependencies. Learn more about Socket for GitHub.
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The address you want is often published on the company's own domain somewhere the crawler never looked: inside a capability statement, or on a team page that names the individual rather than the shared inbox.
Ordered by cost
Verified on a real document
None of which appears in the linking page's HTML.
Opened sparingly
A PDF is a slow, large fetch yielding at most a couple of addresses. Only files linked from a page already being read, only the first few, only up to 12 MB.
Which few is decided by filename — a capability statement earns the download, a terms-and-conditions doesn't:
A
.pdfhref that answers with HTML is a login wall or a 404 in disguise, dropped before parsing.One implementation of "what is an address"
Extracted text is wrapped and handed to
discoverContactEmailsrather than re-implementing extraction — so obfuscation handling and same-domain ranking stay in one place instead of drifting into two.Also: the spend alert is now scheduled
Hourly via pg_cron, matching every other job here (
crawlproof-ai-spend, active). Hourly rather than daily so a runaway day is caught while it's still running; the alert de-duplicates per(day, threshold), so extra runs cost a query and send nothing.Checks
tsc --noEmitclean · 949/949 tests pass, 11 new · build compilesunpdf(no native deps)🤖 Generated with Claude Code