Skip to content

feat(leads): read team pages and linked PDFs before paying to search - #150

Merged
ralyodio merged 1 commit into
masterfrom
feat/pdf-and-team-contacts
Jul 28, 2026
Merged

ralyodio merged 1 commit into
masterfrom
feat/pdf-and-team-contacts

Conversation

@ralyodio

Copy link
Copy Markdown
Contributor

The address you want is often published on the company's own domain somewhere the crawler never looked: inside a capability statement, or on a team page that names the individual rather than the shared inbox.

Ordered by cost

  1. site crawl (existing)
  2. team pages — same host, one fetch
  3. linked PDFs — same domain, no SERP spend
  4. search fallback — costs money
  5. guessed role address — last, where it belongs

Verified on a real document

https://datacenter.5nines.com/.../5NINES_Data_Center.pdf
parsed : 4396 chars
emails : inquiry@5nines.com (sameDomain=true)
phone  : 608.512.1000

None of which appears in the linking page's HTML.

Opened sparingly

A PDF is a slow, large fetch yielding at most a couple of addresses. Only files linked from a page already being read, only the first few, only up to 12 MB.

Which few is decided by filename — a capability statement earns the download, a terms-and-conditions doesn't:

pdf order: [capability-statement.pdf, terms.pdf]

A .pdf href that answers with HTML is a login wall or a 404 in disguise, dropped before parsing.

One implementation of "what is an address"

Extracted text is wrapped and handed to discoverContactEmails rather than re-implementing extraction — so obfuscation handling and same-domain ranking stay in one place instead of drifting into two.

Also: the spend alert is now scheduled

Hourly via pg_cron, matching every other job here (crawlproof-ai-spend, active). Hourly rather than daily so a runaway day is caught while it's still running; the alert de-duplicates per (day, threshold), so extra runs cost a query and send nothing.

Checks

  • tsc --noEmit clean · 949/949 tests pass, 11 new · build compiles
  • new dependency: unpdf (no native deps)

🤖 Generated with Claude Code

The address you want is often published on the company's own domain in a
place the crawler never looked. A capability statement or a media kit
carries it while the page linking to that file says nothing, and a team
page names the individual rather than the shared inbox.

Both are tried after the site crawl finds nothing and before the search
fallback, because both stay on the prospect's own domain and neither
spends a SERP call. Guessing an address remains last, where it belongs.

Documents are opened sparingly. A PDF is a slow, large fetch that yields
at most a couple of addresses, so only files linked from a page already
being read are considered, only the first few, and only up to a size
worth waiting for. Which few is decided by filename: a capability
statement earns a twelve-megabyte download and a terms-and-conditions
does not. A .pdf href that answers with HTML is a login wall or a 404 in
disguise and is dropped before parsing.

Extracted text is wrapped and handed to discoverContactEmails rather than
re-implementing address extraction, so obfuscation handling and
same-domain ranking stay in one place instead of drifting into two.

Verified against a real data-centre PDF: 4,396 characters parsed,
inquiry@5nines.com recovered on-domain, and a phone number the HTML never
mentioned.

Also schedules the AI spend warning hourly via pg_cron, matching how every
other job here is scheduled. Hourly rather than daily so a runaway day is
noticed while it is still running; the alert de-duplicates per day and
threshold, so the extra runs cost a query and send nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

vu1nz Security Review

0 finding(s) in PR #?

No security issues found.

@socket-security

Copy link
Copy Markdown

Review the following changes in direct dependencies. Learn more about Socket for GitHub.

Diff Package Supply Chain
Security
Vulnerability Quality Maintenance License
Addedunpdf@​1.8.010010010093100

View full report

@ralyodio
ralyodio merged commit bb5ba42 into master Jul 28, 2026
8 checks passed
@ralyodio
ralyodio deleted the feat/pdf-and-team-contacts branch July 28, 2026 10:44
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant