Skip to content

Leads: find contact addresses, stop leads stranding at new - #124

Merged
ralyodio merged 1 commit into
masterfrom
fix/leads-contact-discovery
Jul 26, 2026
Merged

ralyodio merged 1 commit into
masterfrom
fix/leads-contact-discovery

Conversation

@ralyodio

Copy link
Copy Markdown
Contributor

Every lead showed "no contact address found". Two causes.

1. Nothing revisited a lead after its scan finished

Discovery queues a scan and returns straight away, so a fresh lead has no findings and no contact. Campaign ticks revisit their own leads; hand-added leads had nothing revisiting them — all 9 leads in prod sat at new next to a completed scan.

→ Adds a "Check scans (n)" button that re-researches whatever is waiting.

2. Contact discovery was too naive

Measured on the 9 real agency sites in the pipeline: 2/9. Diagnosing the pages showed why:

  • not one had a mailto: link — all contact forms
  • two hid the address behind Cloudflare data-cfemail
  • path guessing missed: /contact 404s where /contact-us works, and one address lived only on /privacy-policy

→ Decodes Cloudflare addresses, reads hello (at) example (dot) com and HTML entities, and follows the site's own contact links instead of guessing paths — ranked, not first-six-wins, because the naive crawl burned its budget on /about/are-we-fit and /company/block-inc.

Same 9 sites: 2/9 → 6/9. The 3 misses are directories, not agencies — wrong leads anyway.

642 tests pass (9 new), typecheck clean, build green. No migration.

🤖 Generated with Claude Code

Every lead showed "no contact address found". Two separate causes.

1. Nothing came back after the scan finished.

Discovery queues a free scan and returns immediately, so a new lead has no
findings and no contact yet. A campaign tick revisits its own leads; leads
added by hand from the finder had nothing revisiting them, so they sat at
"new" beside a completed scan forever. All nine leads in production were in
exactly that state. Adds a "Check scans" button that re-researches whatever
is waiting, and shows how many that is.

2. Contact discovery was too naive for how sites actually publish addresses.

Measured against the nine real agency sites in the pipeline, it found 2/9.
Diagnosing those pages showed three causes, none of which were guesses:

  - not one of them had a mailto: link; they all use contact forms
  - two hid the address behind Cloudflare's data-cfemail encoding
  - path guessing missed: /contact 404s where /contact-us works, and one
    address existed only on /privacy-policy

So it now decodes Cloudflare addresses, reads "hello (at) example (dot) com"
and HTML-entity forms, and follows the site's own contact-ish links instead
of guessing paths — the homepage nav knows where its contact page is and we
do not. Links are ranked, not first-six-wins: the naive version spent its
whole budget on /about/are-we-fit and /company/block-inc while the address
sat on /privacy-policy. Legal pages rank high deliberately, since a site
that hides its address everywhere else still has to print it there.

Same nine sites: 2/9 -> 6/9. The three misses are directories rather than
agencies, so they are the wrong leads regardless.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@ralyodio
ralyodio merged commit 94a777e into master Jul 26, 2026
8 checks passed
@ralyodio
ralyodio deleted the fix/leads-contact-discovery branch July 26, 2026 17:46
@github-actions

Copy link
Copy Markdown

vu1nz Security Review

0 finding(s) in PR #?

No security issues found.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant