Skip to content

feat(leads): filter the industry out of portfolio searches, and mine it instead - #138

Merged
ralyodio merged 1 commit into
masterfrom
feat/tighten-prospect-filter
Jul 28, 2026
Merged

ralyodio merged 1 commit into
masterfrom
feat/tighten-prospect-filter

Conversation

@ralyodio

Copy link
Copy Markdown
Contributor

Searching for artists returns the industry around artists. A live run for "3d artist portfolio"-shaped queries produced:

vanarts.com (art school) · therookies.co (competition) · blenderartists.org (forum) · blog.wingfox.com (tutorial blog) · adobe.com · artstation.com · unrealengine.com

Every one has a contact address, so every one sailed through discovery. Left alone, the campaign would have cold-emailed a university about an equity-only job.

Not prospects

Explicit host list plus patterns for the recurring categories — education, forums, wikis, trade press, job boards. Anchored on separators so an ordinary word inside a domain doesn't cost a real prospect: schoonerdesign.com contains "school" and is fine; newsomstudio.com contains "news" and is fine.

Patterns can't catch a school whose domain doesn't say school. vanarts.com is the Vancouver Institute of Media Arts and nothing about the string says so, so the well-known opaque ones are named explicitly. Anything similar that turns up will need adding the same way — that's a real limitation, not a solved problem.

Mined, not discarded

A forum thread isn't noise — it's where the personal sites are. An artist's own domain rarely out-ranks the community discussing their work, so these results are opened for their outbound links instead of dropped.

Keeps the exact result URL, because the thread carries the links, not the site's front page. School alumni and showcase pages work the same way.

Marketplaces and tooling vendors are deliberately not mineable: their pages link to profiles on their own domain, not to sites anyone owns.

Capped at 5 pages per tick and only while the funnel has room, so a query returning nothing but forums can't spend the whole tick on page loads.

Checks

  • tsc --noEmit clean
  • 801/801 tests pass, 50 new — the reject list is the actual hosts from the live run
  • production build compiles

🤖 Generated with Claude Code

…it instead

Searching for artists returns the industry around artists. A live run for
"3d artist portfolio" produced an art school, a competition site, a
Blender forum, a tutorial blog, Adobe, ArtStation and Unreal — every one
with a contact address on it, so every one sailed through discovery. Left
alone, the campaign would have cold-emailed a university about an
equity-only job.

Those hosts are no longer prospects. Alongside the explicit list there
are patterns for the categories that keep recurring — education, forums,
wikis, trade press, job boards — anchored on separators so an ordinary
word inside a domain does not cost a real prospect. "schoonerdesign.com"
contains "school" and is fine.

Patterns cannot catch a school whose domain does not say school, so the
well-known ones are named: vanarts.com is the Vancouver Institute of
Media Arts, and nothing about the domain says so. Anything similar that
turns up will need adding the same way.

But a forum thread is not merely noise — it is where artists' own sites
actually appear, since a personal domain rarely out-ranks the community
discussing the work. So the mineable ones are opened for their outbound
links rather than discarded, keeping the exact result URL because the
thread carries the links, not the front page. Marketplaces and tooling
vendors are deliberately excluded: their pages link to profiles on their
own domain, not to sites anyone owns.

Mining is capped per tick and only runs while the funnel still has room,
so a query returning nothing but forums cannot spend the whole tick on
page loads.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

vu1nz Security Review

0 finding(s) in PR #?

No security issues found.

@ralyodio
ralyodio merged commit ff36500 into master Jul 28, 2026
8 checks passed
@ralyodio
ralyodio deleted the feat/tighten-prospect-filter branch July 28, 2026 03:38
ralyodio added a commit that referenced this pull request Jul 28, 2026
The filter added in #138 was a list of 3D-art hostnames — Blender
forums, CG trade press, art schools, ArtStation. It worked for exactly
one campaign. The community hubs for dentists or accountants have
nothing in common with those, so the list was both endless and stale the
moment anyone pointed a campaign at a niche nobody anticipated. It is
gone.

What is left decides from the shape of a hostname rather than from
knowing an industry: .edu and .ac.*, and the structural subdomains that
mean the same thing everywhere — forum, community, wiki, jobs, support,
docs, blog, news. A forum is a forum whether the subject is character
modelling or root canals. Freelance marketplaces and link-in-bio hosts
stay, because those are cross-niche by nature, and the giants were
already covered by lib/leadCampaign.

Mining works the same way now, so a community page is opened for the
links it carries without anyone having listed it first.

The other half is simply giving up. A prospect whose site published no
address, and which the search fallback could not find either, was left
at "new" — so every tick researched it again, forever, and the funnel
filled with businesses that could never be contacted. Those are skipped
now, with the reason recorded.

The guard is a test asserting none of those hostnames appears in
discover.ts, and that the patterns contain no industry vocabulary. It
exists because the pull toward naming the site in front of you is strong
and the cost only shows up in someone else's niche.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant