Repository navigation
feat(leads): actually record who a lead is, once, across every campaign - #147
Merged
Merged
Conversation
The contacts table shipped empty and stayed empty. Discovery found people and threw them away on every run, which left the source of truth it was built to be as a schema with nothing in it. Two things had to change before anything could be written. Email was required and was the key, which assumed every person arrives with an address. Person discovery does the reverse: a directory gives a name, a title and a LinkedIn profile, and the address is what the pipeline then goes looking for. Those people could not be stored at all. Identity is now the email when there is one — two records with the same address are the same person by definition — and the normalised name and employer when there is not. Weaker, but the alternative is a fresh row for the same human every run. The key is maintained by trigger rather than generated, because finding an address for a known name has to move that row from the name form to the email form, which is exactly the transition a generated column cannot make. Merging fills gaps and does not overwrite. A scraped company name must not replace one a human typed, so every field records where it came from and a better-sourced value wins; the loser is kept in `alternates` rather than dropped, because "we saw something else" is information and a wrong overwrite is otherwise unrecoverable. Socials merge per network, so a run that finds only GitHub does not discard a known LinkedIn. People are recorded even when no address was found, which is the case that motivated all of it: a directory publishes the name and withholds the email, and waiting for an address means rediscovering the same person forever. isContactBlocked returns false when the lookup itself fails. A failed read is not permission to send, but neither is it grounds to silently block a campaign — the caller's own suppression checks still run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
vu1nz Security Review0 finding(s) in PR #? No security issues found. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The contacts table from #140 shipped empty and stayed empty — 0 rows. Discovery found people and discarded them every run, leaving the source of truth as a schema with nothing in it.
Two blockers, both in the schema I wrote
Email was required, and was the key. That assumed every person arrives with an address. Person discovery does the reverse — a directory gives a name, a title, and a LinkedIn profile, and the address is what the pipeline then goes looking for. Those people couldn't be stored at all.
Identity is now:
email:…when there's an address — two records with the same address are the same person by definitionname:…@companyotherwise — weaker, but the alternative is a fresh row for the same human every runMaintained by trigger, not generated. Finding an address for a known name has to move that row from the name form to the email form — exactly the transition a generated column can't make.
Merge fills gaps, never overwrites
A scraped company name must not replace one a human typed. Every field records its source; a better-sourced value wins, and the loser is kept in
alternatesrather than dropped — "we saw something else" is information, and a wrong overwrite is otherwise unrecoverable.Socials merge per network, so a run finding only GitHub doesn't discard a known LinkedIn.
People without addresses are recorded
The case that motivated the whole thing: a directory publishes the name and withholds the email. Waiting for an address means rediscovering the same person forever.
Checks
tsc --noEmitclean🤖 Generated with Claude Code