Stop the autoblog spam: every subject, none researched alone - #212
Merged
Merged
Conversation
coinpayportal.com published nineteen consecutive articles about peptide vendors — "skye peptides", "pure peptide labs", "wolverine stack peptides". Those are competitor storefronts in an industry the site sells payment processing to. Two independent defects produced them. Truncation. The site listed ten subjects. buildSeeds() sliced them to five, the DataForSEO loop took three of those, and subject #1 was handed to the buyer-journey model as its entire query. marijuana, dispensary, weed, iptv and torrents had never produced a single keyword. Platform-wide that is thirty-two subjects across nine sites that have never once been written about. The top-up sweep re-ran the same truncated set every time the queue drained, so the skew compounded instead of averaging out. An unanchored gate. A candidate survived if it contained the seed token. Expanding the bare word "peptide" returns the peptide industry's own vocabulary, and all of it passed a test that only ever asked "is this about peptides?" — never "is this about what we sell to them?". Both answers are the same: a subject is only ever researched crossed with a modifier. "peptide" is not a topic; "peptide merchant account" is. The cross is also the floor — it is built from two operator-controlled columns, so a blog that cannot reach any upstream still publishes, and still publishes about itself. An anchorless site now errors rather than falling back to the expansion that caused this. Allocation is fair-share weighted toward whatever is behind, and emission is interleaved: fairness measured over a quarter still reads as spam if it arrives as six peptide posts in a row. Duplicate detection moved from exact string match to a stemmed, order-independent fingerprint, which is what lets it see that "peptide payments" and "peptide payment" are the same article. Both shipped, nine days apart, in May. Covering ten subjects costs less than the three it replaced: the crosses go into one keywordIdeas call, which takes 200 seeds. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013cbcXdpLFUgJZuVfCB4Gzv
Published articles now carry an ad unit and a short list of related posts — partner blogs from the exchange, plus real posts from the RSS Amplifier directory so the block is not visibly the same three domains every time. Splitting "show ads" from "join the link network" would produce four states, two of them incoherent (take partner backlinks, refuse to give them) and all four needing a human to reason about. So it is one column, ads_enabled, defaulting true. Nobody is asked. A network everybody has to opt into is one that stays empty. The slot is provisioned rather than requested. ad_slots.status defaults to 'inactive' and serveAd() returns null for a non-active slot before the house-ad fallback, so a slot created at the default renders an empty div for ever and looks exactly like a broken embed. These are created active, once per project, reused on redelivery. The block is built at delivery, not generation, because the hosting site is only known once a guest post's target is resolved — and it is the host's readers who see the ad and the host's owner who opted in, not the author's. Titles and links come from RSS feeds crawled off the open web and land in HTML on a customer's domain. Everything interpolated is escaped, hrefs are allowlisted to http(s), and directory links are nofollow ugc: those publishers agreed to nothing, and passing them ranking signal would be spending someone else's reputation. Partner links stay followed — that reciprocity is the point of the exchange and is recorded on both sides. A failure anywhere in here costs the block, never the article. Publishing without an ad costs an impression; failing to publish costs the customer the thing they pay for. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013cbcXdpLFUgJZuVfCB4Gzv
The autoblog cites real posts from RSS Amplifier topic feeds. Those were fetched live, inside article delivery, on a three-second timeout — a third-party HTTP call on the critical path of the one operation a customer pays for. And the source was invisible: a topic feed going missing showed up as an empty block and nothing else, indistinguishable from a topic nobody had configured. Both have the same fix. The worker now reads feeds on its own 30-minute tick (each source at most every 6h, 25 per sweep, least-recently-fetched first, so work per tick is fixed however many subjects the platform grows to), caches what came back, and keeps the record of every attempt. Delivery reads rows. A cache miss costs one article its citation block rather than blocking a publish on somebody else's server. The source list is derived from every active site's master keywords on each sweep, never curated. A blog that adds a subject has that feed crawled on the next tick with nobody filing a request — which is what "no human intervention" actually requires, since a curated list is a queue of requests waiting for somebody. That is also why the status page has no "add a feed" button: it would imply a step that does not exist. Two judgements worth keeping. A 200 carrying no items is not a success — it is what the directory returns for a topic nobody publishes under, and counting it as one parks an empty topic at the front of the least-recently-fetched queue for ever, crowding out real feeds. And a source that gives up is parked, not deleted: a deleted row is re-derived from the same master keyword on the very next sweep and starts failing again, which is an infinite retry wearing the costume of a clean table. The page says "stalled" rather than naming a cause, because it cannot tell a stopped worker from an unreachable directory and those have different fixes. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013cbcXdpLFUgJZuVfCB4Gzv
Pulling the real lx_keyword rows for all seventeen active sites showed the peptide articles were not a coinpayportal problem. Every blog had it. vu1nz.com, CI/CD supply chain security, was queued to write about "adt home security", "brinks home security" and "security public storage". threatcrush, a SOC blog, had "evap line first response" and "emergency response liberty county" — pregnancy tests and a Roblox game. pairux had "garage door opener remote" and "remote control lawn mower". bittorrented had "micron technology" and "palantir technologies". crawlproof itself had "bayesian optimization", "scipy optimization minimize" and, genuinely, "xoloitzcuintli price". Two fixes, both found by running the gate against those rows rather than against fixtures. Anchors now exclude subject words. vu1nz's subject "devops security" and its niche both contain "security", so "adt home security" matched the subject on `security` and then matched the anchor on the same word — satisfying a two-part test with one token. An anchor has to be evidence the subject match did not already provide, or it is not a second test. And a site whose niche says nothing its subjects do not gets a commercial vocabulary instead of an empty anchor set. Mining vu1nz's niche yields only words its own subjects contain, and the alternative to a fallback there was either erroring out thirteen live sites or admitting everything. The vocabulary is narrower than the obvious one. "teams" and "business" were in the first draft and both had to come out: "community emergency response team" is a real queued keyword and it passed on the token `team`. A generic English noun cannot carry half of a two-part test however natural it reads. All 65 of those real strings are now regression cases, with the on-niche keywords each site should still accept alongside them. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013cbcXdpLFUgJZuVfCB4Gzv
Running the gate against all 744 live queued rows rather than fixtures found
three things wrong with it, all of them over-rejection.
The stemmer was broken. "-es" always lost two characters, so "codes" became
"cod" while "code" stayed "code" and the two never matched — which meant a
subject of "promo codes" could not match the keyword "promo code", and a
coupon blog's entire queue was condemned. It also made several peptide
rejections pass for the wrong reason: "skye peptides" was failing the subject
match, not the anchor. Now strips one character by default and two only after
a sibilant, so "boxes" still gives "box".
A complete match on a multi-word subject is now its own evidence and needs no
anchor. Every bad keyword found on live sites matched exactly one generic word
out of a multi-word subject — "security" from "devops security", "remote" from
"remote control". None matched a subject in full. Meanwhile "abercrombie promo
code" matches "promo codes" completely and is exactly right for a coupon blog,
which has no narrowing term to offer because its subject IS its topic.
Single-word subjects are excluded, since that exemption is the original bug.
And the commercial vocabulary is unioned into every anchor set rather than
only used as a fallback. bl0ggers' niche yields {human, loop} once its own
subjects are removed — non-empty, so the fallback never fired, and a thin
anchor set rejected "ai writing tools".
Together these took the verdict from 92 keywords kept to 279, with all 65
pinned junk strings still rejected.
Then applied it: 465 off-niche queued rows deleted across sixteen sites.
Published rows are never touched — those are URLs that exist on somebody's
blog, and deleting the keyword would not unpublish the article, only lose the
record of it. khipu-agency has no master keywords, so it was left alone
rather than emptied on the strength of an opinion the gate cannot form.
Most sites are now under the top-up threshold and will refill through the
anchored pipeline on the next cron tick.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013cbcXdpLFUgJZuVfCB4Gzv
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The autoblog was publishing off-topic spam on all 17 active sites. Two
independent defects, neither a content feedback loop.
Seed truncation.
buildSeeds()sliced the subject list to 5, theDataForSEO loop took 3 of those, and
seeds[0]was also the buyer-journeymodel's entire query. 32 seeds across 9 sites had never produced a single
keyword. The top-up sweep re-ran the same truncated set, so skew compounded.
An unanchored gate. A candidate survived if it merely contained the seed
token. Real queued rows this produced:
adt home securityon vu1nz (CI/CDsupply-chain security),
evap line first responseon threatcrush,garage door opener remoteon pairux,palantir technologieson bittorrented,xoloitzcuintli priceon crawlproof itself.What changed
lx_site.master_keywords— durable 3-12 subject list, allocated over evenlyand weighted toward whatever is behind; emission interleaved so the blog
alternates topics.
lx_keyword.master_keywordcarries provenance.is also the offline floor, so the queue cannot empty when upstreams are down.
complete multi-word subject match is its own evidence; single-word subjects
always need the anchor.
-esstripped two chars always, socodesnever matchedcode, silently breaking every plural comparison.ads_enabled(default true) — one tick joins the network: an ad unit pluspartner and RSS Amplifier links in published articles. Slot auto-provisioned
active, since the column default of
inactiverenders an empty div.in the publish path, with a status page at
/dashboard/autoblog/crawlerand JSON at
/api/lx/feed-crawl/status.Verification
1911 tests pass,
tsc --noEmitclean,next buildcompiles. 65 real garbagestrings from 6 live sites are pinned as regression cases.
Both migrations are already applied to prod, and 465 off-niche queued
keywords were purged (published rows untouched). Until this merges, the old
pipeline refills those queues with the same junk.
🤖 Generated with Claude Code
https://claude.ai/code/session_013cbcXdpLFUgJZuVfCB4Gzv