Skip to content

Stop the autoblog spam: every subject, none researched alone - #212

Merged
ralyodio merged 5 commits into
masterfrom
worktree-autoblog-antispam
Aug 28, 2026
Merged

ralyodio merged 5 commits into
masterfrom
worktree-autoblog-antispam

Conversation

@ralyodio

Copy link
Copy Markdown
Contributor

The autoblog was publishing off-topic spam on all 17 active sites. Two
independent defects, neither a content feedback loop.

Seed truncation. buildSeeds() sliced the subject list to 5, the
DataForSEO loop took 3 of those, and seeds[0] was also the buyer-journey
model's entire query. 32 seeds across 9 sites had never produced a single
keyword. The top-up sweep re-ran the same truncated set, so skew compounded.

An unanchored gate. A candidate survived if it merely contained the seed
token. Real queued rows this produced: adt home security on vu1nz (CI/CD
supply-chain security), evap line first response on threatcrush, garage door opener remote on pairux, palantir technologies on bittorrented,
xoloitzcuintli price on crawlproof itself.

What changed

  • lx_site.master_keywords — durable 3-12 subject list, allocated over evenly
    and weighted toward whatever is behind; emission interleaved so the blog
    alternates topics. lx_keyword.master_keyword carries provenance.
  • A subject is never expanded alone, only crossed with a modifier. That cross
    is also the offline floor, so the queue cannot empty when upstreams are down.
  • Gate requires a subject match AND an anchor match on different words. A
    complete multi-word subject match is its own evidence; single-word subjects
    always need the anchor.
  • Stemmer fix: -es stripped two chars always, so codes never matched
    code, silently breaking every plural comparison.
  • ads_enabled (default true) — one tick joins the network: an ad unit plus
    partner and RSS Amplifier links in published articles. Slot auto-provisioned
    active, since the column default of inactive renders an empty div.
  • RSS Amplifier feeds are crawled by the worker daemon instead of fetched live
    in the publish path, with a status page at /dashboard/autoblog/crawler
    and JSON at /api/lx/feed-crawl/status.

Verification

1911 tests pass, tsc --noEmit clean, next build compiles. 65 real garbage
strings from 6 live sites are pinned as regression cases.

Both migrations are already applied to prod, and 465 off-niche queued
keywords were purged (published rows untouched). Until this merges, the old
pipeline refills those queues with the same junk.

🤖 Generated with Claude Code

https://claude.ai/code/session_013cbcXdpLFUgJZuVfCB4Gzv

ralyodio and others added 5 commits August 28, 2026 15:00
coinpayportal.com published nineteen consecutive articles about peptide
vendors — "skye peptides", "pure peptide labs", "wolverine stack peptides".
Those are competitor storefronts in an industry the site sells payment
processing to. Two independent defects produced them.

Truncation. The site listed ten subjects. buildSeeds() sliced them to five,
the DataForSEO loop took three of those, and subject #1 was handed to the
buyer-journey model as its entire query. marijuana, dispensary, weed, iptv
and torrents had never produced a single keyword. Platform-wide that is
thirty-two subjects across nine sites that have never once been written
about. The top-up sweep re-ran the same truncated set every time the queue
drained, so the skew compounded instead of averaging out.

An unanchored gate. A candidate survived if it contained the seed token.
Expanding the bare word "peptide" returns the peptide industry's own
vocabulary, and all of it passed a test that only ever asked "is this about
peptides?" — never "is this about what we sell to them?".

Both answers are the same: a subject is only ever researched crossed with a
modifier. "peptide" is not a topic; "peptide merchant account" is. The cross
is also the floor — it is built from two operator-controlled columns, so a
blog that cannot reach any upstream still publishes, and still publishes
about itself. An anchorless site now errors rather than falling back to the
expansion that caused this.

Allocation is fair-share weighted toward whatever is behind, and emission is
interleaved: fairness measured over a quarter still reads as spam if it
arrives as six peptide posts in a row. Duplicate detection moved from exact
string match to a stemmed, order-independent fingerprint, which is what
lets it see that "peptide payments" and "peptide payment" are the same
article. Both shipped, nine days apart, in May.

Covering ten subjects costs less than the three it replaced: the crosses go
into one keywordIdeas call, which takes 200 seeds.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013cbcXdpLFUgJZuVfCB4Gzv
Published articles now carry an ad unit and a short list of related posts —
partner blogs from the exchange, plus real posts from the RSS Amplifier
directory so the block is not visibly the same three domains every time.

Splitting "show ads" from "join the link network" would produce four states,
two of them incoherent (take partner backlinks, refuse to give them) and all
four needing a human to reason about. So it is one column, ads_enabled,
defaulting true. Nobody is asked. A network everybody has to opt into is one
that stays empty.

The slot is provisioned rather than requested. ad_slots.status defaults to
'inactive' and serveAd() returns null for a non-active slot before the
house-ad fallback, so a slot created at the default renders an empty div for
ever and looks exactly like a broken embed. These are created active, once
per project, reused on redelivery.

The block is built at delivery, not generation, because the hosting site is
only known once a guest post's target is resolved — and it is the host's
readers who see the ad and the host's owner who opted in, not the author's.

Titles and links come from RSS feeds crawled off the open web and land in
HTML on a customer's domain. Everything interpolated is escaped, hrefs are
allowlisted to http(s), and directory links are nofollow ugc: those
publishers agreed to nothing, and passing them ranking signal would be
spending someone else's reputation. Partner links stay followed — that
reciprocity is the point of the exchange and is recorded on both sides.

A failure anywhere in here costs the block, never the article. Publishing
without an ad costs an impression; failing to publish costs the customer the
thing they pay for.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013cbcXdpLFUgJZuVfCB4Gzv
The autoblog cites real posts from RSS Amplifier topic feeds. Those were
fetched live, inside article delivery, on a three-second timeout — a
third-party HTTP call on the critical path of the one operation a customer
pays for. And the source was invisible: a topic feed going missing showed up
as an empty block and nothing else, indistinguishable from a topic nobody
had configured.

Both have the same fix. The worker now reads feeds on its own 30-minute
tick (each source at most every 6h, 25 per sweep, least-recently-fetched
first, so work per tick is fixed however many subjects the platform grows
to), caches what came back, and keeps the record of every attempt. Delivery
reads rows. A cache miss costs one article its citation block rather than
blocking a publish on somebody else's server.

The source list is derived from every active site's master keywords on each
sweep, never curated. A blog that adds a subject has that feed crawled on
the next tick with nobody filing a request — which is what "no human
intervention" actually requires, since a curated list is a queue of requests
waiting for somebody. That is also why the status page has no "add a feed"
button: it would imply a step that does not exist.

Two judgements worth keeping. A 200 carrying no items is not a success — it
is what the directory returns for a topic nobody publishes under, and
counting it as one parks an empty topic at the front of the
least-recently-fetched queue for ever, crowding out real feeds. And a source
that gives up is parked, not deleted: a deleted row is re-derived from the
same master keyword on the very next sweep and starts failing again, which
is an infinite retry wearing the costume of a clean table.

The page says "stalled" rather than naming a cause, because it cannot tell a
stopped worker from an unreachable directory and those have different fixes.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013cbcXdpLFUgJZuVfCB4Gzv
Pulling the real lx_keyword rows for all seventeen active sites showed the
peptide articles were not a coinpayportal problem. Every blog had it.

vu1nz.com, CI/CD supply chain security, was queued to write about "adt home
security", "brinks home security" and "security public storage". threatcrush,
a SOC blog, had "evap line first response" and "emergency response liberty
county" — pregnancy tests and a Roblox game. pairux had "garage door opener
remote" and "remote control lawn mower". bittorrented had "micron technology"
and "palantir technologies". crawlproof itself had "bayesian optimization",
"scipy optimization minimize" and, genuinely, "xoloitzcuintli price".

Two fixes, both found by running the gate against those rows rather than
against fixtures.

Anchors now exclude subject words. vu1nz's subject "devops security" and its
niche both contain "security", so "adt home security" matched the subject on
`security` and then matched the anchor on the same word — satisfying a
two-part test with one token. An anchor has to be evidence the subject match
did not already provide, or it is not a second test.

And a site whose niche says nothing its subjects do not gets a commercial
vocabulary instead of an empty anchor set. Mining vu1nz's niche yields only
words its own subjects contain, and the alternative to a fallback there was
either erroring out thirteen live sites or admitting everything.

The vocabulary is narrower than the obvious one. "teams" and "business" were
in the first draft and both had to come out: "community emergency response
team" is a real queued keyword and it passed on the token `team`. A generic
English noun cannot carry half of a two-part test however natural it reads.

All 65 of those real strings are now regression cases, with the on-niche
keywords each site should still accept alongside them.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013cbcXdpLFUgJZuVfCB4Gzv
Running the gate against all 744 live queued rows rather than fixtures found
three things wrong with it, all of them over-rejection.

The stemmer was broken. "-es" always lost two characters, so "codes" became
"cod" while "code" stayed "code" and the two never matched — which meant a
subject of "promo codes" could not match the keyword "promo code", and a
coupon blog's entire queue was condemned. It also made several peptide
rejections pass for the wrong reason: "skye peptides" was failing the subject
match, not the anchor. Now strips one character by default and two only after
a sibilant, so "boxes" still gives "box".

A complete match on a multi-word subject is now its own evidence and needs no
anchor. Every bad keyword found on live sites matched exactly one generic word
out of a multi-word subject — "security" from "devops security", "remote" from
"remote control". None matched a subject in full. Meanwhile "abercrombie promo
code" matches "promo codes" completely and is exactly right for a coupon blog,
which has no narrowing term to offer because its subject IS its topic.
Single-word subjects are excluded, since that exemption is the original bug.

And the commercial vocabulary is unioned into every anchor set rather than
only used as a fallback. bl0ggers' niche yields {human, loop} once its own
subjects are removed — non-empty, so the fallback never fired, and a thin
anchor set rejected "ai writing tools".

Together these took the verdict from 92 keywords kept to 279, with all 65
pinned junk strings still rejected.

Then applied it: 465 off-niche queued rows deleted across sixteen sites.
Published rows are never touched — those are URLs that exist on somebody's
blog, and deleting the keyword would not unpublish the article, only lose the
record of it. khipu-agency has no master keywords, so it was left alone
rather than emptied on the strength of an opinion the gate cannot form.

Most sites are now under the top-up threshold and will refill through the
anchored pipeline on the next cron tick.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013cbcXdpLFUgJZuVfCB4Gzv
@ralyodio
ralyodio merged commit dba4b95 into master Aug 28, 2026
3 of 8 checks passed
@ralyodio
ralyodio deleted the worktree-autoblog-antispam branch August 28, 2026 15:35
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant