Skip to content

feat(slop): free Slop Score engine — 50-page content/code/design sweep - #114

Merged
ralyodio merged 1 commit into
masterfrom
feat/slop-score
Jul 25, 2026
Merged

ralyodio merged 1 commit into
masterfrom
feat/slop-score

Conversation

@ralyodio

Copy link
Copy Markdown
Contributor

Adds Slop Score — a free, deterministic scan engine that sweeps up to 50 same-origin pages and scores how careless a site looks. 0 is pristine, 100 is maximum slop.

Output is the headline score, a per-page fix list, and systemic rollups — when 38 pages share a defect it leads with "this lives in your template, fix it once" instead of 38 separate to-dos.

What it checks

Dimension Checks
Content filler-phrase density, no first-party evidence (no numbers/quotes/tables/code/original images), thin pages, near-duplicate bodies (5-word shingle Jaccard), repeated boilerplate intros, placeholders, stale copyright, high-confidence misspellings
Code leaked {{template}} / undefined / [object Object], dev-staging hosts, console.* and TODO leftovers, duplicate + missing titles/descriptions, dead links, deprecated tags
Design missing viewport, unsized images (layout shift), placeholder alt text, stock-only imagery, inline-style density, palette / typography / !important sprawl read from fetched stylesheets

Crawl discovery is sitemap.xml first (following sitemap indexes), then breadth-first from the entry page, bounded by a 3-minute deadline so it stays inside the worker's stuck-audit cutoff.

Two design rules that are load-bearing

1. It reports observable defects, never "this was written by AI." An AI-probability score is unfalsifiable, the classifiers are unreliable (they systematically flag non-native English writers), and it would accuse paying customers. Every finding is something an owner can verify in ten seconds and fix.

2. Ambiguous markers only count in unambiguous positions. Dogfooding turned up four false positives, each now covered by a regression test:

Flagged Reality Fix
"a feature is coming soon, while your product page…" legitimate prose standalone-only matching
"seeing your brand name in an AI answer" legitimate prose standalone-only
"Does [product] use E2E encryption?", "sign me up at [your site]" editorial shorthand [insert …] only
devdrafts.netlify.app, aetherfall-ten.vercel.app scan-result text on /recent; outbound HN link attribute-only scanning; preview hosts count only in resource positions

Palette sprawl also now ignores custom-property definitions, so having design tokens (a Tailwind theme ramp) is no longer penalised as "107 distinct colors".

Wiring

  • Engine registered at cost: 0 (lib/credits.ts), worker dispatch, engines panel, report view with a SlopMeter hero stat
  • audits_engine_check migration — without it every slop insert violates the constraint, the same trap the fugu and zai migrations document. Needs applying to prod before slop scans will insert.
  • Anonymous visitors can now choose between the AEO audit and the Slop Score on the hero form (previously hardcoded to rule). Both are free, self-hosted and page-budgeted; the per-IP anonymous daily limit still applies
  • Runs no LLM, so it can't be stalled by the shared-provider-quota outages that periodically stop Autoblog

Verification

38 new unit tests, full suite green (472 passing), clean typecheck, successful production build, plus live sweeps:

Site Score
crawlproof.com 24 — Clean
ugig.net 29 — Some slop
news.ycombinator.com 58 — Sloppy (all genuine: <center>, no viewport, / ≈ /news at 88%)

It found a real duplicate-meta-description bug on three of our own pages and unsized images across 38.

Branched from master rather than the current working branch so the diff is self-contained; the DeepSeek V3→V4 label change stays with its own PR.

🤖 Generated with Claude Code

Adds a free, deterministic scan engine that sweeps up to 50 same-origin
pages and scores how careless a site looks: 0 is pristine, 100 is maximum
slop. Output is the headline score, per-page fix lists, and systemic
rollups for defects that live in a shared template.

Discovery is sitemap.xml first (following sitemap indexes), then
breadth-first from the entry page, bounded by a 3-minute deadline so it
stays inside the worker's stuck-audit cutoff. A few same-origin
stylesheets are fetched so the design checks can see palette,
typography, and !important sprawl, which is invisible from HTML alone.

Analyzers (lib/audit/checks/slop.ts) cover three dimensions:
  content — filler-phrase density, no first-party evidence, thin pages,
            near-duplicate bodies (5-word shingle Jaccard), boilerplate
            intros, placeholders, stale copyright, misspellings
  code    — leaked template variables, dev/staging hosts, console.* and
            TODO leftovers, duplicate/missing metadata, dead links,
            deprecated tags
  design  — missing viewport, unsized images, placeholder alt text,
            stock-only imagery, inline-style density, style sprawl

Two design rules are load-bearing:

It reports observable defects, never "this was written by AI". An
AI-probability score is unfalsifiable, the classifiers are unreliable
(they systematically flag non-native English writers), and it would
accuse paying customers. Every finding is something an owner can verify
in ten seconds and fix.

Ambiguous markers only count in unambiguous positions. Dogfooding on our
own blog and on Hacker News turned up four false positives, each now
covered by a regression test: "a feature is coming soon" and "your brand
name in an AI answer" in running prose, "[product]"/"[your site]" as
editorial shorthand, and .netlify.app/.vercel.app appearing as scan-result
text or as an outbound link rather than a leaked URL. Hence
standalone-only phrase matching and attribute-only host scanning, with
preview hosts counted only in resource positions. Palette sprawl also
ignores custom-property definitions, so having design tokens is no longer
penalised.

Wiring: engine registered at cost 0 (lib/credits.ts), worker dispatch,
engines panel, report view with a SlopMeter hero stat, and an
audits_engine_check migration — without it every 'slop' insert would
violate the constraint, the same trap the fugu and zai migrations
document. Anonymous visitors can now pick between the AEO audit and the
Slop Score on the hero form; both are free, self-hosted, and
page-budgeted, and the per-IP anonymous limit still applies. Runs no LLM,
so it is immune to the shared-provider-quota outages that stall Autoblog.

Verified with 38 unit tests plus live sweeps of crawlproof.com (24,
Clean), ugig.net (29) and news.ycombinator.com (58, Sloppy). Found a real
duplicate-meta-description bug on three of our pages and unsized images
across 38.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

vu1nz Security Review

0 finding(s) in PR #?

No security issues found.

@ralyodio
ralyodio merged commit 6c1a886 into master Jul 25, 2026
8 checks passed
@ralyodio
ralyodio deleted the feat/slop-score branch July 25, 2026 19:36
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant