Skip to content

feat(autoblog): quality gate, repair-retry, and E-E-A-T delivery fields - #196

Merged
ralyodio merged 1 commit into
masterfrom
worktree-autoblog-quality
Aug 13, 2026
Merged

ralyodio merged 1 commit into
masterfrom
worktree-autoblog-quality

Conversation

@ralyodio

Copy link
Copy Markdown
Contributor

Batch 1 of the autoblog quality work. Three self-contained changes to the generate → deliver path.

Why

generateArticle() validated exactly one thing — that the internal links the model claimed to place were physically present — then wrote the row at status='ready'. Everything else about the draft shipped unread. Any failure called failKeyword, permanently consuming the keyword and discarding a finished 4,000-word draft over one missing URL.

Meanwhile the repo already ships a slop detector we point at other people's sites (lib/audit/checks/slop.ts), and the autoblog SDK ships the heuristic gate a receiver applies to posts we deliver (@profullstack/autoblog/quality). Neither was ever aimed at our own output.

What changed

lib/lx/qualityGate.ts (new) — runs both existing gates against a draft, plus the cross-article checks a single-page scan structurally cannot do:

  • in-repo slop checks, filtered to the keys meaningful for a body fragment (filler phrasing, placeholders, misspellings, first-party evidence) — host-template concerns like viewport meta and deprecated tags are excluded, since those belong to the receiver's page, not our markdown
  • the SDK's receiver heuristics, so we fail on our own terms instead of getting 4xx'd on delivery
  • near-duplicate and repeated-opening detection against the site's recent posts

Deterministic — no LLM call. The SDK's scoreQuality() would add one, but every signal here is free and instant.

lib/lx/articleGen.ts — link validation and the gate now produce model-readable violations rather than terminal failures, and a rejected draft is regenerated once with those violations appended to the brief. Gating runs before image generation, so a rejected draft never pays for four gpt-image-2 "high"-tier calls. The accepted draft's score is stored on the row.

lib/lx/webhookDeliver.ts — posts went out with author: null, no structured data, and both dates stamped with now() at delivery time. Two consequences: a retry silently moved the article's publication date, and our own audit (content.author, content.date_signal) would flag the content we deliver on the customer's own domain. Now carries a configurable byline, BlogPosting JSON-LD, and dates read from the row.

Notes for review

  • The near-duplicate threshold is 0.55, tighter than the audit's 0.70. The audit is judging a stranger's site and wants to be sure before making the accusation; here we're judging our own generator, where that much overlap already means the post competes with one we published last week.
  • JSON-LD is emitted inside html rather than as a new Post field — that shape is shared by four Profullstack consumers, and a receiver that already renders html picks this up with no change.
  • Retry count is 2 by design. The common failure is a dropped link or one filler phrase, which one round of specific feedback fixes; past that, more sampling usually won't help and each attempt is a full long-form generation.

Verification

  • npx tsc --noEmit — clean
  • npx vitest run — 1451 passed, 0 failures (15 new)

The repo has no working lint setup: the lint script calls next lint, which Next 15 removed, and there's no eslint.config.*. Not addressed here.

Not in this PR

The larger items from the analysis: SERP-grounded research so articles carry real citations (the system prompt currently instructs the model to hedge instead of cite, and nothing in lib/lx/ ever supplies sources), and breaking the structural/phrase sameness the system prompt hard-codes into every post on a site.

Incidental finding: stripInPageAnchorLinks in articleGen.ts is dead code — never called. Its premise is also stale; the SDK's link counter already skips href="#…", so the table of contents was never counting against link density. Left in place to keep this diff focused.

🤖 Generated with Claude Code

…ip E-E-A-T fields

generateArticle() validated exactly one thing — that the internal links the
model claimed to place were physically present — and any failure marked the
keyword 'failed', throwing away a finished 4,000-word draft. Everything else
about the post shipped unread.

Meanwhile this repo already ships a slop detector we point at other people's
sites, and the autoblog SDK ships the heuristic gate a receiver applies to
posts we deliver. Neither was ever aimed at our own output.

- lib/lx/qualityGate.ts: runs the in-repo slop checks (filler, placeholders,
  misspellings, first-party evidence) plus the SDK's receiver heuristics
  against a draft, and adds the cross-article checks a single-page scan
  structurally cannot do — near-duplicate and repeated-opening detection
  against the site's recent posts. Deterministic, no LLM call.

- articleGen: link validation and the gate now produce model-readable
  violations instead of terminal failures, and a rejected draft is
  regenerated once with those violations appended to the brief. Gating runs
  before image generation, so a rejected draft never pays for four
  gpt-image-2 calls. The accepted draft's score is stored on the row.

- webhookDeliver: posts went out with author: null, no structured data, and
  both dates stamped with the delivery time — so a retry moved the article's
  publication date, and our own audit would flag the content we deliver for
  missing attribution. Now carries a configurable byline, BlogPosting
  JSON-LD, and dates from the row.

Only the near-duplicate threshold differs from the audit's (0.55 vs 0.70):
we are judging our own generator, where that much overlap already means the
post competes with one we published last week.

Typecheck clean; 1451 tests pass, 15 of them new.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

ThreatCrush Security Scan

35 finding(s)

HIGH/CRITICAL: 5 | MEDIUM: 27 | LOW: 3

Severity Rule Location
HIGH secret-generic-credential app/(marketing)/docs/autoblog-webhook/page.tsx:145
HIGH tls-verification-disabled lib/onion.ts:47
HIGH secret-generic-credential lib/sp/platforms/facebook.ts:32
HIGH secret-generic-credential lib/sp/platforms/linkedin.ts:25
HIGH sh-remote-script-execution prober/deploy/provision.sh:30
MEDIUM js-unescaped-html-sink app/(app)/admin/email-broadcast/EmailBroadcastForm.tsx:125
MEDIUM js-unescaped-html-sink app/(app)/projects/[id]/autoblog/articles/[articleId]/page.tsx:214
MEDIUM js-unescaped-html-sink app/(marketing)/blog/[slug]/page.tsx:67
MEDIUM js-unescaped-html-sink app/(marketing)/blog/[slug]/page.tsx:97
MEDIUM js-unescaped-html-sink app/(marketing)/blog/[slug]/page.tsx:104
MEDIUM js-unescaped-html-sink app/(marketing)/blog/[slug]/page.tsx:110
MEDIUM js-unescaped-html-sink app/(marketing)/recent/page.tsx:186
MEDIUM js-unescaped-html-sink app/(marketing)/recent/page.tsx:190
MEDIUM js-unescaped-html-sink app/c/[project]/[slug]/page.tsx:77
MEDIUM js-unescaped-html-sink app/c/[project]/page.tsx:57
MEDIUM js-unescaped-html-sink app/careers.js/route.ts:228
MEDIUM js-unescaped-html-sink app/careers.js/route.ts:285
MEDIUM js-unescaped-html-sink app/layout.tsx:129
MEDIUM js-open-redirect app/login/form.tsx:39
MEDIUM js-unescaped-html-sink app/r/[token]/page.tsx:176
MEDIUM js-open-redirect app/signup/form.tsx:43
MEDIUM js-open-redirect components/billing/buy-credits-modal.tsx:98
MEDIUM js-unescaped-html-sink components/json-ld.tsx:8
MEDIUM js-unescaped-html-sink components/report/markdown-view.tsx:15
MEDIUM redos-nested-quantifier lib/careers/jobs.ts:139
MEDIUM js-unescaped-html-sink lib/careers/page-templates.ts:198
MEDIUM redos-nested-quantifier lib/emailMarkdown.ts:130
MEDIUM redos-nested-quantifier lib/lx/articleGen.ts:98
MEDIUM js-dynamic-code-execution tests/careers-page-templates.test.ts:21
MEDIUM js-dynamic-code-execution tests/careers-widget-script.test.ts:69
MEDIUM js-dynamic-code-execution tests/contract/ad-visitor-id.test.ts:51
MEDIUM js-dynamic-code-execution tests/contract/ad-visitor-id.test.ts:52
LOW secret-generic-credential tests/contract/coinpay.test.ts:4
LOW secret-generic-credential tests/contract/posthog-integration.test.ts:13
LOW secret-generic-credential tests/lead-campaign.test.ts:16

Snippets are redacted; ThreatCrush never prints matched credential material.

@ralyodio
ralyodio marked this pull request as ready for review August 13, 2026 13:40
@ralyodio
ralyodio merged commit e9a1378 into master Aug 13, 2026
8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant