Skip to content

fix(audit): isolate the link crawl so linkinator can't kill the worker - #198

Merged
ralyodio merged 1 commit into
masterfrom
fix/links-engine-process-crash
Aug 17, 2026
Merged

ralyodio merged 1 commit into
masterfrom
fix/links-engine-process-crash

Conversation

@ralyodio

Copy link
Copy Markdown
Contributor

Why

Scan run c6c19e9b reported 13 of 15 audits as Engine timed out (no response in 7 minutes). None of them timed out — none of them ran. The worker process was hard-killed about one second in.

From the Railway logs for crawlproof.com, right after spec (0.6s) and dns (0.7s) completed:

node:events:487
      throw er; // Unhandled 'error' event
DOMException [TimeoutError]: The operation was aborted due to timeout
    at Timeout._onTimeout (node:internal/abort_controller:154:9)
Emitted 'error' event on Readable instance at:
Node.js v24.15.0
[start] launching worker on :9080     <-- restart

The links engine is the source. linkinator applies its per-link timeout as an AbortSignal.timeout on the fetch, wraps the body with Readable.fromWeb(), then pipes it to the HTML parser with the 'error' handler on the destination only — linkinator/build/src/links.js:158:

source.pipe(parser).on('finish', resolve).on('error', reject);

pipe() does not forward source errors, so when the 10s abort fires mid-body the source Readable emits an unhandled 'error' event and Node hard-exits. Not catchable from linksAudit()'s try/catch. And because start.sh supervises the worker and Next.js with wait -n, the crash took the container down along with every in-flight audit.

Those 13 then sat in running — sweep() only picks up queued, so nothing retried them — until auditStuckSweep() failed and refunded them minutes later. The message was the reaper's guess, not an engine report.

The correlation holds across every recent multi-engine run:

run started total failed links
c6c19e9b Aug 17 06:22 15 13 failed
1a4366f7 Aug 17 04:49 6 0 complete
9c47420c Aug 16 11:56 15 12 failed
53793fae Aug 11 10:14 8 1 complete
60256969 Aug 10 11:16 8 0 complete

It's a race on body timing, which is why the same 15-engine fan-out succeeded on Aug 16 00:00.

What changed

  • lib/audit/links-crawl.ts — the raw crawl and its budgets, extracted. Fills an accumulator from linkinator's link/pagestart events instead of reading check()'s return value, so a crash partway through still leaves usable results.
  • lib/audit/links-crawl-child.ts — forked entry. An unhandled 'error' event arrives here as an uncaught exception, where it is genuinely recoverable: salvage the partial crawl, print it as JSON, exit 0.
  • lib/audit/links-engine.ts — forks the child and builds findings from its report. A child that dies, wedges past its budget, or emits garbage becomes a finding rather than a dead worker. A partial sweep is now reported as partial instead of claiming full coverage.
  • worker/index.ts — recoverOrphanedAudits() on boot. Any crash previously guaranteed the loss of every in-flight audit; now they are re-dispatched if still inside their stuck-timeout budget. Bounded by the existing created_at cutoff, so a crash loop cannot retry forever.

fork() inherits execArgv, so tsx's loader carries into the child under start.sh; childExecArgv() adds it explicitly for runtimes that transform TypeScript another way (vitest).

Verification

  • New tests/links-engine-crash-isolation.test.ts serves a body that stalls mid-stream — the exact fatal case. Reaching the assertions at all is the regression check.
  • Negative control: running the same crawl in-process (the old path) dies with exit 1 and the identical Unhandled 'error' event stack.
  • Verified on the real production path too — parent under npx tsx, as start.sh runs it: worker survives, audit completes with score 60 and a links.crawl_incomplete=warn finding.
  • Full suite: 1454 passed, 7 skipped. Root and worker typechecks clean. (npm run lint is broken at HEAD — next lint was removed in Next 16 — and CI does not run it.)

Follow-ups, not in this PR

  • Boot recovery could double-process an audit if a redeploy leaves two containers briefly overlapping. A heartbeat column would close that; it needs a migration. Today's behaviour loses 100% of crash-orphaned audits, so this is still strictly better.
  • The missing source.on('error') is worth sending upstream to linkinator.
  • start.sh's wait -n means any worker crash also restarts Next.js. Supervising them independently would shrink the blast radius further.
  • Unrelated: the claude engine 400s until 2026-09-01 on the Anthropic org spend cap, and dns-engine logs AI analysis failed; using baseline.

🤖 Generated with Claude Code

A slow response body took down the whole container and every in-flight
audit with it. linkinator applies its per-link `timeout` as an
AbortSignal.timeout on the fetch, wraps the body with Readable.fromWeb(),
then pipes it to the HTML parser with the 'error' handler attached to the
destination only (linkinator/build/src/links.js:158). pipe() does not
forward source errors, so when the abort fires mid-body the source
Readable emits an unhandled 'error' event and Node hard-exits. That is
not catchable from linksAudit()'s try/catch, and because start.sh
supervises the worker and Next.js with `wait -n`, the container restarted.

Scan run c6c19e9b lost 13 of 15 audits this way: only spec (0.6s) and dns
(0.7s) finished before the crash, and the remaining 13 sat in 'running'
until auditStuckSweep failed and refunded them, reporting "Engine timed
out (no response in 7 minutes)" for engines that never made a request.
Every recent multi-engine run shows the same signature — links failed
means the batch died; links completed means the batch was fine.

- lib/audit/links-crawl.ts: the raw crawl plus its budgets, filling an
  accumulator as it goes so a crash partway through keeps its results.
- lib/audit/links-crawl-child.ts: forked entry that catches the uncaught
  stream error, emits the partial crawl as JSON and exits 0.
- lib/audit/links-engine.ts: forks the child and builds findings from its
  report; a child that dies, wedges or emits garbage becomes a finding
  instead of a dead worker. Reports a partial sweep honestly rather than
  claiming full coverage.
- worker/index.ts: recover audits orphaned by a restart on boot. sweep()
  only looks at 'queued', so a crash previously guaranteed their loss.
  Re-dispatch is bounded by the existing created_at stuck cutoff, so a
  crash loop cannot retry forever.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

ThreatCrush Security Scan

35 finding(s)

HIGH/CRITICAL: 3 | MEDIUM: 23 | LOW: 9

Severity Rule Location
HIGH tls-verification-disabled lib/onion.ts:47
HIGH secret-generic-credential lib/sp/platforms/facebook.ts:32
HIGH sh-remote-script-execution prober/deploy/provision.sh:30
MEDIUM js-unescaped-html-sink app/(app)/dashboard/admin/email-broadcast/EmailBroadcastForm.tsx:125
MEDIUM js-unescaped-html-sink app/(app)/dashboard/projects/[id]/autoblog/articles/[articleId]/page.tsx:214
MEDIUM js-unescaped-html-sink app/(marketing)/blog/[slug]/page.tsx:67
MEDIUM js-unescaped-html-sink app/(marketing)/blog/[slug]/page.tsx:97
MEDIUM js-unescaped-html-sink app/(marketing)/blog/[slug]/page.tsx:104
MEDIUM js-unescaped-html-sink app/(marketing)/blog/[slug]/page.tsx:110
MEDIUM js-unescaped-html-sink app/(marketing)/recent/page.tsx:186
MEDIUM js-unescaped-html-sink app/(marketing)/recent/page.tsx:190
MEDIUM js-unescaped-html-sink app/c/[project]/[slug]/page.tsx:77
MEDIUM js-unescaped-html-sink app/c/[project]/page.tsx:57
MEDIUM js-unescaped-html-sink app/careers.js/route.ts:228
MEDIUM js-unescaped-html-sink app/careers.js/route.ts:285
MEDIUM js-unescaped-html-sink app/layout.tsx:129
MEDIUM js-open-redirect app/login/form.tsx:39
MEDIUM js-unescaped-html-sink app/r/[token]/page.tsx:176
MEDIUM js-open-redirect app/signup/form.tsx:43
MEDIUM js-open-redirect components/billing/buy-credits-modal.tsx:98
MEDIUM js-unescaped-html-sink components/json-ld.tsx:8
MEDIUM js-unescaped-html-sink components/report/markdown-view.tsx:15
MEDIUM redos-nested-quantifier lib/careers/jobs.ts:139
MEDIUM js-unescaped-html-sink lib/careers/page-templates.ts:198
MEDIUM redos-nested-quantifier lib/emailMarkdown.ts:130
MEDIUM redos-nested-quantifier lib/lx/articleGen.ts:98
LOW secret-generic-credential app/(marketing)/docs/autoblog-webhook/page.tsx:145
LOW secret-generic-credential lib/sp/platforms/linkedin.ts:25
LOW js-dynamic-code-execution tests/careers-page-templates.test.ts:21
LOW js-dynamic-code-execution tests/careers-widget-script.test.ts:19
LOW js-dynamic-code-execution tests/careers-widget-script.test.ts:69
LOW js-dynamic-code-execution tests/contract/ad-visitor-id.test.ts:51
LOW js-dynamic-code-execution tests/contract/ad-visitor-id.test.ts:52
LOW secret-generic-credential tests/contract/posthog-integration.test.ts:13
LOW secret-generic-credential tests/lead-campaign.test.ts:16

Snippets are redacted; ThreatCrush never prints matched credential material.

@ralyodio
ralyodio merged commit c28a151 into master Aug 17, 2026
8 checks passed
@ralyodio
ralyodio deleted the fix/links-engine-process-crash branch August 17, 2026 07:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant