Skip to content

Charge the crawler that does not announce itself - #363

Merged
ralyodio merged 4 commits into
masterfrom
worktree-rpc-fallback-fix
Sep 24, 2026
Merged

ralyodio merged 4 commits into
masterfrom
worktree-rpc-fallback-fix

Conversation

@ralyodio

Copy link
Copy Markdown
Contributor

The crawl gateway judges by user agent, and the traffic that actually costs us does not lie in one.

A headless Chromium in a Singapore cloud region was 95.8% of this site's reported HUMAN traffic — 152,131 hits in a month — and passed every check we have, because it is not pretending to be a browser. It is one. It renders the page and fires the analytics beacon, which is exactly why it was counted as a reader.

Finding it

It needed the countries/cities panels added to CrawlProof's stats API (crawlproof.com#279) to see at all:

HUMANS (158,736)                  BOTS (298)
  Singapore   152,131  95.8%        China          114  38.3%
  Japan         2,431   1.5%        India           64  21.5%

Singapore never appears in the bot panel. And of those 152,131 hits, 267 resolve to the city of Singapore — 0.18%. Country resolves, city does not, which is what a datacenter allocation looks like in a geo database.

The approach

Judge the one thing it cannot dress up. AWS, GCP and DigitalOcean each publish their address space as a machine-readable file, so the ranges are knowable exactly — no IP reputation service, no per-request lookup, no quota, and without our storing anybody's address to work it out.

Verified against the live files: 12,619 prefixes merge to 1,194 ranges; 200k lookups take 177ms (~0.9µs each).

It asks for money rather than refusing. A datacenter address is not misconduct — agents, integrations and corporate egress live there too, and an agent that wants the pages can buy the same pass a declared crawler buys. 402 is a question a caller can answer; 403 is not.

Never charged

Exclusion Why it would have hurt
/api/* Stripe, Column and Plaid call us from cloud addresses by nature, as do the x402 endpoints themselves
No client IP Railway runs on a cloud and healthchecks /. Charging that marks the deploy unhealthy and breaks every future deploy
Signed in They already pay us
Search crawlers Googlebot and the retrieval half send readers back; charging them de-indexes the site

Ships inert

CLOUD_CHARGE=pages turns it on. This can answer a real customer with a payment demand if a rule is wrong, on the live payment platform, so it is switched on deliberately with the logs watched rather than by the act of merging it. Before the first range download lands the list is empty and matches nobody, so a slow or failed fetch degrades to today's behaviour.

Also here

  • @profullstack/footprint — extracted from launchpadder's ip-protection-service. Three real defects fixed on the way out: fetch(url, {timeout: 5000}) was a no-op (not a fetch option, so the call had no deadline); google/amazon/microsoft were scored trusted, exactly inverted for scraper detection; and a failed lookup returned proxy: false, indistinguishable from clean, so an outage read as "everyone is innocent".
  • proxy.runtime.test.ts — the fixture-app file list now includes cloud-gate.ts. A missing entry there does not fail to compile; the fixture cannot resolve the module and every request answers 500, which is how that test earned its keep here.

Testing

  • 43 tests in packages/footprint, 12 in cloud-gate (every exclusion pinned individually), 5 in proxy.runtime.
  • Full suite: 5,996 pass; 3 failures in src/lib/banking/service.test.ts, verified identical on clean origin/master — pre-existing and unrelated.
  • tsc --noEmit clean.

🤖 Generated with Claude Code

ralyodio and others added 4 commits September 24, 2026 13:28
…ror"

A misconfigured Infura project took down every ETH and POL operation. It had
"require API key secret" switched on while ETHEREUM_RPC_URL carried only the
project id, so every call came back

    403  private key only is enabled in Project ID settings

Nothing failed over, because the endpoint was a single string. Balance reads
logged "checkEVMBalance: lookup failed" and sends from the wallet extension
died outright.

What made it expensive was the reporting, not the outage. estimateFees() was
called OUTSIDE the try in prepareTransaction, so the throw escaped to the
route's outer catch, which answers serverError() with no argument: the payer
saw the bare string "Internal server error", with no chain, no status and no
cause. A dead provider was indistinguishable from a bug in our own code.

So:

- evm-rpc.ts resolves an endpoint LIST per chain (configured first, then
  keyless public fallbacks) and walks it. Only transport failures fail over,
  a network error or a non-2xx status, which mean the provider is broken. A
  JSON-RPC error inside a 200 is the chain talking and is returned as-is:
  every provider would repeat it, and for a send retrying it would mean
  resubmitting.
- Errors name each host and status, never the URL, which carries the API key.
- estimateFees() moved inside the try, so a fee failure is PREPARE_FAILED
  plus the real message instead of "Internal server error".
- USDC_BASE resolves to Base, not Ethereum. These are substring matches and
  the wrong one reads a nonce from the wrong chain without failing loudly.
- "already known" from eth_sendRawTransaction is now success, not failure. It
  means the node already holds these exact signed bytes, so the money moved;
  marking the row failed was the worst of both records. The hash is keccak256
  of bytes we already hold, so it needs no provider to report.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Extracted from launchpadder-web's ip-protection-service, which had been doing
VPN/proxy/Tor detection against ip-api for a while. Not yet wired to anything:
landing it on its own so the extraction is reviewable separately from any
policy change it might later inform.

Three defects were fixed on the way out, each real rather than stylistic:

- `fetch(url, { timeout: 5000 })` does nothing. `timeout` is not a fetch
  option in any runtime, so the call had NO deadline. Now an AbortSignal.
- The reputation table scored google/amazon/microsoft/cloudflare as TRUSTED.
  For scraper detection that is exactly inverted: those are the providers
  scraping runs from. Hosting is now a positive signal of automation.
- A failed lookup returned `proxy: false, vpn: false`, indistinguishable from
  a clean result, so a dead upstream silently read as "everyone is innocent".
  Failures are marked `unknown` and score nothing.

Header signals (./headers.js) carry the cheap half and need no upstream: the
load-bearing one is that every Chromium since 76 sends Sec-Fetch-Mode and
cannot suppress it, so a request claiming Chrome/148 without it is an HTTP
client wearing a copied string. Self-declared agents are exempt from that
check — Googlebot claims Chrome and sends no Sec-Fetch-Mode, and convicting
it would cut off indexing.

Scoring is three-valued (human / unclear / automated) because two is not
enough to act well, and hosting alone never convicts: corporate VPNs,
university egress and privacy relays all resolve to infrastructure while
carrying real readers.

21 tests.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The crawl gateway judges by user agent, and the traffic that actually costs us
does not lie in one. A headless Chromium in a Singapore cloud region was 95.8%
of this site's reported HUMAN traffic — 152,131 hits in a month — and passed
every check we have, because it is not pretending to be a browser. It is one.
It renders the page and fires the analytics beacon, which is exactly why it
was counted as a reader.

The tell was geography, and it needed the countries/cities panels added to
CrawlProof's stats API to see at all: 152,131 hits from Singapore, of which
267 resolve to the city of Singapore. Country resolves, city does not, which
is what a datacenter allocation looks like in a geo database.

So judge the one thing it cannot dress up. AWS, GCP and DigitalOcean each
publish their address space as a machine-readable file, so the ranges are
knowable exactly — no IP reputation service, no per-request lookup, no quota,
and without our storing anybody's address to work it out. Live against the
real files: 12,619 prefixes merge to 1,194 ranges, and 200k lookups take
177ms.

It ASKS FOR MONEY rather than refusing. A datacenter address is not
misconduct: agents, integrations and corporate egress live there too, and an
agent that wants the pages can buy the same pass a declared crawler buys.
402 is a question a caller can answer; 403 is not.

Never charged, each for a reason that would otherwise have hurt:

  /api/*        Stripe, Column and Plaid call us from cloud addresses by
                nature, as do the x402 endpoints themselves.
  no client IP  Railway runs ON a cloud and healthchecks `/`. Charging that
                marks the deploy unhealthy and breaks every future deploy.
  signed in     They already pay us.
  search bots   Googlebot and the retrieval half send readers back; charging
                them de-indexes the site.

Ships INERT. `CLOUD_CHARGE=pages` turns it on. This can answer a real customer
with a payment demand if a rule is wrong, on the live payment platform, so it
is switched on deliberately with the logs watched rather than by the act of
merging it. Before the first range download lands the list is empty and
matches nobody, so a slow or failed fetch degrades to today's behaviour.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The runtime test builds a throwaway Next app from an explicit list of
src/lib files and runs the real HTTP pipeline against it. A module the
interception files import but the list omits does not fail to compile — the
fixture cannot resolve it, the proxy module fails to load, and EVERY request
answers 500. That is what four of these tests were reporting.

Adding cloud-gate.ts, and a comment saying why the list has to track the
imports, since the failure it produces points nowhere near the cause.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown

ThreatCrush Security Scan

331 finding(s)

HIGH/CRITICAL: 33 | MEDIUM: 40 | LOW: 258

Severity Rule Location
HIGH secret-private-key .env.example:236
HIGH secret-generic-api-key docs/API.md:430
HIGH secret-generic-api-key docs/API.md:585
HIGH secret-generic-credential docs/FIX_VERIFY_SIGNATURE.md:156
HIGH secret-generic-credential docs/integration-examples/nodejs-bot.md:225
HIGH secret-generic-api-key docs/sdk/getting-started.md:36
HIGH secret-generic-api-key docs/sdk/getting-started.md:318
HIGH secret-generic-credential packages/extension/scripts/make-screenshots.mjs:283
HIGH secret-generic-api-key packages/sdk/README.md:99
HIGH secret-generic-credential packages/sdk/README.md:122
HIGH secret-generic-credential packages/sdk/README.md:848
HIGH sh-remote-script-execution public/install.sh:167
HIGH sh-remote-script-execution public/install.sh:407
HIGH sh-remote-script-execution public/install.sh:412
HIGH sh-remote-script-execution public/install.sh:416
HIGH sh-remote-script-execution public/install.sh:761
HIGH sh-remote-script-execution public/install.sh:762
HIGH sh-remote-script-execution public/install.sh:802
HIGH sh-remote-script-execution public/install.sh:803
HIGH sh-remote-script-execution public/install.sh:804
HIGH secret-generic-credential scripts/setup-droplet.sh:609
HIGH secret-generic-api-key src/app/docs/sdk/page.tsx:135
HIGH secret-generic-api-key src/app/docs/sdk/page.tsx:214
HIGH secret-generic-credential src/app/docs/sdk/page.tsx:1001
HIGH secret-generic-credential src/app/docs/sdk/page.tsx:1022
HIGH secret-generic-api-key src/app/docs/sdk/page.tsx:1401
HIGH secret-generic-credential src/app/docs/sdk/page.tsx:1482
HIGH secret-generic-credential src/app/docs/sdk/page.tsx:1491
HIGH secret-generic-api-key src/app/docs/sdk/page.tsx:1533
HIGH secret-generic-credential src/components/docs/AuthenticationDocs.tsx:37
HIGH secret-generic-credential src/components/docs/OAuthDocs.tsx:262
HIGH secret-generic-credential supabase/config.toml:255
HIGH secret-generic-credential supabase/config.toml:287
MEDIUM manifest-install-lifecycle-script package.json:28
MEDIUM js-dynamic-code-execution packages/extension/scripts/make-screenshots.mjs:256
MEDIUM js-dynamic-code-execution packages/extension/scripts/make-screenshots.mjs:265
MEDIUM js-shell-exec-interpolation packages/sdk/bin/coinpay.js:49
MEDIUM js-shell-exec-interpolation packages/sdk/test/cli-issuer.test.js:23
MEDIUM js-shell-exec-interpolation packages/sdk/test/cli-reputation.test.js:23
MEDIUM js-shell-exec-interpolation packages/sdk/test/cli-subscription.test.js:22
MEDIUM js-shell-exec-interpolation packages/sdk/test/cli-subscription.test.js:33
MEDIUM js-shell-exec-interpolation packages/sdk/test/cli-subscription.test.js:48
MEDIUM js-shell-exec-interpolation packages/sdk/test/cli-subscription.test.js:63
MEDIUM js-shell-exec-interpolation packages/sdk/test/cli-subscription.test.js:78
MEDIUM js-shell-exec-interpolation packages/sdk/test/wallet-backup.test.js:82
MEDIUM js-shell-exec-interpolation packages/sdk/test/wallet.test.js:249
MEDIUM js-shell-exec-interpolation packages/sdk/test/wallet.test.js:280
MEDIUM insecure-temp-file public/install.sh:108
MEDIUM sh-unquoted-expansion-destructive public/install.sh:718
MEDIUM js-unescaped-html-sink public/payments.js:93

…and 281 more. Full results in the Security tab.

Snippets are redacted; ThreatCrush never prints matched credential material.

@ralyodio
ralyodio merged commit 8c3fca0 into master Sep 24, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant