Skip to content

Let crawlers reach the content - #57

Merged
lbesecker195 merged 1 commit into
mainfrom
fix/crawler-access
Sep 24, 2026
Merged

lbesecker195 merged 1 commit into
mainfrom
fix/crawler-access

Conversation

@lbesecker195

Copy link
Copy Markdown
Owner

Closes #56. Found by reading the nginx logs rather than guessing.

Googlebot, last 24 hours

Path Requests
/live/longpoll?... 39
/robots.txt 64
/ 8
any /servers/* or tool page 0

Plus 403s on /sitemaps/servers-2.xml, /sitemaps/servers-3.xml and /book.

For contrast, ClaudeBot made 3,908 content requests in the same window and hit the trap zero times. The pages were always crawlable; the problem was what Googlebot was allowed and pointed at.

Cause one: an endless URL supply

LiveView's long-poll fallback carries a fresh CSRF token in the query string. A crawler cannot hold a websocket, so it falls back to long-polling — and every fetch mints a URL that has never been seen before. robots.txt said Disallow: with nothing after it, excluding nothing, so the crawl budget drained into the transport layer.

/api/ is excluded too: it is for programs, it restates what the listing pages say, and every request spent there is one not spent on a page that can rank. llms.txt stays allowed — it is written for agents to read.

Cause two: the geo block was refusing Googlebot

Some of Google's own addresses geolocate to Mountain View. Since the block went live at 13:32 it has served Googlebot 403 on two sitemap files and the book page. The site was removing itself from search while working perfectly for every human visitor — exactly the risk flagged when this was enabled, now measured rather than predicted.

Known search and AI crawlers now skip the geo check. Matching on user agent is spoofable, but this is a coarse filter on human traffic rather than a security control — anyone willing to forge a Googlebot header could equally use a VPN.

138 tests pass, including one that a browser in the blocked city is still refused while four crawler agents are not.


Pages affected:

🤖 Generated with Claude Code

Googlebot made 39 requests to /live/longpoll and 8 to real pages in 24 hours,
and was refused on two sitemap files and /book.

- Disallow /live/ in robots.txt. LiveView's long-poll fallback carries a fresh
  CSRF token per request, so each fetch mints a URL never seen before; a
  crawler that cannot hold a websocket finds an endless supply of new pages
  there. robots.txt had excluded nothing at all.
- Disallow /api/ too: it is for programs, it restates the listing pages, and
  every request spent there is one not spent on a page that can rank
- Exempt known search and AI crawlers from the geo block. Some of Google's own
  addresses geolocate to Mountain View, so with the block on the site was
  quietly removing itself from search while looking perfect to every human.

ClaudeBot made 3,908 content requests in the same window and hit the trap
zero times, so the pages were always crawlable -- the problem was what
Googlebot was allowed and pointed at.

---

Pages affected:

- [MCP Registry](https://ai.mcpharbor.dev/) — the Model Context Protocol server directory.
- [Browse MCP servers](https://ai.mcpharbor.dev/servers) — the catalogue crawlers should be reaching.
- [Sitemap](https://ai.mcpharbor.dev/sitemap.xml) — which Googlebot was being refused.
- [MCP Server Optimization](https://ai.mcpharbor.dev/book) — also refused.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@lbesecker195
lbesecker195 merged commit b0a0d91 into main Sep 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Crawlers are not reaching the content: a LiveView crawl trap and a geo-block 403

1 participant