Let crawlers reach the content - #57
Merged
Merged
Conversation
Googlebot made 39 requests to /live/longpoll and 8 to real pages in 24 hours, and was refused on two sitemap files and /book. - Disallow /live/ in robots.txt. LiveView's long-poll fallback carries a fresh CSRF token per request, so each fetch mints a URL never seen before; a crawler that cannot hold a websocket finds an endless supply of new pages there. robots.txt had excluded nothing at all. - Disallow /api/ too: it is for programs, it restates the listing pages, and every request spent there is one not spent on a page that can rank - Exempt known search and AI crawlers from the geo block. Some of Google's own addresses geolocate to Mountain View, so with the block on the site was quietly removing itself from search while looking perfect to every human. ClaudeBot made 3,908 content requests in the same window and hit the trap zero times, so the pages were always crawlable -- the problem was what Googlebot was allowed and pointed at. --- Pages affected: - [MCP Registry](https://ai.mcpharbor.dev/) — the Model Context Protocol server directory. - [Browse MCP servers](https://ai.mcpharbor.dev/servers) — the catalogue crawlers should be reaching. - [Sitemap](https://ai.mcpharbor.dev/sitemap.xml) — which Googlebot was being refused. - [MCP Server Optimization](https://ai.mcpharbor.dev/book) — also refused. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #56. Found by reading the nginx logs rather than guessing.
Googlebot, last 24 hours
/live/longpoll?.../robots.txt//servers/*or tool pagePlus 403s on
/sitemaps/servers-2.xml,/sitemaps/servers-3.xmland/book.For contrast, ClaudeBot made 3,908 content requests in the same window and hit the trap zero times. The pages were always crawlable; the problem was what Googlebot was allowed and pointed at.
Cause one: an endless URL supply
LiveView's long-poll fallback carries a fresh CSRF token in the query string. A crawler cannot hold a websocket, so it falls back to long-polling — and every fetch mints a URL that has never been seen before.
robots.txtsaidDisallow:with nothing after it, excluding nothing, so the crawl budget drained into the transport layer./api/is excluded too: it is for programs, it restates what the listing pages say, and every request spent there is one not spent on a page that can rank.llms.txtstays allowed — it is written for agents to read.Cause two: the geo block was refusing Googlebot
Some of Google's own addresses geolocate to Mountain View. Since the block went live at 13:32 it has served Googlebot 403 on two sitemap files and the book page. The site was removing itself from search while working perfectly for every human visitor — exactly the risk flagged when this was enabled, now measured rather than predicted.
Known search and AI crawlers now skip the geo check. Matching on user agent is spoofable, but this is a coarse filter on human traffic rather than a security control — anyone willing to forge a Googlebot header could equally use a VPN.
138 tests pass, including one that a browser in the blocked city is still refused while four crawler agents are not.
Pages affected:
🤖 Generated with Claude Code