Skip to content

Chunk the tools sitemap after the probe outgrew it - #55

Merged
lbesecker195 merged 2 commits into
mainfrom
fix/tools-sitemap-scale
Sep 24, 2026
Merged

lbesecker195 merged 2 commits into
mainfrom
fix/tools-sitemap-scale

Conversation

@lbesecker195

Copy link
Copy Markdown
Owner

Closes #54.

The probe took the catalogue from 134 tools to 188,948 across 10,628 listings in a few hours. The tools sitemap was written as a single unchunked file back when tool coverage was a few hundred names, and that assumption is now false:

Measured Limit
URLs in one file 63,495 50,000
Size 9.6 MB 50 MB
Generation time 180 s crawlers give up long before

It also loaded full Server rows — article_content included — for every listing just to build URLs, which is where the three minutes went.

Changes

  • Chunk by listing, 150 a file (~19,000 URLs), so a listing's tools stay together and each file is comfortably legal
  • Paginate the query and select a partial struct instead of the row
  • Add Clients.ids/1, so the sitemap can name the six client pages without JSON-encoding 60,000 configs to throw them away

On indexing

An earlier revision of this branch put noindex on the client pages, reasoning that 1.1 million of them was a doorway pattern. Logan overruled that, and on inspection he is right: a VS Code page emits servers with ${input:} prompts, Zed emits context_servers, Claude Desktop bridges remote servers via mcp-remote — different config format, different path, different caveats. That is distinct content, not boilerplate, so all of it is indexed and listed.

The only noindex left in the silo is the pre-existing rule for pending and deprecated listings, and there is now a test pinning both halves: every live page indexable, pending ones not.

135 tests pass.


Pages affected:

🤖 Generated with Claude Code

Logan Besecker and others added 2 commits September 24, 2026 10:40
The probe took the catalogue from 134 tools to 188,948 across 10,628 listings
in a few hours, and the tools sitemap was a single unchunked file written when
tool coverage was a few hundred names. It reached 63,495 URLs against a 50,000
limit, 9.6 MB, and three minutes to generate, holding a web process the whole
time while loading full rows including article text.

- Chunk by listing, 1,000 a file, which is about 19,000 URLs each
- Paginate the query and select name, tools and updated_at rather than the row
- Drop client pages from the sitemap and mark them noindex

That last one is a judgement worth stating: six client pages per tool across
188,948 tools is 1.1 million near-identical URLs. That is the doorway pattern
rather than coverage, on the domain whose value is that it ranks. The pages
stay linked and useful for a reader who wants "how do I do this in Cursor" --
they simply no longer ask to be ranked. Tool pages stay indexed, and each of
those is a real tool on a real server.

---

Pages affected:

- [MCP Registry](https://ai.mcpharbor.dev/) — the Model Context Protocol server directory.
- [Sitemap](https://ai.mcpharbor.dev/sitemap.xml) — now lists chunked tool files.
- [Browse MCP servers](https://ai.mcpharbor.dev/servers) — the catalogue behind it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The probe took the catalogue from 134 tools to 188,948 across 10,628 listings
in a few hours. The tools sitemap was a single unchunked file written when tool
coverage was a few hundred names: it reached 63,495 URLs against a 50,000
limit, 9.6 MB, and three minutes to generate, loading full rows including
article text to do it.

- Chunk by listing, 150 a file, which is about 19,000 URLs each
- Paginate the query and select a partial struct rather than the row
- Add Clients.ids/1 so the sitemap can name the six client pages without
  JSON-encoding sixty thousand configs to throw them away

Client pages stay indexed and listed. I had made them noindex on a
near-duplicate argument; that was wrong. A VS Code page emits "servers" with
${input:} prompts, Zed emits context_servers, Claude Desktop bridges remote
servers through mcp-remote -- different format, path and caveats each. The only
noindex left is the pre-existing rule for pending and deprecated listings.

---

Pages affected:

- [MCP Registry](https://ai.mcpharbor.dev/) — the Model Context Protocol server directory.
- [Sitemap](https://ai.mcpharbor.dev/sitemap.xml) — now lists chunked tool files.
- [Browse MCP servers](https://ai.mcpharbor.dev/servers) — the catalogue behind it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@lbesecker195
lbesecker195 merged commit a21ec5a into main Sep 24, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

The tools sitemap outgrew its own file within hours

1 participant