Skip to content

Raise the service memory ceiling and record the new env vars - #52

Merged
lbesecker195 merged 1 commit into
mainfrom
ops/raise-memory-cap
Sep 24, 2026
Merged

lbesecker195 merged 1 commit into
mainfrom
ops/raise-memory-cap

Conversation

@lbesecker195

Copy link
Copy Markdown
Owner

The unit capped the registry at MemoryHigh=300M / MemoryMax=400M, with a comment explaining the box had "about 1 GB of memory". The box has been 12 GB for some time; the cap never moved with it, leaving ~87 MB of headroom against a 213 MB working set.

That is what caused the 19 September outage. Loading the ~180 MB geo-IP database pushed the BEAM through MemoryHigh into MemoryMax; the cgroup throttled it, so every database-backed route hung while routes touching no database kept answering 301.

Both of my earlier explanations were wrong, and I'm recording that here so the next person doesn't repeat them:

Theory Verdict
Host ran out of memory ❌ 12 GB total, 9.5 GB free, no swap
Postgres connection exhaustion ❌ steady 71/100 across 9 databases, never the constraint
cgroup MemoryMax on this service ✅ measured: 213 MB working set, 300 MB soft cap

Changes

  • Ceiling raised to 768M/1G, in line with siblings (agentads 512M, claudeslist 768M; csuite-finder and email-provider are uncapped)
  • GEO_BLOCK_ENABLED, PROBE_BATCH_SIZE, PROBE_CONCURRENCY recorded in deploy/.env.example

That second part matters: remote-install.sh installs both mcp-registry.service and mcp-registry.env over the server's copies, so settings I applied by hand would be silently reverted on the next install. They are now in the repo rather than living only on the box.

Geo blocking is enabled again on the strength of this — 180 MB inside a 1 GB ceiling is comfortable, where inside 400 MB it was not.


Pages affected:

🤖 Generated with Claude Code

The unit capped the registry at MemoryHigh=300M / MemoryMax=400M with a
comment explaining the box had about 1 GB. The box has been 12 GB for a while
and the cap never moved, leaving roughly 87 MB of headroom against a 213 MB
working set.

That is what took the site down on 19 September: loading the ~180 MB geo-IP
database pushed the BEAM through MemoryHigh into MemoryMax, the cgroup
throttled it, and every database-backed route hung while routes touching no
database kept answering. Neither of my earlier explanations -- host memory,
then Postgres connections -- was right; the host has 9 GB free and connections
sit at 71 of 100.

- Raise the ceiling to 768M/1G, in line with siblings at 512M-768M, two of
  which are uncapped
- Record GEO_BLOCK_ENABLED, PROBE_BATCH_SIZE and PROBE_CONCURRENCY in the env
  example, since remote-install.sh installs both files over the server's copies
  and settings made by hand would be lost

---

Pages affected:

- [MCP Registry](https://ai.mcpharbor.dev/) — the Model Context Protocol server directory this serves.
- [Browse MCP servers](https://ai.mcpharbor.dev/servers) — the catalogue.
- [Sitemap](https://ai.mcpharbor.dev/sitemap.xml) — grows as the probe fills in tool data.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant