ops: install a disk alarm as part of box setup - #306
Merged
Merged
Conversation
dev2 runs crawlproof and nichedb on one Postgres, so disk is a SHARED failure mode: a full filesystem stops writes for every database on it and takes both sites down together. Nothing was watching it. Folded into system_setup rather than shipped as a separate script, because setup-supabase.sh is deployed as a single scp'd file and a second file is a step someone forgets. Three checks: warn at 70%, critical at 85%, and — the one that earns its keep — critical whenever the observed growth rate would fill the disk inside 48 hours. A percentage threshold is useless against a burst: this box grew 149 -> 247 GB in four hours (~22 GB/h) during nichedb's catalogue import and 15% never looked alarming. It runs from cron on the box, not from an agent session. The first version of this lived in a session monitor, died with that session, and ~100 GB of overnight growth went unnoticed as a result. Monitoring that is not on the box is not monitoring. The comments also record the estimate that nearly misled us: nichedb's "6 GB/day" was measured on Railway with a throttled disk, and the same importers ran ~40x faster on NVMe. Growth figures do not transfer between hosts. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ThreatCrush Security Scan48 finding(s) HIGH/CRITICAL: 2 | MEDIUM: 31 | LOW: 15
Snippets are redacted; ThreatCrush never prints matched credential material. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
dev2 runs crawlproof and nichedb on one Postgres, so disk is a shared failure mode — a full filesystem stops writes for every database on it and takes both sites down together. Nothing was watching it.
Folded into
system_setuprather than shipped as a separate script, sincesetup-supabase.shis deployed as a single scp'd file and a second file is a step someone forgets.Three checks
That third one is the one that earns its keep. A percentage threshold is useless against a burst: this box grew 149 → 247 GB in four hours (~22 GB/h) during nichedb's catalogue import, and "15% used" never looked alarming at any point.
Why cron and not a session watch
The first version of this lived inside an agent session's monitor. It died with the session, and ~100 GB of overnight growth went unnoticed. Monitoring that isn't on the box isn't monitoring.
The estimate that nearly misled us
The comments record it deliberately: nichedb's "6 GB/day" was measured on Railway with a throttled disk. The same importers ran roughly 40x faster on dev2's NVMe. Growth figures do not transfer between hosts — that mistake made days of headroom look like months.
Already installed and logging on dev2 (
/var/log/dev2-disk.logdoubles as the growth record, which is the right place to read a real steady-state rate off once nichedb's importers finish). The nichedb session has asked this session to own it.