ops: make the disk alarm say what is eating the disk, and not invent rates - #308
Merged
Merged
Conversation
…rates Two fixes from the alarm's first real firing. ATTRIBUTION. It fired CRITICAL at 57 GB/h and said nothing about the cause. The cause turned out to be 38 GB of docker build cache from an image build, not the databases at all, and establishing that cost a round trip. The alarm now reports docker usage and the largest databases inline, to syslog as well as mail, so the first question anyone asks is already answered. Parsing it needed --format rather than positional awk: "Build Cache" is two whitespace-separated fields in `docker system df`'s table output, which shifts every column and silently prints counts where you expect sizes. The first version reported "build-cache 40" — an object count — as though it were a size. RATE SANITY. Two samples seconds apart produce numbers like 166 GB/h from ordinary write jitter. That pollutes the growth record and can fire a spurious CRITICAL, since the 48h projection divides by it. Rates now need a 300s interval, and a too-short sample leaves the baseline alone so the next real sample still measures from a sensible point. Verified on dev2: attribution prints real sizes (Images 17.41GB, Build Cache 7.156GB, nichedb 163 GB), and two back-to-back runs both report rate=? instead of nonsense. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
| fi | ||
| # Sizes straight from the running cluster, largest first. | ||
| if docker exec -i supabase-db psql -U postgres -h localhost -d postgres -X -At \ | ||
| -c "select string_agg(datname || ' ' || pg_size_pretty(pg_database_size(datname)), ', ' order by pg_database_size(datname) desc) from pg_database where not datistemplate" >/tmp/.dbsz 2>/dev/null; then |
ThreatCrush Security Scan49 finding(s) HIGH/CRITICAL: 2 | MEDIUM: 32 | LOW: 15
Snippets are redacted; ThreatCrush never prints matched credential material. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Two fixes from the disk alarm's first real firing (#306 merged ~8 minutes before it fired).
Attribution
It fired
CRITICAL: will fill in ~26h at 57 GB/hand said nothing about the cause. The cause turned out to be 38 GB of docker build cache from an image build — not the databases at all — and establishing that cost a round trip with another session.It now reports docker usage and the largest databases inline, to syslog as well as mail:
Parsing needed
--formatrather than positional awk: "Build Cache" is two whitespace-separated fields indocker system df's table output, which shifts every column. The first version printedbuild-cache 40— an object count — as though it were a size.Rate sanity
Two samples seconds apart produce numbers like 166 GB/h from ordinary write jitter. That pollutes the growth record and can fire a spurious CRITICAL, since the 48h projection divides by the rate. Rates now require a 300s interval, and a too-short sample leaves the baseline untouched so the next real sample still measures from a sensible point.
Verified on dev2
rate=?instead of nonsense/var/log/dev2-disk.log; the legitimate 3 and 57 GB/h entries remain