Skip to content

mq: validate WH_MQ_MAX_BYTES_GB upper bound to prevent disk over-reservation #138

Description

@EricAndrechek

Update (2026-09-24): #583 story 5b (#612, in review) changes what this issue is about.

  • The budget is per tenant. mq.max_bytes_gb has been a settings-directory key since the settings-directory work (each folder of a nested directory has its own), not WH_MQ_MAX_BYTES_GB. With feat(mq): give every tenant a queue of its own #612, each tenant's queue is a pair of JetStream streams: an ingest stream capped at the budget and a dead-letter stream at a tenth of it. The broker takes a budget per tenant through mq.Broker.SetMaxBytes(ctx, tenant, bytes); NewEmbedded no longer takes one.
  • JetStream guards this today, and feat(mq): give every tenant a queue of its own #612 turns that off. JetStream counts every stream's cap as reserved disk and refuses a stream once the caps together pass its store limit, which defaults to 75% of the free disk at boot. So an oversized budget refuses boot (insufficient storage resources available), contrary to the premise below. Per tenant, that check caps how many tenants a disk holds: the 569th at the 1 GB minimum on an 834 GB disk, the 12th at the seed's 50 GB. So feat(mq): give every tenant a queue of its own #612 sets the embedded server's store limit out of reach, making a budget a cap and never a reservation. Nothing checks the budgets against the disk after that, which makes the premise below true.
  • What to check is the sum. That means every tenant's budget plus its dead-letter stream's cap, removed and rejected tenants included, since their queues are kept on disk until a purge route exists. The dead-letter cap is a tenth of the budget, or what the stream held when a smaller budget arrived if that is more: the mq(dlq): shrinking mq.max_bytes_gb silently deletes the oldest dead letters — investigate how the reload should treat a non-empty DLQ #532 interim guard never caps a dead-letter stream below what it holds.
  • What a full disk does today (nats-server 2.14.6, verified for feat(mq): give every tenant a queue of its own #612's docs): the failed store write is logged and never answered, so the publish times out and ingest answers 500 for every tenant on the volume, not a 503. The stream's store keeps refusing writes until WaveHouse restarts. settings-directory.mdx ("Sizing the volume") documents this and the sizing rule.
  • If a fix brings JetStream's own limit back rather than checking in WaveHouse: when a stream's store fails to open, nats-server 2.14.6 releases a reservation it never made, so its reserved count can go negative. A limit near the top of the int64 range then overflows and refuses every later stream (see the comment on JetStreamMaxStore in internal/mq/embedded.go).

Problem

internal/mq/embedded.go:NewEmbedded accepts maxBytes (configured via WH_MQ_MAX_BYTES_GB → cfg.MQ.MaxBytesGB, default 50) and passes it straight to NATS JetStream's store-reservation knob. There's no upper-bound validation against the actual filesystem capacity of storeDir.

Operational footgun: an operator setting WH_MQ_MAX_BYTES_GB=1000 on a host with 100GB of free disk on the data volume gets a JetStream that "reserves" 1TB. NATS will accept writes happily until the filesystem actually fills, then fail in less obvious ways (write errors during ingest spikes, partial stream corruption on crash) instead of refusing at boot with a clear diagnostic.

Realistic for distroless deploys where the operator can easily mis-size the volume relative to the env-var default.

Proposed Solution

At NewEmbedded time, statfs the storeDir and compare available bytes vs maxBytes:

  • maxBytes > free * 0.9 → log a WARN with both numbers and a "this will fail when filling" hint
  • maxBytes > free * 1.0 → return an error (binary won't boot, operator gets immediate signal)
  • Escape hatch: WH_MQ_OVERSUBSCRIBE_OK=1 to downgrade the error to a WARN (some operators deliberately oversubscribe knowing pruning will keep them under)

Lean toward the explicit error over silent capping — silent capping hides config errors and the actual storage knob ends up being min(env, free) rather than the documented env.

Acceptance criteria

  • Unit test: maxBytes > free * 0.9 emits WARN with both numbers in the structured fields
  • Unit test: maxBytes > free returns a non-nil error from NewEmbedded
  • Unit test: WH_MQ_OVERSUBSCRIBE_OK=1 downgrades error to WARN
  • Document under mq.max_bytes_gb in docs/src/content/docs/configuration.md
  • Document the escape hatch in docs/src/content/docs/deployment.md under the "Distroless Permission Traps" neighborhood (same footgun shape)

Context

TODO comment landed in internal/mq/embedded.go:60 as part of PR #125. Out of scope for the boot-non-fatal fix tracked by #95 but worth pulling into a follow-up. Sibling issue tracks sync_interval as a separate MQ tunable.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/configConfig file, config knobs, hot-reloadarea/docsDocumentation, site/, READMEarea/ingestIngest pipeline (Bento, batching, DLQ)documentationImprovements or additions to documentationenhancementNew feature or request

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions