feat(settings): nested settings directory with a per-tenant registry - #598
Conversation
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Organization UI Review profile: ASSERTIVE Plan: Advanced Run ID: 📒 Files selected for processing (49)
Included review availability: Your plan provides up to 1 included review per hour; 0 remain after this review. 📜 Recent review details🧰 Additional context used📓 Path-based instructions (2)**Never hard-wrap prose.📄 CodeRabbit inference engine (AGENTS.md) Files:
In MDX, leave a blank line between a JSX tag and a code fence.📄 CodeRabbit inference engine (AGENTS.md) Files:
🧠 Learnings (4)📓 Common learnings📚 Learning: 2026-06-26T12:23:22.696ZApplied to files:
📚 Learning: 2026-05-23T01:23:59.268ZApplied to files:
📚 Learning: 2026-08-13T12:17:52.620ZApplied to files:
🪛 LanguageTooldocs/src/content/docs/architecture.md[style] ~79-~79: Since ownership is already implied, this phrasing may be redundant. (PRP_OWN) [style] ~92-~92: This phrase is redundant. Consider writing “last”. (LAST_OF_ALL) [style] ~92-~92: Since ownership is already implied, this phrasing may be redundant. (PRP_OWN) [style] ~186-~186: This word has been used in one of the immediately preceding sentences. Using a synonym could make your text more interesting to read, unless the repetition is intentional. (EN_REPEATEDWORDS_WHOLE) CHANGELOG.md[typographical] ~13-~13: Consider using an em dash in dialogues and enumerations. (DASH_RULE) [style] ~13-~13: This sentence is over 40 words long. Consider splitting it up, as shorter sentences make the text easier to read. (TOO_LONG_SENTENCE) [style] ~13-~13: This word has been used in one of the immediately preceding sentences. Using a synonym could make your text more interesting to read, unless the repetition is intentional. (EN_REPEATEDWORDS_WHOLE) [style] ~13-~13: This word has been used in one of the immediately preceding sentences. Using a synonym could make your text more interesting to read, unless the repetition is intentional. (EN_REPEATEDWORDS_WHOLE) [style] ~67-~67: Consider a more concise word here. (IN_ORDER_TO_PREMIUM) docs/src/content/docs/deployment.md[style] ~335-~335: Since ownership is already implied, this phrasing may be redundant. (PRP_OWN) [style] ~383-~383: This word has been used in one of the immediately preceding sentences. Using a synonym could make your text more interesting to read, unless the repetition is intentional. (EN_REPEATEDWORDS_WHOLE) [style] ~387-~387: Since ownership is already implied, this phrasing may be redundant. (PRP_OWN) [style] ~393-~393: Consider using a more formal/concise alternative here. (OTHER_THAN) docs/src/content/docs/api.md[style] ~803-~803: This word has been used in one of the immediately preceding sentences. Using a synonym could make your text more interesting to read, unless the repetition is intentional. (EN_REPEATEDWORDS_WHOLE) docs/src/content/docs/sdk/reference.md[style] ~59-~59: This word has been used in one of the immediately preceding sentences. Using a synonym could make your text more interesting to read, unless the repetition is intentional. (EN_REPEATEDWORDS_USUALLY) [style] ~59-~59: Consider an alternative for the overused word “exactly”. (EXACTLY_PRECISELY) 🔇 Additional comments (48)
📝 SummarySummary by CodeRabbit
WalkthroughThe change adds nested tenant settings directories with per-tenant validation, reloads, routing, and failure handling. It updates application wiring, admin APIs, the TypeScript SDK, authentication query handling, tests, CLI validation, and documentation. ChangesNested tenant settings
Priority: ⬇️ Low Estimated code review effort: 5 (Critical) | ~120 minutes Change: Feature Sequence Diagram(s)sequenceDiagram
participant Client
participant TenantMW
participant Registry
participant Handler
Client->>TenantMW: Send request with tenant identifier
TenantMW->>Registry: Resolve tenant store
Registry-->>TenantMW: Served, rejected, or unknown result
TenantMW->>Handler: Forward request with tenant store
Handler-->>Client: Return response
Merge Risk: ⚪ Minimal · up to No actionable current-head risk remains; the PR is ready to merge after normal checks. 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
Full details: Docstring CoverageExplanation Docstring coverage is 73.39% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 124 functions across 37 files. (12 skipped: 12 unsupported.)
✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
✨ Simplify code
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
📚 Docs preview is live → https://33a3f65b-wavehouse-docs.wave-rf.workers.dev |
Code Coverage OverviewLanguages: Go GoThe overall line coverage in commit b0eedbc in the Show a line coverage summary of the most impacted files.
Updated |
|
@coderabbitai review |
✅ Action performedReview finished.
|
…nt (#610) ## Summary The second slice of story 6 of the multi-tenant epic: each tenant reads its own ClickHouse. No behavior change for a settings directory that holds the four files beyond the three noted at the end. - One native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, a plain comparable value; `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`. `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own: the HTTP target is composed per tenant over the pool's TLS config, and the deadline is per call. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a changed largest ask resizes with the same grace. - The ceiling across pools: `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together. Boot refuses naming the sum and the ceiling. At a reload the walk keeps the ceiling at every step, tenants no longer served leaving first and what was refused placed once more at the end: a resize above it is refused and the pool keeps its size; a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or a pool the driver refuses to open (opened as the walk places it, so every one of these is undone in place) — leaves its tenants on the pool they had, `Params` and all (the keep-previous-wiring rule of #603), or on none when they had none. Both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. - One `discovery.SchemaRegistry` per served tenant over its pool (a `discovery.Source`: the pool's connection and the database that pool was opened for, read together once per refresh, so a move the ceiling refused keeps discovering the database the tenant's queries and inserts still use), each with a loop of its own: `RetryRefresh` until the first success, then `StartAutoRefresh` at the tenant's `schema.refresh_interval`, first tick at a random point within the interval so tenants adopted together do not refresh together. Created and stopped from `AfterAdopt`, all stopped under `App.Close` within the release budget. `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. The hub and the ingest worker resolve the registry and the HTTP target through tenant-keyed getters called with the tenant each message's topic names (#609), so the worker inserts each (tenant, table) batch into its own tenant's ClickHouse. - Probes. Flat mode unchanged: tenant 0's synchronous boot refresh drives `BootState`, `/readyz` pings the one pool. Nested: `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; nested boot never waits on a tenant's ClickHouse. `/readyz` pings every open pool at once and is ready at the first answer — concurrently, since the driver waits up to its 30 s dial timeout on a host that does not answer, past a kubelet probe's 1 s — and names every pool that did not answer when none does. `SchemaRegistry.Lookup` answers `ErrNotLoaded` before the first success, mapped at the three former 404 sites (and the schema list, where `[]` would read as no tables) to `503` with `Retry-After: 5`. - `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (#598): absent is tenant 0, `400` malformed, `404` unknown, `503` rejected. The SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. - `sharedTables` narrowed: the worker's invalidation fans out to the tenants on the same ClickHouse address and database (`Pools.SharingTables`), whatever their user or `tls` block, since they read the same tables. A tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step: `Cache.InvalidateTenant`, a tenant generation leading every version key (an enumeration of the index would miss a table no bump ever keyed); a pipe result names no table and keeps its TTL, as after any insert (#343). - `HTTPClients` keeps one client per TLS config: the proxy now serves tenants on different configs in alternation. - Passwords: `Identity` carries the field, but until #529 every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - Three changes reach the single-tenant directory: a table lookup before the first discovery is `503` with `Retry-After: 5` rather than `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema`; the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` orphans the structured-query results cached before it, where they were served until their TTL. ## Test plan - [x] `make ci` passes locally on the merge of main (#609) and on the commits before it - [x] `chconn`: different tuples get different pools and a shared tuple one, sized to the max; a tuple change repoints one tenant while the other keeps its `Manager`, resized with the grace close; a released tuple closes after its grace; boot refused over the ceiling naming sum and ceiling; a reload refuses a third tuple and the next reload opens it once a sharer shrinks; a refused resize keeps the size and is retried; a shared pool's refusal is reported once; a refused move keeps the previous pool and `Params`; a move to another address or database is reported stale, a username change is not; a move at the ceiling is allowed; an unreadable certificate refuses the tuple; a pool the driver refuses to open keeps the mover on its previous pool and `Params`; `SharingTables` by address and database; `Ping` first success, all failing (naming each), none open; `Close`; per-tenant `Target`; `HTTPClients` one client per config; `Identity` comparability - [x] `discovery`: `Lookup` before and after the first refresh; a nil connection getter; the connection read once per refresh; the first tick within the interval - [x] `api`: the two 503s on every route, the strict `?tenant=` table on the three ops routes, the ClickHouse getters receiving the request's store through the real router, the health `Ping` - [x] `app`, booting the real wiring on closed ports: two tenants on different tuples, then shared, then a username change repointing one; the ceiling at boot and at reload with the recovery reload; a rejected and a removed tenant releasing pool and registry; nested `/livez` degraded with the tenant diagnostic, then sticky `200`, `/readyz` naming every closed pool; `ErrNotLoaded` → `503` in both shapes; loops stopped by `Close`; the narrowed fan-out; a readmitted tenant's cache orphaned - [x] `cache`: the tenant generation in the keys, `BumpTenant`/`InvalidateTenant` - [x] `stream`/`ingest`: a registry source yielding nil reads as no schema and is asked for the tenant the lookup names; two tenants' interleaved batches each reach their own tenant's target; a batch with no target fails the insert naming its tenant - [x] Integration: the boot resilience test on the getters; a nested directory with two tenants on two databases of the one container — each discovers only its own tables, `/readyz` 200, structured queries and the raw-SQL proxy run against each tenant's own database - [x] SDK unit tests and an e2e case for the `tenant` option on the schema routes and `sql()` - [ ] Manual: a nested directory with two tenants on different ClickHouse users, reload one to a new address, watch the pools log and `/readyz` ## Follow-ups - A served tenant on no pool (refused by the ceiling) has its inserts fail into the existing failure path: parked on the DLQ, or left for redelivery when its `dlq.enabled` is off, until a reload gives it a pool. - Story 3 owns what a removed tenant's async paths do; a rejected tenant's registry is dropped and rebuilt on repair, so its first lookups after a repair are a sub-second `503`. - Per-tenant ClickHouse passwords wait for #529. - A tenant moved to another database keeps its previous schema until its next refresh (`schema.refresh_interval`, or `POST /v1/ops/schema/refresh?tenant=`); the `stale` list the pools reconcile returns is the natural trigger for an immediate refresh. - Over a nested directory, `/livez` keeps the last boot-time discovery error of a tenant removed before any tenant loaded, until another tenant loads. - The codegen CLI has neither an operator-key option nor a `--tenant` flag, so it cannot read a nested directory's schemas (documented as unsupported on the SDK reference page). ## Related Issues Part of #583 (story 6).
#611) ## Summary Story 3 of the multi-tenant epic: what removing or rejecting a tenant at runtime does. A settings directory that holds the four files sees no change beyond the dedupe store moving back to `<data_dir>/pebble`. - **A tenant's open streams end with it.** A `GET /v1/stream` used to outlive a tenant a reload removed or rejected: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served (`Hub.Prune`, through a close-once `Subscriber.Evict`), and the handler ends the stream. That covers a gap-fill in progress (`replayContext` now also watches the eviction) and a stream `TenantMW` admitted just before the reload but registered with the hub just after the prune: the handler checks right after registering (`StreamHandler.Served`), and since the registry swaps its map before its hooks run, one of the two always sees it. The reconnect gets `404` (removed) or `503` (rejected). The SDK needs no change: it reconnects after a server-ended stream, stops on a `404`, and retries a `503`, resuming from `Last-Event-ID` once the folder is back — in a browser going cross-origin, only where the refusal passes CORS (see Follow-ups). The stale comment in `sse.ts` that read a `404` as a missed route is fixed. - **The last tenant can be removed too.** An emptied nested directory used to read as a change of shape (`Validate`'s flat directory missing its files), which rejected the reload whole and left that tenant served. A registry started with tenant folders now reads it as no folder left, so the reload removes the tenant like any other. Boot still reads an empty directory as the four files, missing. Reloading a deleted folder by name still leaves its tenant rejected (`503`), as `registry_test.go` has pinned since #598. The flip side: a whole-directory reload that finds the directory emptied by accident drops every tenant until the folders are back and reloaded, as it already drops any one tenant whose folder is deleted. `wavehouse validate` still reads an empty directory as boot does, as the four files missing, so a server restarted before a folder is written back refuses to boot. - **`/v1/health` keeps resolving a tenant, deliberately.** Its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, which needs the tenant's verifier; exempting it would take its CORS answer from tenant `0`'s list, which a nested directory need not have. Recorded at the route and in `api.md`. - **A removed or rejected tenant's queued rows go to the dead-letter queue.** With no pool the insert fails and `dlqFor` reads the registry miss as on, so the rows are parked under `dlq.<tenant>.<table>`: never dropped, never inserted into another tenant's ClickHouse, and not left unacked, where they would hold the ack floor, stop the sweeper, and fill the one shared stream toward `mq.max_bytes_gb` until every tenant's ingest got `503`. A batch with no ClickHouse connection (its tenant no longer served, or no pool could be opened for it, such as by the connection ceiling) now also skips the row-by-row retry it could never pass and meets its dead-letter switch once, whole, logged once per batch rather than twice per row; a served tenant on no pool still honors its `dlq.enabled`. - **`/livez` drops a gone tenant's diagnostic.** Over a nested directory, a discovery error naming a tenant that stops being served before any tenant has loaded goes back to `no tenant has completed a first discovery yet` (a #610 follow-up). The schema side needed nothing new: story 6 already stops a gone tenant's refresh loop, now pinned by the removal test. - **Dedupe: one Pebble instance for every tenant** (its own commit). The wiring hands the embedded implementation `data_dir` once (`dedupe.NewEmbedded`) and asks it for each tenant's store (`Embedded.Tenant`, the `Factory`); the implementation keeps every tenant's seen ids in one instance at `<data_dir>/pebble`, each key `<tenant>\x00<id>`. Each tenant behaves as before: the instance is open only while some served tenant has dedupe on (a server with dedupe off opens nothing), a tenant switched off, rejected or removed keeps its seen ids, and the Pebble gauges are one set (`Embedded.Stats`; `Stats` leaves the per-tenant `Deduplicator`, `Managed` and `Stores`). Measured with Pebble v1.1.5, built like the release: an instance per tenant cost 13 goroutines, 7 open files and about 0.5 MB of heap each, so 1,000 tenants meant 13,000 goroutines, 7,000 files, about 500 MB of heap and 7.6 s of opens at boot, against 13, 7, about 4 MB and 14 ms for one instance; synced writes were 1.7x faster shared. There is nothing to migrate: `moveLegacyDedupeStore` and the old `<data_dir>/pebble` handling are deleted, and `tenant.Parse` no longer reserves `nats` and `pebble`, which it did only because a tenant's state lived at `<data_dir>/<tenant>`. One consequence of sharing: an instance that cannot open fails closed every tenant with dedupe on, not one. - **Sweep.** Every sentence that deferred to story 3 or that this change makes false is updated: the `perTenant` comment, the CHANGELOG entry on a removed-and-restored tenant's cache (settled by #610's `Cache.InvalidateTenant`), `deployment.md`'s "What a lost tenant `0` costs", and the per-tenant dedupe layout across the wiring, config, compose, Dockerfiles, `AGENTS.md`, the CHANGELOG entry for #602, and the architecture, configuration, deployment and settings-directory pages. ## Test plan - [x] `make ci` passes locally - [x] Pre-push reviewers: five rounds; the code review ended at `ship_it`, and the docs review's last round (three wording fixes) is applied without a further round - [x] `stream`: `Evict` is close-once; `Prune` evicts every subscriber of a tenant no longer served, on each topic and under each role, and no one else's - [x] `api`: a stream ends on its own when its tenant stops being served, whether idle, mid-gap-fill, or admitted before the reload and registered after the prune - [x] `app`, over a real listener: the stop ends every stream; a reload that rejects one tenant and then removes another ends each one's stream at once and cleanly while the other's stays open, and the reconnects get `503` and `404`; a reload a flat directory rejects leaves the stream open - [x] `app`, remove then restore: a rejected and then a removed tenant release pool, registry, refresh loop, verifier and dedupe store; the removed one's routes answer `404`, the worker gets no target and the switch on, and `/livez` goes back to the no-tenant line; restored, it is served over a fresh pool, registry and verifier (the JWKS is fetched again), and an id sent before the removal is still a duplicate - [x] `settings`: removing every folder removes every tenant and runs the hooks, and a folder written back is a tenant again; an emptied flat directory is still a rejected reload; an empty directory still refuses boot - [x] `ingest`: a batch with no ClickHouse connection meets its switch once, parked under its own topic and acked, or left unacked, with no request made, while the served tenant beside it inserts its own rows alone - [x] `dedupe`: tenants' ids never meet in the shared instance, however their keys would join; the instance opens with the first store and closes with the last, its ids kept; `Stats` is the instance's one set; an unopenable instance fails every tenant's store and the next apply retries - [ ] Manual: a nested directory with two tenants and dedupe on, a `wh.from(table).stream()` open on each; remove one folder and reload, watch its stream end and the SDK stop on the `404`; restore it and send an id it sent before ## Follow-ups - Over a nested directory serving no tenant `0`, a removed tenant's `404` carries no CORS headers (refusals read tenant `0`'s list, #600), so the browser SDK sees a network error and keeps re-dialing instead of stopping. Whatever turns away an unknown subdomain in front of WaveHouse should answer with CORS headers, the preflight included. If browser clients should stop on their own, one option is a final event on the evicted stream before it closes, which that connection's own CORS lets the browser read; that is a design question for Eric. - `sdk/streaming.md` and `sdk/reference.md` say the tenant-resolution `404` comes "when you send `X-Tenant-ID`"; over a nested directory with no `0` folder a client that sends no header gets `404 unknown tenant: 0` too, and removing a `0` folder now ends such a client's stream into it. The wording predates this branch. ## Related Issues Part of #583 (story 3).
## Summary Story 5b of the multi-tenant epic: every tenant gets a message queue of its own inside the embedded NATS server, and the `internal/mq` interfaces speak per tenant, never per stream, so nothing outside the package assumes that layout (an external implementation, where the Kubernetes operator owns stream config, can keep one shared stream). A settings directory that holds the four files sees no change beyond the stream names: tenant `0`'s own window and budget, as before. - **A stream pair per tenant.** `INGEST_<tenant>` (`ingest.<tenant>.>`, `DiscardNew`) at the tenant's own `mq.max_bytes_gb`, and `DLQ_<tenant>` (`dlq.<tenant>.>`, `DiscardOld`) at a tenth of it. The prefixes differ in their first letter, so no tenant id makes one kind's name the other's (pinned by `TestStreamNames_NeverCollide`). Subjects are unchanged. - **Opened with the first budget, kept on removal.** `SetMaxBytes(ctx, tenant, bytes)` opens a tenant's pair the first time (dead-letter stream first, so no row is queued that could not be parked) and resizes it after; the wiring hands every served tenant's budget at boot and after every reload, replacing the tenant-`0`-only `onDefaultAdopt` hook and its boot warning — and with it the tracking of tenant `0`'s last adopted store (`defaultSetting`), whose one reader left, a flat directory's ops gate, now reads tenant `0` through the registry (`defaultPolicy`). A publish or park that finds a stream missing reopens it at the budget last asked for, and so does a publish to a queue the broker has not recorded open: an open that timed out can leave a stream JetStream creates after all, which no consumer would hold. The publishes and parks that find a queue not open share one attempt, and after one fails the tenant's publishes and parks are refused at once for five seconds — so a queue that cannot open never holds the broker's lock for every retrying client, which every other tenant's reload would wait on — while a reload retries regardless; with none asked yet (the instant between a reload adopting a new tenant and its budget arriving) the publish is refused as `ErrQueueFull`, a `503`. A removed or rejected tenant keeps its pair at the last budget, and boot takes stock of the pairs on disk, so its queued rows are still delivered and parked on its own dead-letter queue. - **Isolation.** The ingest worker's and the hub bridge's durables are held on every tenant's stream, those opened later included; each tenant has its own ack floor, its own `MaxAckPending`, and a delivery goroutine of its own, so one tenant's backlog or stuck handler holds back no other. A publish that reopens a missing queue does it detached from the request's cancellation, and the consumers join on a budget of their own, so a client that goes away mid-open cannot leave a queue no consumer holds — which would fail the worker and stop the process. The worker's prefetch, and the hub bridge's fetch-ahead (the client default of 500), are each shared across the tenants' streams (at least one each) rather than multiplied by them. The durables are looked up before anything is written, so a boot over many queues writes nothing it need not (measured: 0.6 s at 1,000 tenants, against 15 s rewriting every stream and durable). - **Sweeper.** `PurgeAcked` takes each tenant's cutoff; the sweeper hands it every tenant's own `stream.gap_window_minutes` (`gapWindows` replaces `longestGapWindow`). A rejected tenant keeps the window its folder last had — a rejection is the common reload failure, and its clients resume from `Last-Event-ID` once the folder is fixed (`settings.Registry.Known` now yields each tenant's last adopted store for this), and one rejected since boot, whose window this process never read, keeps all of its history until its folder validates — while a removed tenant's stream keeps no acknowledged history. Per-tenant purge lines are Debug, with one Info summary per sweep. - **`GET /v1/ops/dlq/stats?tenant=`**, parsed strictly with `opsTenant`. Absent is tenant `0`, the ops-read convention, and for a flat directory exactly the previous answer; no parameter no longer sums every tenant. The tenant is looked up in the MQ, not the settings, so a rejected or removed tenant's parked rows are read by name; a tenant with no dead-letter queue is a `404`. The SDK's `wh.dlq.list()` and `.table()` take a `tenant` option, as the pipe reads did in #598. - **Dead-letter shrink guard (#532, interim).** A reload that shrinks a budget never caps a dead-letter stream below the bytes it holds: it keeps what it has, and a warning names the tenant and both figures. #532's ClickHouse-backed dead-letter table stays there. - **Disk accounting (decided: option A).** JetStream counts every stream's cap as reserved disk and refuses a stream once they pass 75% of the free disk at boot; per tenant that refused the 569th tenant at the 1 GB minimum on an 834 GB disk, and would refuse the 12th at the seed's 50 GB. The embedded server's limit is set out of reach, so a budget is a cap and never a reservation; what the budgets add up to against the disk is #138's. The one flat difference beyond the stream names: a budget above three quarters of the free disk now boots where it was refused. - **Old data directories.** `WAVEHOUSE` and `WAVEHOUSE_DLQ` claim `ingest.>` and `dlq.>`, which JetStream will not let overlap a new stream, so boot deletes them, logging what they held. What that made dead is gone: `parseTopicKey`'s one-token subjects and `deployment.md`'s "Upgrading across the tenant subject token". The v2-envelope upgrade runbook now says the old queue is deleted rather than carried over, so draining and replaying the parked rows has to happen before the upgrade, and the worker's comments and dead-letter hint no longer assume an undrained pre-v2 row can reach it. - **Sweep.** Every "story 5b" and "until 5b", and every sentence this makes false, is updated: the `mq` interface comments, the sweeper, `wire.go` and `app.go`, `app_test.go`, the hub subscriber's one-goroutine comment, the settings comments, and `deployment.md`, `settings-directory.mdx`, `api.md`, `architecture.md`, `ingest-pipeline.md` (diagrams included), `durability.md`, `configuration.mdx`, `why-wavehouse.md`, the SDK admin and reference pages, `AGENTS.md` and the CHANGELOG. ### Cost check Measured in a scratch program against the same nats-server 2.14.6 and server options (fsync on every write), with both durables consuming on each tenant's stream, at 10 / 100 / 1,000 tenants, against the one shared pair: - Resident memory, idle: 25 / 53 / 280–315 MB, against 21–35 MB shared — about 0.3 MB per tenant. Goroutines: 323 / 2,123 / 20,123, against 43. - Under load, a tenant written in the last ~10 s holds a block buffer: at 1,000 tenants one write to each took the heap from 172 to 439 MB, back to 172 MB within 15 s. One open file per stream written, closed within about two minutes of quiet. - Boot with data: 12 / 59 / 611 ms. A tenant's first creation costs about 20 ms (the pair, two durables, delivery), so a first boot of 1,000 new tenants pays about 20 s once. - Publish throughput: 32 concurrent publishers 1,565 / 7,155 / 2,390 msg/s against ~1,170 shared; one publisher round-robin across tenants 999 / 560 / 503 against ~1,100. - One sweep: 31 ms / 296 ms / 2.75 s, once a minute. - Not measured, but by construction: while ClickHouse stalls, the worker can hold up to `maxAckPending` (10,000) unacknowledged rows per tenant with traffic — about 10 million at 1,000 tenants — where the shared stream held 10,000 for the whole process (see Follow-ups). ### An upstream quirk When a stream's store fails to open, nats-server 2.14.6 releases a disk reservation it never made, so its reserved count goes negative. Harmless under the default limit, but with the limit at `math.MaxInt64` the subtraction overflows and every later stream is refused until restart. The limit is half the int64 range instead, and `TestEmbeddedNATS_SetMaxBytes_AQueueThatCannotOpen` pins it: one tenant's failed open, then the next tenant's queue opens. ## Test plan - [x] `make ci` passes locally - [x] Pre-push reviewers (round 1: both iterate — fixed: the reopen detach, the v2 upgrade runbook, the remaining sweeper and `create stream` lines, the `503` rows, the disk-full sizing rule, the per-tenant consume callbacks in the goroutine topology; round 2: both iterate — fixed: a test for the boot rule of a queue that cannot open, the boot pass honoring `New`'s context, the last pre-v2 mentions, the resume hole after a rejected folder, the runbook's replay step the old build cannot carry out, the boot warning's count; round 3: code iterate on one SHOULD — boot now re-applies a tenant's budget to a pair it finds split, or missing its dead-letter stream, as every boot rewrote both streams before; docs iterate — the boot/reload contexts in `architecture.md`, the runbook's opening, the per-tenant consumer line, and a full disk needing a restart; round 4: code iterate — a failed resize's undo now restores the ingest stream's actual cap (after a boot that found a pair split it would have restored 0, which JetStream reads as no cap), and a test pins the report of a queue a consumer cannot join; docs iterate — both open-timeout errors and when a later boot writes, three garden-path sentences; round 5: code iterate — a rejected tenant keeps its replay history at its folder's last window instead of losing it at the next sweep, and all of it while rejected since boot; docs iterate — disk sizing counts every tenant ever served, since removed tenants' queues are kept; round 6: code iterate — the hub bridge's fetch-ahead is shared across the tenants like the worker's prefetch, instead of 500 per tenant; docs iterate — a stale "stop the sweeper" sentence from the shared stream, and a dead-letter stream is opened, not guaranteed to exist, when its tenant is first served; round 7: code iterate — a publish goes by the broker's record of an open queue rather than a stream answering, so a stream an open gave up on but JetStream created anyway is joined before rows land in it, and the queue-full and publish-failure log lines name the tenant; docs iterate — the disk-sizing guidance gets its own paragraph, and the leftover tenant-`0` tracking clause goes, with the tracking itself; round 8: code ship it; docs iterate — what a consumer that cannot join a queue opened at runtime does (the worker exits the process, the hub logs and the tenant's streams get no live rows until a restart), and a failed purge lookup ends that tenant's purge, not the sweep; round 9: code iterate — a tenant whose queue cannot open no longer holds the broker's lock for every retrying publish (publishes share one attempt, and after one fails are refused at once for five seconds; a reload retries regardless), and the worker's per-tenant memory ceiling is recorded here, for #583's deferred worker rework; docs iterate — two wording fixes; round 10: code iterate — tests for one tenant's failed purge stopping no other and for a durable reused across a restart; docs iterate — the SDK replay note names the upgrade that deletes the old queue, and each tenant's own gap window; round 11: code ship it; docs iterate — disk sizing counts a dead-letter stream the shrink guard kept above a tenth, and the stats route says where the tenant is looked up; round 12: docs ship it — the round-12 commit is docs prose only, so the code reviewer's round-11 ship it stands for it, recorded as a logged skip) - [x] CodeRabbit round 1 (changes requested, two findings, fixed): the sweeper no longer logs other tenants' purge failures at `WARN` beside one tenant's missing consumer, and a park that finds the dead-letter stream missing is paced like a publish - [x] CodeRabbit round 2: approved - [x] Flake fix after Eric's heads-up from the stacked #613 builds (`f5d8f484`): `TestEmbeddedNATS_PacesTheRetriesOfAQueueThatCannotOpen` re-created its obstacle while nats-server's post-failure goroutine was removing the emptied `streams` and account directories (EINVAL on APFS, ENOENT on Linux; 3 of 25 runs here); the test and the one app test with the pattern now keep another tenant's streams in the directory. CodeRabbit round 3: approved. `eb9f7409` closes the same goroutine's much narrower window in `TestEmbeddedNATS_SetMaxBytes_AQueueThatCannotOpen` with a non-stream directory under `streams/` (reserves nothing, so the test still pins the reservation overflow); Eric's second measurement (9 of 10 on `feat/mq-nats-wiring`) was on a branch forked before `f5d8f484` — verified by applying the fix there (5 of 20 → 0 of 20). CodeRabbit round 4: approved - [x] `mq`: a reopen outlives the caller's cancellation and the consumer still joins (fails without the detach); a tenant's pair opens with its first budget, names and caps as above, and no other tenant's; a publish reopens a missing pair at the last budget and is refused with none asked; a publish opens a queue whose open gave up on a stream JetStream created anyway, and its row reaches a running consumer (fails without the gate); after a failed attempt a publish or park is refused without waiting on the broker's lock, and a publish tries again once the window has passed (fails without the pacing); one tenant at its budget is refused while another publishes; a tenant at `MaxAckPending` and one with a stuck handler hold back no other, and a tenant's messages arrive in order; both consumer paths join a queue opened after they started; one tenant's deleted durable reports on `failed`, and a stop is not a failure; the prefetch share, and the hub bridge's shared fetch-ahead; per-tenant resize, rollback and cancellation; the shrink guard keeps every parked row; per-tenant purge at each cutoff, and no history for a tenant not named; one tenant's failed purge — its durable or its stream gone — stops no later tenant's, and an ended sweep touches none (fails with either `continue` turned into a `break`); a durable on disk is reused across a restart, or updated in place when its settings differ, and delivery resumes past what it acknowledged; per-tenant dead-letter counts, `ErrNoDeadLetterQueue`, and a missing dead-letter stream reopened; a queue that cannot open costs that tenant alone; boot deletes the old shared pair; boot over existing pairs reads their budgets back and delivers a queued row of a tenant given no budget, and re-applies the budget to a split pair or one missing its dead-letter stream while leaving a guarded one alone; a failed resize's undo restores the ingest stream's own cap; a consumer that cannot join a queue opened later reports it on `failed` - [x] `ingest`: the sweeper hands each tenant's own cutoff, re-read every sweep, and logs a failed sweep at `WARN` only when every tenant's failure is a missing buffer consumer - [x] `settings`: `Known` yields a rejected tenant's last adopted store, and none for a folder that has not validated since boot - [x] `api`: `dlq/stats` without the parameter reads tenant `0`, names a tenant's queue alone, `404`s without a queue, and refuses a malformed, empty, repeated or misparsed `tenant` - [x] `app`: a queue that cannot open refuses a flat boot and costs a nested directory that tenant alone, whose next publish opens it; a cancelled boot opens no queue; each tenant's budget follows its own folder, and a rejected or removed folder keeps its queue at the last budget; `gapWindows` names each served tenant, a rejected one at its folder's last window (unbounded when rejected since boot), and no removed one - [x] SDK: `tenant` on `wh.dlq.list()` and `.table()` - [ ] Manual: a nested directory with two tenants; fill one tenant's queue and see only its ingest answer `503`; remove it and reload, then read its parked rows with `GET /v1/ops/dlq/stats?tenant=` ## Follow-ups - The nats-server reservation quirk above is worth reporting upstream. - A consumer that cannot join a queue opened at runtime fails the ingest worker, so the process restarts. Now that a publish goes only into a queue recorded open, `apply` could leave such a queue unrecorded instead, so the tenant answers `503` and publishes and reloads retry the join, with no restart. - The worker's in-memory bound is per tenant now: while ClickHouse stalls it can hold up to `maxAckPending` (10,000) rows for every tenant with traffic, documented at the constant and in the ingest pipeline's backpressure list. A process-wide bound would bring back one tenant holding the others back; it belongs with #583's deferred worker rework. - nats-server logs three Info lines per stream at every boot (restore and consumer recovery), about 6,000 at 1,000 tenants, through `slogNATSLogger`. - `sdk/admin.md`'s older DLQ example declares `const { data }` twice in one block, and shows a `"users": 0` count the server never reports (both predate this branch). - `api.md`'s `?table=` row says it returns only that table's count; `total` stays the tenant's whole count (predates this branch). - `settings-directory.mdx`'s DLQ switch still calls the unreadable-envelope exception "new in this release", which goes stale with the next one (predates this branch). ## Related Issues Part of #583 (story 5b).
#711) ## Summary A `settings.dir` holding folders but none of the four files is read as nested. When every folder was rejected, or none was named by a tenant id (a fresh volume whose only entry is `lost+found`, a parent directory mounted by mistake), `settings.Open` still returned a registry and the server came up serving no tenant, answering every tenant route `unknown tenant`. Before #598 that root refused boot. Boot now refuses a nested root that would serve no tenant, with the findings plus one naming the rule, through the same path a flat invalid root takes, so the error still points at `wavehouse validate` and `wavehouse bootstrap`. The rule is boot's alone: a whole-directory reload that finds every folder gone or broken still drops every tenant and keeps running (#611), and `wavehouse validate` is unchanged. A root with one valid folder beside a rejected or stray one still boots and serves that tenant, and an empty or missing root still refuses as the four files, missing. ## Test plan - [x] `settings.Open` refuses a root whose every folder is rejected, and a root whose only folder is `lost+found`, each with the findings - [x] `settings.Open` on one valid folder beside a broken one and a stray one opens and serves the valid tenant - [x] Existing pins unchanged: an emptied directory on reload removes every tenant and keeps running; empty and missing roots refuse as "config.json: missing"; `wavehouse validate` exit codes - [x] `make ci` passes locally, integration and e2e included ## Related Issues Closes #599 Part of #583 <!-- Checklist for the author (not kept in the squash commit message): - `make ci` passes locally - Docs updated per AGENTS.md "Documentation & Consistency Sync" rules - CHANGELOG.md [Unreleased] entry added - Tests cover new / changed behavior (70 % minimum, 80 %+ preferred) The PR title is the squash commit subject — use Conventional Commits (`feat:`, `fix:`, `docs:`, `refactor:`, `test:`, `chore:`, `ci:`, `deps:`, `build:`, `perf:`, `revert:`, `style:`). The PR body below is the squash commit message, so keep it tight. -->
Summary
Story 2 of the multi-tenant epic: a settings directory can hold one folder per tenant, and the registry owns every reload. No behavior change for a directory that holds the four files — the same findings byte for byte, the same boot refusal, the same keep-previous on a rejected reload, the same watcher and
SIGHUP, the samewavehouse validateexit codes.settings.Validatereads the shape off the root: any of the four file names makes it flat (ValidateDir, the oldValidate, untouched); otherwise a folder makes it nested — folder names through the tenant-id grammar, each folder throughValidateDir, findings led by the folder (acme/policies.json). The shapes never mix. A folder whose name is not a tenant id is reported and skipped; an entry that cannot be stat'ed (a dangling symlink) is a finding about the root.settings.Openreturns theRegistry, which ownsReload,ReloadTenant, theAfterAdopthooks and the watcher;Storeis a passive holder. A nested directory fails closed per tenant, at boot and on reload alike: a rejected folder stops its tenant (a bare503on tenant routes, which resolve before authentication), the rest carry on, and there is no previous-snapshot fallback. A whole-tree reload mirrors the folders and names a removed tenant in the log. A finding about the directory itself refuses boot and rejects a reload whole. A nested directory gets no watcher (Registry.Watchrefuses one);SIGHUPreloads the whole tree in both shapes.POST /v1/ops/settings/reloadandGET /v1/ops/pipes[/{name}]take an optional?tenant=, parsed strictly (400for a query that does not parse or an empty, repeated or malformed id;404unknown; absent means the whole tree on the reload and tenant0on the reads). It lands with theauth.bearerTokenfix deferred from feat(tenant): resolve X-Tenant-ID before auth and thread the tenant #593 — the?token=strip no longer repairs a query that does not parse — pinned throughNewRouterwith the real authenticator./v1/ops/*gate admits the operator key alone and a token admin gets403;NewRouterdecides that from the registry's shape, whatever policy source was wired. Boot warns when a nested directory has no operator key, sinceSIGHUPis then the only reload.0last adopted, and their hooks run only when tenant0is adopted, so a rejected or removed0folder leaves them as they were. The keepalive wheel runs at the shortestkeepalive_intervalamong the tenants being served (feat(stream): honor each tenant's keepalive settings on the wheel #597).wh.pipes.list(),wh.pipes.get()andwh.settings.reload()take atenantoption (OpsRequestOptions), sent as?tenant=.503is generic; the reload body is unchanged and a422can mean adopted in part; a failure about the directory itself rejects the reload whole; one shared keepalive wheel, with feat(stream): honor each tenant's keepalive settings on the wheel #597 tracking per-tenant keepalive; a badly named folder is skipped; a whole-tree reload mirrors the folders;Dependencies.PolicySourcestays and is ignored when the registry is nested.Test plan
make cipasses locally (unit, integration, e2e, coverage gates)ValidateequalsValidateDiracross six root shapes, findings and document alike*Store; folder mirroring; a root-level failure changing nothing (four shapes); a shape change refused in both directions; symlinked and dangling tenant folders; an unreadable folderTestNewRouter_HandlersReceiveTheRequestTenantsStore:assert.Sameon the store each handler's getters received, two tenants alternating through one router (fails under two wrong-store mutations)?tenant=table on all three routes;TestNewRouter_MalformedTenantSurvivesTheTokenStrip(fails against the pre-fixbearerToken)internal/app: a nested boot without a0folder; the operator reloading one tenant by name; the process-wide resources following tenant0through another tenant's reload, a rejected0folder and a removed one (dedupe, MQ budget, CORS); the no-operator-key boot warningtenant: "0", unknown404, malformed and empty400)acme/+globex/,curl -H 'X-Tenant-ID: acme' /v1/pipes/x→404 pipe not found; no header →404 unknown tenant: 0Follow-ups
/v1/streamdelivery is by table alone and is authorized by tenant0's policy, and the most permissive tenant'sdefault_roleis the effective floor.0boots unconfigured: no ClickHouse address,/livezdegraded, the MQ byte budget uncapped (inert, since nothing publishes). Until story 6.0(its folder rejected or removed by a reload): the five settings read throughdefaultSettingkeep their last adopted values, but the async paths read tenant0as unset — the hub reads no policy (every stream subscriber's rows withheld), the sweeper's gap window is zero (gap-fill history purged), the DLQ switch reads on — andperTenantlogs an error per event. Story 3 owns the async paths' miss semantics; left exactly as story 1 left them.GET /v1/ops/schema,POST /v1/ops/schema/refresh,GET /v1/ops/dlq/statsandPOST /v1/ops/queryignore?tenant=(stories 5 and 6).0's admin role name (story 9). The dedupe store keys on the event id alone (story 7). The pre-auth400/404/503split distinguishes tenant ids,/v1/healthincluded (story 3).settings.dirpointed at a directory of unrelated subfolders boots as nested with nothing served;wavehouse validateexits 1 for a nested directory with one bad folder where boot would still serve the rest.0folder must still carry every required key, andconfig.yaml's flat-onlysettings:comment.Related Issues
Part of #583 (story 2). Lands the
?tenant=andbearerTokenwork deferred from #593.