Skip to content

epic(settings): multi-tenant settings directory and tenant-scoped runtime #583

Description

@taitelee

Area: settings / api / streaming — multi-tenant preparation epic · decided with @EricAndrechek, 2026-09-10 → 2026-09-15

WaveHouse runs one settings snapshot per process. WaveHouse Cloud needs one process to serve many tenants, each with its own roles, policies, pipes, and config, delivered by the cloud fan-out as files on disk exactly like standalone. This epic makes the settings directory, the in-memory snapshots, and every request path tenant-aware without changing anything for a standalone operator. Split into the stories below as they are promoted to Ready (the #194 pattern).

Decisions

  • Two directory shapes, no hybrid. Flat: the four files (config.json, policies.json, pipes.json, roles.json) sit in the root and are tenant 0. Nested: the root holds only subfolders, each folder name is a tenant id, each folder holds the four files. Loose files and subfolders in the same root is an error at boot and at reload. Switching shapes is stop, restructure, start.
  • Tenant id is a validated string. One rule: safe as a folder name and safe as a NATS subject token (letters, digits, _, -; no spaces, dots, slashes, or wildcards; capped length). The folder name is the id. 0 is the reserved default. Not an integer: 19-digit ids already round when treated as numbers (internal/auth/auth.go).
  • Requests resolve to a tenant before auth. The id comes from a request header (name is a constant, TBD); absent means 0. In flat mode tenant 0 always exists. In nested mode a 0 folder is allowed but advised against; the control plane won't create one. Putting the header on every request is the client's or proxy's job.
  • Settings owns the tenant map. The registry maps tenant id → *settings.Store. The middleware resolves the header against it and 404s on a miss; what goes into the request context is the resolved store (the long-lived handle whose getters read the latest adopted document), not the bare id. Handlers extract it once and pass it down as an explicit parameter; nothing below a handler reads context. Async paths (ingest worker, sweeper, DLQ, stream hub) carry the tenant in the MQ subject.
  • Invalid tenant folder: fail closed per tenant, per pod. At boot and on reload alike, a folder that fails Validate drops that tenant on that pod (its requests 503 with the findings); every other tenant on that pod, and every other pod, keeps serving. No previous-snapshot fallback in nested mode. The control plane runs the same settings.Validate (pure, no deps) before fan-out, so an invalid folder on disk is a bug, not a normal event. Flat mode keeps today's behavior: refuse to boot, keep the previous snapshot on a rejected reload. No status route for now.
  • Reload triggers. Flat mode keeps fsnotify + SIGHUP + POST /v1/ops/settings/reload. Nested mode does not use fsnotify: the control plane writes the folder and calls the reload route. SIGHUP works in both modes and always reloads the whole tree. The route takes an optional tenant id: absent reloads the whole tree, present looks the tenant up and reloads that folder. Ops routes stay tenant-exempt (/livez, /readyz, metrics, version, /v1/ops/*).
  • Ops routes take the operator key only in nested mode. Today the /v1/ops/* gate admits the operator key or a token whose role is the policy's admin role. In nested mode there is no single policy to read an admin role from, and no tenant's admin may act on another tenant, so the gate admits the operator key alone and a token admin gets 403: tenant admins never touch the ops routes. The operator key is boot-level and reaches every ops route for every tenant, whether the call names one (?tenant=) or none (the whole tree). Flat mode is unchanged — the gate still admits tenant 0's admin role — so nothing changes for a standalone operator. Decided for now (2026-09-21); a per-tenant admin surface would be a later story.
  • Scope stays as is. scope is a within-table partition (an org inside a shared table); tenant selects the settings and the tables. Different axes. Scope remains inert and tenant goes in as a separate leading token: cache namespaces tenant.table.version.scope, stream topic tenant.table[.scope], ingest subjects ingest.<tenant>.<table> / dlq.<tenant>.<table>.
  • Every folder carries a full config.json. ClickHouse, auth, dedupe, DLQ, query, schema, stream, MQ, CORS are all per tenant, so the runtime resources behind them become per tenant too (stories 6–9).
  • Cache is one global pool. One ristretto instance, tenant-prefixed keys. Heavier tenants hold more of it. Redis is the next layer, later.
  • Settled with Eric on 2026-09-21, landed in feat(settings): nested settings directory with a per-tenant registry #598. A rejected tenant's 503 is generic. The reload response body is unchanged and a 422 can mean adopted in part; a failure about the directory itself rejects the reload whole. One shared keepalive wheel at the shortest interval among served tenants (feat(stream): honor each tenant's keepalive settings on the wheel #597 tracks per-tenant keepalive). A badly named folder is skipped, never fatal. A whole-tree reload mirrors the folders (new adopted, missing dropped). Dependencies.PolicySource stays and is ignored when the registry is nested. /v1/ops/* admits the operator key alone in nested mode; boot warns when a nested directory has no operator key.
  • Landed in feat(mq): lead every subject with the tenant #609 (story 5a), decided 2026-09-23; Eric's to revisit. The tenant token goes in verbatim, since the id grammar already makes it one subject token. A subject without one, written before the upgrade, reads as tenant 0, so an upgrade needs no drain and rows parked before it count as tenant 0's. Until 5b, the sweeper's longest gap window among the tenants served is the keepalive wheel's rule: the one value that satisfies every tenant. Say acme keeps 15 minutes and globex 60: acme's history stays 60 minutes too, counted against the one shared mq.max_bytes_gb.
  • Landed in feat(clickhouse): tls, headers, pool sizes and a connection ceiling #603 and feat(clickhouse): one pool per tuple and one schema registry per tenant #610 (story 6), decided 2026-09-21 → 2026-09-23; Eric's to revisit. One native ClickHouse pool per distinct address, database, user, password and tls tuple among the tenants served, shared by the tenants naming it and sized to their largest ask. clickhouse.max_total_conns caps the open pools together: boot refuses above it, and a reload that would exceed it is refused and logged, the tenant keeping its previous pool. A tenant on no pool answers 503 with Retry-After: 30, and a table lookup before its tenant's first discovery answers 503 with Retry-After: 5 where it was a 404. Over a nested directory /livez is 503 until some tenant completes a first discovery and 200 for the rest of the process lifetime, and /readyz is ready when any open pool answers. An insert invalidates cached results for every tenant on the same address and database, and a tenant readmitted after a rejection or removal, or moved to another address or database, has its cached structured-query results dropped at once. Per-tenant ClickHouse passwords wait for auth(config): where boot secrets live and how they rotate (jwt_secret, operator_key, clickhouse.password) #529.
  • Landed in feat(app): end a removed tenant's streams, one Pebble for every tenant #611 (story 3), decided with Eric 2026-09-24. A tenant that stops being served, removed or rejected, has its open streams ended; the reconnect gets 404 or 503, and the SDK stops on the first and retries the second, resuming from Last-Event-ID. /v1/health keeps resolving a tenant: every tenant route answers an unknown tenant before authentication anyway, and a tenant's existence is not a secret (Cloud's proxy is meant to turn away unknown subdomains). A removed or rejected tenant's queued rows go to the dead-letter queue rather than staying unacked, where they would hold the shared stream's ack floor, and Eric expects that to hold with per-tenant streams too. Dedupe is one Pebble instance for every tenant, keys led by the tenant, the layout decided inside the embedded implementation.
  • Landed in feat(mq): give every tenant a queue of its own #612 (story 5b), talked through with Eric 2026-09-24. Every tenant has a stream pair of its own inside the embedded implementation: INGEST_<tenant> at the tenant's own mq.max_bytes_gb and DLQ_<tenant> at a tenth of it, opened with the tenant's first budget and kept on removal, so a removed or rejected tenant's queued rows are still delivered and parked on its own dead-letter queue. The sweeper purges each tenant at its own gap window: a rejected tenant keeps the window its folder last had (all of its history when rejected since boot), and a removed tenant keeps no acknowledged history. GET /v1/ops/dlq/stats reads one tenant, tenant 0 when ?tenant= is absent, looked up in the MQ rather than the settings. A budget is a cap and never a disk reservation; what the budgets add up to against the disk is mq: validate WH_MQ_MAX_BYTES_GB upper bound to prevent disk over-reservation #138's. A reload that shrinks a budget never caps a dead-letter stream below what it holds (mq(dlq): shrinking mq.max_bytes_gb silently deletes the oldest dead letters — investigate how the reload should treat a non-empty DLQ #532, interim). Boot deletes the old shared WAVEHOUSE and WAVEHOUSE_DLQ streams. One tenant's failed queue join costs that tenant alone (fix(mq): keep one tenant's failed queue join from ending ingest for all #680).
  • Deferred by decision. Boot secrets (auth(config): where boot secrets live and how they rotate (jwt_secret, operator_key, clickhouse.password) #529). Dedupe remote backend (Scylla/Dynamo), dedupe admin routes, purge and rename of tenants, the per-tenant status route, a ClickHouse socket ceiling beyond a process-wide cap, and a scheduler or worker model for the background loops — discovery, JWKS refresh, the sweeper, batch inserts (per-tenant loops with jitter instead for now; Eric, 2026-09-24, wants them reworked together later; a library like gocron v2 if a cap is ever needed).

Stories

  • 0. Prerequisite: refactor(cmd): extract app.App to dedupe main.go wiring + setup_test.go::buildServer #140, extract app.App. Done in refactor(app): extract the process wiring from main.go into internal/app #585. Pull the ~550 lines of hand wiring and reload hooks out of main.go into a package with consistent init/cleanup per component under an errgroup, so the registry lands into clean wiring. First.
  • 1. Tenant id extracted from the header and threaded everywhere. Done in feat(tenant): resolve X-Tenant-ID before auth and thread the tenant #593. New internal/tenant package (ID type, grammar check, reserved Default = "0", header name constant, FromContext/WithStore helpers; no imports from the rest of the repo). A router middleware reads the header, validates it, defaults to 0, resolves it against the settings registry, 404s on a miss, stores the resolved *settings.Store in context, and skips the exempt routes. Runs before auth. Every handler extracts once and passes it as an explicit argument; the func() T / func(table) T getter fields in internal/api, internal/ingest, internal/discovery and the async workers gain the parameter, wired with the default tenant in internal/app (wire.go). No behavior change: the registry holds one store keyed 0. The general-notes refactors (slog default instead of a logger parameter, dependency structs for constructors with more than a few params, context on everything with context-aware logging) ride along if the diff stays reviewable; otherwise they split into their own PR first. This PR also owes the slog-default cleanup deferred from refactor(mq): seal the MQ boundary behind an intent-level broker API #586 (drop the *slog.Logger param from mq.NewEmbedded, the handlers, NewAuthenticator, NewSchemaRegistry, NewSweeper, and the worker).
  • 2. Boot and reload detect the directory shape; registry. Done in feat(settings): nested settings directory with a per-tenant registry #598. settings.Validate detects flat vs nested from the root; hybrid is an error; each folder gets the existing per-directory checks; folder names are checked against the grammar; findings carry the tenant folder in their file path. settings.Open returns a registry keyed by tenant with For(id) (*Store, bool); stores are passive holders (document pointer + swap), the registry owns reload. Per-tenant fail-closed on a rejected folder. Reload route gains the optional tenant parameter. Flat mode keeps the fsnotify watcher; nested mode starts none. wavehouse validate [dir] and its exit codes unchanged. The admin pipe reads (GET /v1/ops/pipes[/{name}]) gain the same optional ?tenant=, pulled out of feat(tenant): resolve X-Tenant-ID before auth and thread the tenant #593: auth.bearerToken must first stop rewriting a query string that does not parse (the rewrite erases the malformed pair a strict ?tenant= parse exists to refuse, so ?tenant=acme;x=1&token=… reads the default tenant), pinned by a regression test through NewRouter; the SDK's wh.pipes.list() / .get() gain an option to send it (options.headers cannot set a query parameter). RequireAdmin admits the operator key alone in nested mode (Decisions) in the same PR — the parameter must never ship behind a gate that reads tenant 0's policy, or tenant 0's admins read every tenant's pipes. Once a two-tenant registry can be built, one test drives NewRouter + TenantMW with two tenants and asserts (assert.Same) the store a recording PolicySource was handed — feat(tenant): resolve X-Tenant-ID before auth and thread the tenant #593's doubles discard it.
  • 3. Tenants added and removed at runtime. Done in feat(app): end a removed tenant's streams, one Pebble for every tenant #611. In nested mode a whole-tree reload already mirrors the folders (feat(settings): nested settings directory with a per-tenant registry #598): a new folder is adopted and a removed one dropped. What remains is the drop's consequences (requests 404, open streams close; refactor(app): extract the process wiring from main.go into internal/app #585's StreamHandler.Closing channel does this for process shutdown, tenant removal needs the per-tenant equivalent). Removing never touches data: the tenant's JetStream stream, DLQ, and dedupe store stay on disk, and restoring the folder restores the tenant. Purge is a later ops route. Rename is out of scope: it is a remove plus an add, the old id's data stays stale and unpurged and is not carried to the new id. Defines what the async paths (worker, hub, schema registry) do for a tenant the registry no longer holds: since feat(tenant): resolve X-Tenant-ID before auth and thread the tenant #593 a miss is logged and read as the getter's zero value (nil policy), with two fail-safes: the worker's DLQ switch reads as on (park, never drop) and the schema refresh keeps its cadence. Since feat(mq): lead every subject with the tenant #609 the worker and hub look up each event's own tenant, so any event still queued for a removed or rejected tenant reaches the miss. The sweeper looks no tenant up (it folds over the tenants served; see Decisions), so it has no miss to define: a removed tenant's acknowledged history ages out at the longest window among those still served, all of it at the next sweep when none is left. Also decides /v1/health: it resolves a tenant before auth, so its 200 vs 404 is an unauthenticated tenant-existence check once ids are customer-derived — exempt it from TenantMW or accept it, deliberately. Also moves dedupe from story 7's one Pebble per tenant to one shared instance with the tenant leading every key, the layout decided inside the Pebble implementation of internal/dedupe rather than by the wiring (Eric, 2026-09-24): measured idle, each instance cost 13 goroutines, 7 open files and about 0.5 MB of heap, so 1,000 tenants meant 13,000 goroutines and 7.6 s of opens at boot. A tenant switched off, rejected or removed keeps its seen ids, and there is nothing to migrate.
  • 4. Seal the MQ boundary. Done in refactor(mq): seal the MQ boundary behind an intent-level broker API #586, which went further than wrapping types: internal/mq is now an intent-level broker API and the only package that imports NATS/JetStream. Addressing is mq.Topic{Table, Scope}; subjects, stream names, and the token encoder are private in internal/mq/subject.go; internal/app holds an mq.Broker, not the concrete embedded type; the sweeper's purge arithmetic lives behind Purger.PurgeAcked; DLQ behind DeadLetterer/DeadLetterStats; a full queue is mq.ErrQueueFull.
  • 5a. MQ tenant addressing and a tenant-keyed stream hub. Done in feat(mq): lead every subject with the tenant #609. After refactor(mq): seal the MQ boundary behind an intent-level broker API #586 the tenant token is one change inside internal/mq: a Tenant field on mq.Topic, encoded in subject.go as ingest.<tenant>.<table> / dlq.<tenant>.<table> and stream topic <tenant>.<table>[.<scope>] (tenant first so one wildcard selects a tenant's traffic). Callers set Topic.Tenant from the tenant they already hold (story 1); the ingest worker, DLQ, and stream replay read it back from Message.Topic(). No caller builds a subject. Stream hub subscriptions and per-role frame projection keyed by tenant live here. Every tenant still shares one ingest stream and one DLQ stream, so until 5b the sweeper purges at the longest gap window among the tenants served and GET /v1/ops/dlq/stats sums a table across tenants.
  • 5b. Per-tenant JetStream streams and DLQ. Done in feat(mq): give every tenant a queue of its own #612. Talked through with Eric on 2026-09-24, so no research spike gates it: stream counts, sharding and distribution are an external implementation's questions, and the embedded one gets a cost check in its own PR. Per-tenant JetStream streams (own byte cap and discard policy, which is what isolates a noisy tenant's disk; sum-of-budgets vs disk is mq: validate WH_MQ_MAX_BYTES_GB upper bound to prevent disk over-reservation #138), living inside the embedded implementation, so the sweeper purges each tenant's stream at that tenant's own gap window. The DLQ goes per tenant with them: the DLQ shrink guard (mq(dlq): shrinking mq.max_bytes_gb silently deletes the oldest dead letters — investigate how the reload should treat a non-empty DLQ #532), and story 2's optional ?tenant= on GET /v1/ops/dlq/stats. Eric (2026-09-24): a stream per tenant is fine for the embedded implementation, the only one built. An external NATS implementation will want one shared stream, perhaps sharded deterministically by subject, since in a distributed deployment the Kubernetes operator owns stream config and policy. So the per-tenant streams live only inside the embedded implementation, and nothing outside internal/mq assumes a layout: the sweeper's one cutoff and the one byte budget read from tenant 0 become per-tenant calls. A removed or rejected tenant's queued rows still land on the dead-letter queue. Reworking the sweeper as a background worker stays deferred (see Decisions).
  • 6. Per-tenant ClickHouse connection and schema registry. Done in feat(clickhouse): tls, headers, pool sizes and a connection ceiling #603 (TLS and connection keys) and feat(clickhouse): one pool per tuple and one schema registry per tenant #610. One chconn manager per distinct address+database+user+password+tls tuple (CH authenticates per connection, so different creds never share a pool), created and closed from the registry's AfterAdopt. One instance of the existing discovery refresh loop per tenant, started with a small random offset so tenants don't fire together; no scheduler and no worker pool (Eric, 2026-09-22: premature; if a concurrency cap is ever needed, adopt a library such as gocron v2 rather than build one — a follow-up, not here). /readyz semantics when some tenants' ClickHouse is unreachable. The tls block under clickhouse in config.json (certs, verification, headers) passed straight to the driver and the HTTP client, plus max_idle_conns / max_open_conns / conn_max_lifetime, shipped first as their own tenant-agnostic PR (feat(clickhouse): tls, headers, pool sizes and a connection ceiling #603). A process-wide socket ceiling and nothing beyond it. feat(dedupe): one store per tenant under data_dir/<tenant>/dedupe #602 added a wiring wrapper that fans a table invalidation out to every known tenant, because every tenant still queries the same ClickHouse tables; once tenants have their own tuples, narrow it to the tenants sharing the same tuple and database (they still share tables), rather than removing it. Gates cleared: story 8 (feat(cache): key the query cache and singleflight by tenant #601) and feat(clickhouse): tls, headers, pool sizes and a connection ceiling #603 are in.
  • 7. Per-tenant dedupe store. Done in feat(dedupe): one store per tenant under data_dir/<tenant>/dedupe #602. One dedupe.Managed per tenant rooted at data_dir/<tenant>/dedupe, each following its own dedupe.enabled, built through a factory that takes the tenant id so a shared backend can later put the tenant in the key without touching callers. Ingest picks the instance off the tenant handle. Remote backends and admin routes deferred (Pebble has no per-record TTL; retention is a store feature to design, not a route).
  • 8. Cache keyed by tenant. Done in feat(cache): key the query cache and singleflight by tenant #601. Tenant prefix on pipe/query cache keys and version namespaces; single global ristretto pool unchanged. Covers the singleflight keys too (pipes.go and structured_query.go reuse the cache key), and is the same edit as the TODO above the pipe cache key. Must land before story 6's implementation: a tenant-blind key is a silent cross-tenant read once tenants have their own ClickHouse connections. feat(pipes): resolve table dependencies for cache invalidation #343 (pipe cache invalidation by table dependencies) rewrites the same pipe cache key; it waits until the epic is done and then rebases onto this story's key.
  • 9. Per-tenant auth verifier (JWKS) and policy source. Done in feat(auth): one token verifier per tenant, fetched off the boot path #604 (with the empty-HMAC-key hotfix fix(auth): refuse tokens when no secret and no JWKS URL are configured #607 landed ahead of it). One verifier per tenant, built and rebuilt from the registry's AfterAdopt, released inside App.Close's budget (refactor(app): extract the process wiring from main.go into internal/app #585). Boot does not block on the URL (no-error-on-first-fetch override; fails closed until a fetch succeeds). Refresh stays managed by the library (its own hourly refresh and unknown-kid refetch per tenant; Eric, 2026-09-22: bring it into a scheduler only if scale ever demands it). Add a response size cap on the fetch client. HMAC secret and operator key stay boot-level. Also threads auth's policy source, still bound to tenant 0's store after feat(tenant): resolve X-Tenant-ID before auth and thread the tenant #593: on a tenant route the operator key's admin role reads the request's resolved store (the ops gate reads no policy in nested mode, feat(settings): nested settings directory with a per-tenant registry #598). No dependency on story 6; runs alongside stories 5, 7, 8, 10.
  • 10. Per-tenant CORS. Done in feat(cors): decide the allowlist from the request's tenant #600. Origin list resolved through the request's tenant store.
  • 11. Docs. settings-directory.mdx layout section (both shapes, the no-hybrid rule, tenant 0 semantics, the header, how to add a tenant folder by hand since bootstrap stays single-tenant), api.md for the header, 404/503, and the reload parameter, configuration.mdx for the exempt routes.

Order

Out of scope

Related: #508 (settings reload), #530 (boot fails loudly), #140, #138, #532, #214, #235, #262, #361.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/apiHTTP handlers, routing, middlewarearea/configConfig file, config knobs, hot-reloadarea/streamingSSE / live-query delivery path (/v1/stream)enhancementNew feature or request

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions