From 3022b926392f9cdfa16931ada447607b94fe09b7 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Thu, 24 Sep 2026 23:27:32 -0400 Subject: [PATCH 01/27] feat(config): choose each layer's implementation at boot mq.backend, cache.backend, dedupe.backend and coord.backend select each layer's implementation; only today's in-process one exists per layer and it is the default. Validate refuses an unknown value, internal/app picks the implementation in one switch per layer, data_dir is probed only when a selected backend keeps state there, and boot logs Config.Warnings. Part of #613. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 1 + cmd/wavehouse/main.go | 19 ++- config.yaml | 10 ++ docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/configuration.mdx | 31 +++- docs/src/content/docs/settings-directory.mdx | 2 +- internal/app/app.go | 4 + internal/app/app_test.go | 27 +++- internal/app/wire.go | 63 ++++++-- internal/config/backends.go | 143 ++++++++++++++++ internal/config/backends_test.go | 161 +++++++++++++++++++ internal/config/config.go | 12 +- internal/config/config_test.go | 10 +- tests/integration/setup_test.go | 5 +- tests/integration/tenants_test.go | 5 +- 15 files changed, 453 insertions(+), 42 deletions(-) create mode 100644 internal/config/backends.go create mode 100644 internal/config/backends_test.go diff --git a/CHANGELOG.md b/CHANGELOG.md index 23c0c715..9f2e821b 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,6 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added +- **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block, which until that backend lands is an unknown key and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name every backend: the zero value is not the default, and `app.New` refuses it. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/cmd/wavehouse/main.go b/cmd/wavehouse/main.go index 042257e2..3a1304f0 100644 --- a/cmd/wavehouse/main.go +++ b/cmd/wavehouse/main.go @@ -171,14 +171,17 @@ func run(ctx context.Context) int { return 1 } - // data_dir must be writable before anything dials out, so the refusal - // (and, for the typical cause — a bind mount owned by root rather than - // UID 65532 — the remediation) lands at the top of the log rather than - // after ClickHouse discovery. NATS and Pebble still fail loud on their - // own if the directory changes underneath us. - if err := config.CheckDataDir(cfg.DataDir); err != nil { - logger.Error("check data_dir", "error", err) - return 1 + // data_dir, when a selected backend keeps state there, must be writable + // before anything dials out, so the refusal (and, for the typical cause — + // a bind mount owned by root rather than UID 65532 — the remediation) + // lands at the top of the log rather than after ClickHouse discovery. + // NATS and Pebble still fail loud on their own if the directory changes + // underneath us. + if cfg.NeedsDataDir() { + if err := config.CheckDataDir(cfg.DataDir); err != nil { + logger.Error("check data_dir", "error", err) + return 1 + } } a, err := app.New(ctx, app.Options{ diff --git a/config.yaml b/config.yaml index 53a43502..65c384ea 100644 --- a/config.yaml +++ b/config.yaml @@ -43,9 +43,19 @@ clickhouse: password: "" max_total_conns: 0 # ceiling on open native connections across pools; 0 = none +# Each layer's implementation, chosen at boot. Only the in-process backend +# exists for each today, and it is the default. +mq: + backend: embedded # NATS JetStream under /nats +dedupe: + backend: pebble # Pebble under /pebble +coord: + backend: local + # In-process L1 cache size. The query time-bucket # (query.timestamp_bucket_seconds) is a settings key. cache: + backend: local l1_max_cost: 67108864 # Auth has no on/off switch — the JWT middleware always runs. A request with no diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 6eaf3d54..938751ff 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -90,7 +90,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability, ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe` builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ` hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 8a709426..8d311adc 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -17,7 +17,7 @@ WaveHouse is configured via a YAML file with environment variable overrides. All 2. Environment variables override any values from the YAML file. 3. If no config file exists, all values are read from environment variables. Every key has a default except `settings.dir` (`WH_SETTINGS_DIR`), which must be set either way. 4. Both sources are **strict**. A YAML key this page doesn't list — a typo, or a tunable that has moved to the settings directory (`dlq.enabled`, `clickhouse.addr`, `stream.*`, a leftover `policy:` or `pipes:` block, …) — refuses to boot and names every offending key, so nothing is read, ignored, and believed. A `WH_*` environment variable that binds to no key on this page (`WH_DEDUPE_ENABLED`, `WH_CH_ADDR`, a misspelling) refuses to boot the same way. Two variables have no YAML key and are exempt because they are not config keys at all but process-level settings `main` reads directly: `WH_CONFIG` (below), which locates the file, and `WH_LOG_LEVEL`. Only the `WH_` prefix is checked, since the environment always carries names that aren't WaveHouse's. One outside source does share the prefix. Kubernetes injects `{SERVICE}_SERVICE_HOST`, `{SERVICE}_PORT`, and similar link variables into every pod in a Service's own namespace, for each Service with a cluster IP that existed before the pod started (a headless Service injects nothing, and a Service in another namespace is harmless). The name is uppercased with `-` mapped to `_`, so a Service named `wh` produces `WH_SERVICE_HOST` and `WH_PORT`, one named `wh-foo` produces `WH_FOO_SERVICE_HOST` and `WH_FOO_PORT`, and either way the pod refuses to boot on its next restart. Set `enableServiceLinks: false` on the pod spec, or name the Service something else. The error says so. -5. Before anything dials out, `data_dir` is probed, and boot refuses on any of these: the value is empty; the path exists but is not a directory; the path, or any component above it, is a dangling symlink (a mount that never came up); the directory exists but the process cannot write to it; the directory is absent and its nearest existing ancestor is not writable, so it could not be created. The probe runs before ClickHouse discovery, so the refusal lands at the top of the log, and a permission denial — on the write probe, or on reaching the path at all through a parent without search permission — carries the UID-65532 remediation, since a bind mount owned by root is the typical cause. +5. Before anything dials out, `data_dir` is probed — when a selected [backend](#backends) keeps state there, as the in-process `mq` and `dedupe` backends do — and boot refuses on any of these: the value is empty; the path exists but is not a directory; the path, or any component above it, is a dangling symlink (a mount that never came up); the directory exists but the process cannot write to it; the directory is absent and its nearest existing ancestor is not writable, so it could not be created. The probe runs before ClickHouse discovery, so the refusal lands at the top of the log, and a permission denial — on the write probe, or on reaching the path at all through a parent without search permission — carries the UID-65532 remediation, since a bind mount owned by root is the typical cause. Boot is the validator for this half of configuration: there is no dry run, and a refused boot with the offending key, variable, or path named in the error is the loud signal. The hot-reloadable half has a dry run — `wavehouse validate` — because it is edited under a running server; boot config only ever takes effect through a restart, so the restart is where it is checked. @@ -37,6 +37,19 @@ This page is boot config only — what the platform operator owns (wiring, lifec | --- | --- | ------- | ----------- | | `data_dir` | `WH_DATA_DIR` | `./data` | Root directory for embedded state. NATS JetStream lives at `/nats`; Pebble, holding every tenant's dedupe store while any tenant has dedupe enabled, at `/pebble`. Subdirectory names are conventions, not config — one knob, one mount. **In a container this MUST resolve to a host-backed volume**; the relative default is for local binary use. WaveHouse logs a startup `WARN` when the directory is missing or empty (no prior state). See [Persistent Storage](/deployment#persistent-storage-required-for-containers). | +### Backends + +Each layer's implementation is chosen once, at boot. Today every layer has one backend, the in-process one, and it is the default, so a config that sets none of these keys runs as it always has. A value this build has no backend for refuses boot and names the valid ones. + +| YAML Key | Env Var | Default | Description | +| --- | --- | ------- | ----------- | +| `mq.backend` | `WH_MQ_BACKEND` | `embedded` | The message queue. `embedded`: NATS JetStream inside this process, under `/nats`. It listens on no port, so no other process can reach its queue. | +| `cache.backend` | `WH_CACHE_BACKEND` | `local` | The query-result cache. `local`: in this process, sized by `cache.l1_max_cost`. | +| `dedupe.backend` | `WH_DEDUPE_BACKEND` | `pebble` | Where ingest dedupe keeps the event ids it has seen. `pebble`: in this process, under `/pebble`, open while any tenant has dedupe on. | +| `coord.backend` | `WH_COORD_BACKEND` | `local` | Where the leases for work only one process may do at a time are held. `local`: in this process. | + +Settings for one backend go in a sub-block named after it, `.`, read only when that backend is selected. A sub-block for a backend this build does not have is an unknown key and refuses boot. `mq` and `dedupe` also appear in the settings directory's `config.json`, with different keys (`mq.max_bytes_gb`, `dedupe.enabled`, …): those are per-tenant tunables and stay there, and one written in `config.yaml` refuses boot as an unknown key. + ### Server | YAML Key | Env Var | Default | Description | @@ -97,6 +110,8 @@ WaveHouse's per-role caps are sent as per-query `SETTINGS` on its connection, so ### Message Queue (NATS) +This section describes the `embedded` [backend](#backends), the only one today. + Each tenant's queue has its own disk budget, `mq.max_bytes_gb`, a hot-reloadable key in the [Settings Directory](/settings-directory#message-queue) — there is no boot-config knob for it. **Durability.** The embedded server runs with JetStream `SyncAlways`, so every event is `fsync`'d to disk before `POST /v1/ingest` returns `200`. This makes your storage's `fsync` latency your ingest latency floor — see [Durability & Storage](/durability) to check whether your substrate can sustain it. There is no knob to relax this today ([#139](https://github.com/Wave-RF/WaveHouse/issues/139) tracks a configurable group-commit interval). @@ -191,9 +206,19 @@ clickhouse: # headers and pool sizes are settings (config.json) max_total_conns: 0 # ceiling on open native connections; 0 = none +mq: + backend: embedded # in-process NATS JetStream under /nats + cache: + backend: local l1_max_cost: 67108864 +dedupe: + backend: pebble # in-process Pebble under /pebble + +coord: + backend: local + auth: jwt_secret: change-me-in-production # jwks_url and role_claim are settings (config.json) operator_key: "" # non-JWT full-access operator credential (Authorization: Operator , or X-Operator-Key); empty disables @@ -239,7 +264,11 @@ WH_SERVER_SHUTDOWN_TIMEOUT=10 WH_CH_PASSWORD= WH_CH_MAX_TOTAL_CONNS=0 +WH_MQ_BACKEND=embedded +WH_CACHE_BACKEND=local WH_CACHE_L1_MAX_COST=67108864 +WH_DEDUPE_BACKEND=pebble +WH_COORD_BACKEND=local WH_AUTH_JWT_SECRET=change-me-in-production WH_AUTH_OPERATOR_KEY= diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index c0e7a19b..56135462 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -179,7 +179,7 @@ The tenant tunables. Every key is required (a missing one is a validation error) } ``` -What stays in boot config is only what cannot change under a running process — resource sizing (`data_dir`, `cache.l1_max_cost`, `clickhouse.max_total_conns`), the listeners, the observability exporters — and the **secrets**: `clickhouse.password`, `auth.jwt_secret`, `auth.operator_key`. Secrets never belong in a tracked JSON file, so they stay in the environment and are combined with the wiring here on every (re)connect; rotating one is a restart. See [Configuration](/configuration). Everything else lives here and reloads. +What stays in boot config is only what cannot change under a running process — the implementation each layer runs on (`mq.backend`, `cache.backend`, `dedupe.backend`, `coord.backend`), resource sizing (`data_dir`, `cache.l1_max_cost`, `clickhouse.max_total_conns`), the listeners, the observability exporters — and the **secrets**: `clickhouse.password`, `auth.jwt_secret`, `auth.operator_key`. Secrets never belong in a tracked JSON file, so they stay in the environment and are combined with the wiring here on every (re)connect; rotating one is a restart. See [Configuration](/configuration). Everything else lives here and reloads. ## Deduplication diff --git a/internal/app/app.go b/internal/app/app.go index dc15bf5d..726dad19 100644 --- a/internal/app/app.go +++ b/internal/app/app.go @@ -170,6 +170,10 @@ func New(ctx context.Context, opts Options) (app *App, err error) { return nil, err } a.wireObservability(ctx) + // After observability, so an OTLP log pipeline carries them too. + for _, w := range a.cfg.Warnings() { + slog.Warn(w) + } if err := a.wireClickHouse(); err != nil { return nil, err } diff --git a/internal/app/app_test.go b/internal/app/app_test.go index 2ecd77d9..10658ddc 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -95,7 +95,10 @@ func testConfig(t *testing.T, settingsDir string) *config.Config { return &config.Config{ DataDir: t.TempDir(), Server: config.Server{Port: closedPort(t), ShutdownTimeout: 2}, - Cache: config.Cache{L1MaxCost: 1 << 20}, + MQ: config.MQ{Backend: config.MQEmbedded}, + Cache: config.Cache{Backend: config.CacheLocal, L1MaxCost: 1 << 20}, + Dedupe: config.Dedupe{Backend: config.DedupePebble}, + Coord: config.Coord{Backend: config.CoordLocal}, Auth: config.Auth{JWTSecret: "unit-test-secret"}, Settings: config.Settings{Dir: settingsDir}, } @@ -535,6 +538,28 @@ func TestNew_NestedDedupeStoreFollowsEachTenant(t *testing.T) { assert.False(t, restored.Open(), "Close releases every open store") } +// Validate refuses a backend no layer has a case for, so the switch's default +// is reached only by a Config built by hand; it must refuse boot, not wire +// nothing. +func TestNew_RefusesALayerWithoutABackend(t *testing.T) { + for _, tc := range []struct { + key string + unset func(*config.Config) + }{ + {"dedupe.backend", func(c *config.Config) { c.Dedupe.Backend = "" }}, + {"mq.backend", func(c *config.Config) { c.MQ.Backend = "" }}, + {"cache.backend", func(c *config.Config) { c.Cache.Backend = "" }}, + } { + t.Run(tc.key, func(t *testing.T) { + guardGlobals(t) + cfg := testConfig(t, writeSettings(t, nil)) + tc.unset(cfg) + _, err := New(t.Context(), Options{Config: cfg}) + require.ErrorContains(t, err, tc.key+` "" has no wiring`) + }) + } +} + // A Pebble instance that cannot open follows the registry's own rule for the // shape: a flat directory refuses boot, like every other store, and a nested // one fails closed for every tenant with dedupe on, since they share the diff --git a/internal/app/wire.go b/internal/app/wire.go index 60cbdcd8..d5372025 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -471,7 +471,18 @@ func (a *App) wireDiscovery(ctx context.Context) { a.add(component{name: "schema discovery", close: d.close}) } -// wireDedupe builds the dedupe stores: one per tenant (#583 story 7), each +// wireDedupe builds the dedupe stores — the one place the implementation is +// chosen. +func (a *App) wireDedupe() error { + switch b := a.cfg.Dedupe.Backend; b { + case config.DedupePebble: + return a.wirePebbleDedupe() + default: + return unreachableBackend("dedupe.backend", b) + } +} + +// wirePebbleDedupe builds the dedupe stores: one per tenant (#583 story 7), each // following its own tenant's hot-reloadable dedupe.enabled, over the // embedded Pebble implementation, which is handed data_dir and decides the // rest: every tenant's seen ids in one instance there, open while any @@ -488,7 +499,7 @@ func (a *App) wireDiscovery(ctx context.Context) { // than silently publishing un-deduped, since the files asked for dedupe; // nested fails closed the same way at boot too, for every tenant with // dedupe on, the next reload retrying, so it never costs the process. -func (a *App) wireDedupe() error { +func (a *App) wirePebbleDedupe() error { nested := a.tenants.Nested() embedded := dedupe.NewEmbedded(a.cfg.DataDir) stores := dedupe.NewStores(embedded.Tenant) @@ -530,10 +541,20 @@ func (a *App) wireDedupe() error { return nil } -// wireMQ starts the MQ — the embedded NATS under data_dir/nats, the one -// place the implementation is chosen; everything after it sees mq.Broker — -// and hands it each served tenant's mq.max_bytes_gb, which opens that -// tenant's queue the first time. The budget is hot-reloadable: after every +// wireMQ starts the MQ — the one place the implementation is chosen; +// everything after it sees mq.Broker. +func (a *App) wireMQ() error { + switch b := a.cfg.MQ.Backend; b { + case config.MQEmbedded: + return a.wireEmbeddedMQ() + default: + return unreachableBackend("mq.backend", b) + } +} + +// wireEmbeddedMQ starts the embedded NATS under data_dir/nats and hands it +// each served tenant's mq.max_bytes_gb, which opens that tenant's queue the +// first time. The budget is hot-reloadable: after every // reload the registry applies, each served tenant's is handed over again, // and the MQ owns how it is split across the tenant's queues and keeps them // consistent (see mq.Broker.SetMaxBytes). A tenant no longer served keeps @@ -543,7 +564,7 @@ func (a *App) wireDedupe() error { // previous budget; a nested directory logs it at boot too, so it never costs // the process — the tenant's ingest answers 503 until a reload opens its // queue. The hook is registered before the boot apply, as the dedupe one is. -func (a *App) wireMQ() error { +func (a *App) wireEmbeddedMQ() error { dir := filepath.Join(a.cfg.DataDir, "nats") config.WarnIfFreshDataDir("nats", dir) var broker mq.Broker @@ -591,16 +612,28 @@ func (a *App) wireMQ() error { return nil } -// wireCache opens the L1 cache — the only tier in standalone mode. +// wireCache opens the query-result cache — the one place the implementation +// is chosen. func (a *App) wireCache() error { - l1, err := cache.NewLocal(a.cfg.Cache.L1MaxCost) - if err != nil { - return fmt.Errorf("cache init: %w", err) + switch b := a.cfg.Cache.Backend; b { + case config.CacheLocal: + l1, err := cache.NewLocal(a.cfg.Cache.L1MaxCost) + if err != nil { + return fmt.Errorf("cache init: %w", err) + } + a.cache = l1 + a.add(component{name: "cache", close: withoutContext(l1.Close)}) + return nil + default: + return unreachableBackend("cache.backend", b) } - // TODO: eventually this is where we can switch between ristretto, redis, tiered (both), etc - a.cache = l1 - a.add(component{name: "cache", close: withoutContext(l1.Close)}) - return nil +} + +// unreachableBackend is each layer switch's default case. config.Validate +// refuses a backend with no case, so reaching it means a Config built by hand +// without one (the zero value is not the default), or a case missing here. +func unreachableBackend[T ~string](key string, got T) error { + return fmt.Errorf("%s %q has no wiring: a Config built without config.Load must name every backend", key, got) } // wireSweeper adds the active sweeper — purges messages that are both diff --git a/internal/config/backends.go b/internal/config/backends.go new file mode 100644 index 00000000..8328cea6 --- /dev/null +++ b/internal/config/backends.go @@ -0,0 +1,143 @@ +package config + +import ( + "fmt" + "slices" + "strings" +) + +// Each layer's implementation is chosen here, once, at boot: `.backend` +// names it, and the default is today's in-process one. Settings for one +// backend go in `.`, a sub-block read only when that backend +// is selected. Adding a backend is its constant in the layer's list, a case +// in the layer's validate for its sub-block, and a case in the layer's +// wire function in internal/app — nothing else in Validate changes. + +// MQBackend names the message queue implementation. +type MQBackend string + +// MQEmbedded is the NATS JetStream server inside this process, under +// /nats. +const MQEmbedded MQBackend = "embedded" + +var mqBackends = []MQBackend{MQEmbedded} + +// MQ selects the message queue. The per-tenant byte budget, mq.max_bytes_gb, +// is a settings-directory key, not this block's. +type MQ struct { + Backend MQBackend `yaml:"backend" env:"WH_MQ_BACKEND" env-default:"embedded"` +} + +func (m MQ) validate() error { + return checkBackend("mq.backend", "WH_MQ_BACKEND", m.Backend, mqBackends) +} + +// CacheBackend names the query-result cache implementation. +type CacheBackend string + +// CacheLocal is the in-process Ristretto cache, sized by cache.l1_max_cost. +const CacheLocal CacheBackend = "local" + +var cacheBackends = []CacheBackend{CacheLocal} + +// Cache selects and sizes the query-result cache. The time-range bucket +// structured queries normalize to is a settings-directory key +// (query.timestamp_bucket_seconds) — query shaping, not process memory. +type Cache struct { + Backend CacheBackend `yaml:"backend" env:"WH_CACHE_BACKEND" env-default:"local"` + L1MaxCost int64 `yaml:"l1_max_cost" env:"WH_CACHE_L1_MAX_COST" env-default:"67108864"` +} + +func (c Cache) validate() error { + return checkBackend("cache.backend", "WH_CACHE_BACKEND", c.Backend, cacheBackends) +} + +// DedupeBackend names where ingest dedupe keeps the ids it has seen. +type DedupeBackend string + +// DedupePebble is the Pebble instance inside this process, under +// /pebble, opened while any tenant has dedupe on. +const DedupePebble DedupeBackend = "pebble" + +var dedupeBackends = []DedupeBackend{DedupePebble} + +// Dedupe selects the dedupe store. Whether a tenant dedupes, and on which +// field, are settings-directory keys, not this block's. +type Dedupe struct { + Backend DedupeBackend `yaml:"backend" env:"WH_DEDUPE_BACKEND" env-default:"pebble"` +} + +func (d Dedupe) validate() error { + return checkBackend("dedupe.backend", "WH_DEDUPE_BACKEND", d.Backend, dedupeBackends) +} + +// CoordBackend names the lease implementation singleton work (the sweeper) +// is elected through. +type CoordBackend string + +// CoordLocal holds leases in this process: correct while no other process +// shares its queue. +const CoordLocal CoordBackend = "local" + +var coordBackends = []CoordBackend{CoordLocal} + +// Coord selects the coordination layer. +type Coord struct { + Backend CoordBackend `yaml:"backend" env:"WH_COORD_BACKEND" env-default:"local"` +} + +func (c Coord) validate() error { + return checkBackend("coord.backend", "WH_COORD_BACKEND", c.Backend, coordBackends) +} + +// checkBackend refuses a backend this build has no implementation for, +// listing the ones it has. env repeats the struct tag's literal: a tag can't +// reference a constant. +func checkBackend[T ~string](key, env string, got T, valid []T) error { + if slices.Contains(valid, got) { + return nil + } + names := make([]string, len(valid)) + for i, v := range valid { + names[i] = string(v) + } + return fmt.Errorf("%s (%s) %q is not a backend this build has; valid: %s", key, env, got, strings.Join(names, ", ")) +} + +// validateBackends checks every layer's backend and its sub-block. +func (c *Config) validateBackends() error { + for _, check := range []func() error{c.MQ.validate, c.Cache.validate, c.Dedupe.validate, c.Coord.validate} { + if err := check(); err != nil { + return err + } + } + return nil +} + +// Distributed reports whether the message queue is shared with other +// processes. The embedded one listens on no port, so while it is selected +// every process is an island: nothing else can reach its queue. +func (c *Config) Distributed() bool { return c.MQ.Backend != MQEmbedded } + +// NeedsDataDir reports whether a selected backend keeps state under data_dir, +// and so whether boot must probe it (CheckDataDir). +func (c *Config) NeedsDataDir() bool { + return c.MQ.Backend == MQEmbedded || c.Dedupe.Backend == DedupePebble +} + +// Warnings returns what a valid configuration is still likely to get wrong, +// one line each, for boot to log at WARN. They are not errors because each is +// correct for a single replica, and one process cannot count its replicas. +func (c *Config) Warnings() []string { + if !c.Distributed() { + return nil + } + var out []string + if c.Cache.Backend == CacheLocal { + out = append(out, "cache.backend=local with a shared mq.backend is correct for one replica only: an event ingested on another replica never invalidates this one's cache, so its reads stay stale until the cached entry expires") + } + if c.Dedupe.Backend == DedupePebble { + out = append(out, "dedupe.backend=pebble with a shared mq.backend dedupes per replica only: an id seen by another replica is not seen by this one") + } + return out +} diff --git a/internal/config/backends_test.go b/internal/config/backends_test.go new file mode 100644 index 00000000..0e70dc7a --- /dev/null +++ b/internal/config/backends_test.go @@ -0,0 +1,161 @@ +package config + +import ( + "os" + "path/filepath" + "testing" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" +) + +// withDefaultBackends sets what Load's env-defaults would: a literal Config +// names no backend, and Validate refuses that. +func withDefaultBackends(c Config) *Config { + c.MQ.Backend, c.Cache.Backend = MQEmbedded, CacheLocal + c.Dedupe.Backend, c.Coord.Backend = DedupePebble, CoordLocal + return &c +} + +func defaultBackends() Config { + return *withDefaultBackends(Config{Server: Server{Port: 8080}, Settings: Settings{Dir: "./settings"}}) +} + +func TestLoad_BackendDefaults(t *testing.T) { + t.Parallel() + cfg, err := Load("nonexistent.yaml") + require.NoError(t, err) + assert.Equal(t, MQEmbedded, cfg.MQ.Backend) + assert.Equal(t, CacheLocal, cfg.Cache.Backend) + assert.Equal(t, DedupePebble, cfg.Dedupe.Backend) + assert.Equal(t, CoordLocal, cfg.Coord.Backend) + assert.False(t, cfg.Distributed()) + assert.True(t, cfg.NeedsDataDir()) + assert.Empty(t, cfg.Warnings()) +} + +func TestLoad_BackendsFromEnv(t *testing.T) { + t.Setenv("WH_MQ_BACKEND", "embedded") + t.Setenv("WH_CACHE_BACKEND", "local") + t.Setenv("WH_DEDUPE_BACKEND", "pebble") + t.Setenv("WH_COORD_BACKEND", "local") + cfg, err := Load("nonexistent.yaml") + require.NoError(t, err) + assert.Equal(t, MQEmbedded, cfg.MQ.Backend) + assert.Equal(t, CoordLocal, cfg.Coord.Backend) +} + +func TestLoad_BackendFromEnvRefusesAnUnknownValue(t *testing.T) { + t.Setenv("WH_MQ_BACKEND", "nats") + _, err := Load("nonexistent.yaml") + require.Error(t, err) + assert.Contains(t, err.Error(), `mq.backend (WH_MQ_BACKEND) "nats" is not a backend this build has; valid: embedded`) +} + +func TestLoad_BackendsFromYAML(t *testing.T) { + t.Parallel() + path := filepath.Join(t.TempDir(), "config.yaml") + require.NoError(t, os.WriteFile(path, []byte(` +mq: + backend: embedded +cache: + backend: local + l1_max_cost: 1024 +dedupe: + backend: pebble +coord: + backend: local +`), 0o600)) + cfg, err := Load(path) + require.NoError(t, err) + assert.Equal(t, MQEmbedded, cfg.MQ.Backend) + assert.Equal(t, CacheLocal, cfg.Cache.Backend) + assert.Equal(t, int64(1024), cfg.Cache.L1MaxCost) + assert.Equal(t, DedupePebble, cfg.Dedupe.Backend) + assert.Equal(t, CoordLocal, cfg.Coord.Backend) +} + +// A sub-block written before its backend exists, and a settings-directory +// key under a block both files share, are unknown keys — not read and ignored. +func TestLoad_BackendBlocksRefuseUnknownKeys(t *testing.T) { + t.Parallel() + path := filepath.Join(t.TempDir(), "config.yaml") + require.NoError(t, os.WriteFile(path, []byte(` +mq: + backend: embedded + max_bytes_gb: 5 + nats: + urls: nats://localhost:4222 +dedupe: + enabled: true +`), 0o600)) + _, err := Load(path) + require.Error(t, err) + assert.Contains(t, err.Error(), "dedupe.enabled, mq.max_bytes_gb, mq.nats") + assert.Contains(t, err.Error(), EnvSettingsDir) +} + +func TestUnboundEnv_KnowsTheBackendVariables(t *testing.T) { + t.Parallel() + assert.Empty(t, unboundEnv([]string{ + "WH_MQ_BACKEND=embedded", "WH_CACHE_BACKEND=local", + "WH_DEDUPE_BACKEND=pebble", "WH_COORD_BACKEND=local", + })) +} + +func TestValidate_UnknownBackend(t *testing.T) { + t.Parallel() + cases := []struct { + name string + set func(*Config) + want string + }{ + {"mq", func(c *Config) { c.MQ.Backend = "kafka" }, `mq.backend (WH_MQ_BACKEND) "kafka" is not a backend this build has; valid: embedded`}, + {"cache", func(c *Config) { c.Cache.Backend = "redis" }, `cache.backend (WH_CACHE_BACKEND) "redis" is not a backend this build has; valid: local`}, + {"dedupe", func(c *Config) { c.Dedupe.Backend = "dynamodb" }, `dedupe.backend (WH_DEDUPE_BACKEND) "dynamodb" is not a backend this build has; valid: pebble`}, + {"coord", func(c *Config) { c.Coord.Backend = "nats" }, `coord.backend (WH_COORD_BACKEND) "nats" is not a backend this build has; valid: local`}, + // The zero value, which a Config built without Load carries. + {"empty", func(c *Config) { c.MQ.Backend = "" }, `mq.backend (WH_MQ_BACKEND) "" is not a backend`}, + } + for _, tc := range cases { + t.Run(tc.name, func(t *testing.T) { + t.Parallel() + cfg := defaultBackends() + require.NoError(t, cfg.Validate()) + tc.set(&cfg) + err := cfg.Validate() + require.Error(t, err) + assert.Contains(t, err.Error(), tc.want) + }) + } +} + +// Every warning keys on a shared queue, which no backend offers yet, so the +// value is set directly: Warnings reads the choice, it doesn't validate it. +func TestWarnings_SharedQueue(t *testing.T) { + t.Parallel() + cfg := defaultBackends() + assert.Empty(t, cfg.Warnings()) + + cfg.MQ.Backend = "shared" + require.True(t, cfg.Distributed()) + got := cfg.Warnings() + require.Len(t, got, 2) + assert.Contains(t, got[0], "cache.backend=local") + assert.Contains(t, got[1], "dedupe.backend=pebble") + + cfg.Cache.Backend, cfg.Dedupe.Backend = "shared", "shared" + assert.Empty(t, cfg.Warnings()) +} + +func TestNeedsDataDir(t *testing.T) { + t.Parallel() + cfg := defaultBackends() + assert.True(t, cfg.NeedsDataDir()) + cfg.MQ.Backend = "shared" + assert.True(t, cfg.NeedsDataDir(), "pebble dedupe still keeps state under data_dir") + cfg.Dedupe.Backend = "shared" + assert.False(t, cfg.NeedsDataDir()) + cfg.MQ.Backend = MQEmbedded + assert.True(t, cfg.NeedsDataDir(), "the embedded mq keeps state under data_dir") +} diff --git a/internal/config/config.go b/internal/config/config.go index 68b0314b..cc756696 100644 --- a/internal/config/config.go +++ b/internal/config/config.go @@ -19,7 +19,10 @@ type Config struct { DataDir string `yaml:"data_dir" env:"WH_DATA_DIR" env-default:"./data"` Server Server `yaml:"server"` ClickHouse ClickHouse `yaml:"clickhouse"` + MQ MQ `yaml:"mq"` Cache Cache `yaml:"cache"` + Dedupe Dedupe `yaml:"dedupe"` + Coord Coord `yaml:"coord"` Auth Auth `yaml:"auth"` OTel OTel `yaml:"otel"` Prometheus Prometheus `yaml:"prometheus"` @@ -132,13 +135,6 @@ type ClickHouse struct { MaxTotalConns int `yaml:"max_total_conns" env:"WH_CH_MAX_TOTAL_CONNS" env-default:"0"` } -// Cache sizes the in-process L1 cache. The time-range bucket structured -// queries normalize to is a settings-directory key -// (query.timestamp_bucket_seconds) — query shaping, not process memory. -type Cache struct { - L1MaxCost int64 `yaml:"l1_max_cost" env:"WH_CACHE_L1_MAX_COST" env-default:"67108864"` -} - // Auth holds the authentication secrets. The verifier wiring — `jwks_url`, // `role_claim` — is the settings directory's `auth` block (hot-reloadable: // a change rebuilds the verifier). There is no on/off switch: the middleware always runs. A request @@ -224,7 +220,7 @@ func (c *Config) Validate() error { } } - return nil + return c.validateBackends() } // Load reads config from a YAML file (if it exists) with env var overrides. diff --git a/internal/config/config_test.go b/internal/config/config_test.go index 822d639d..ee8b07cc 100644 --- a/internal/config/config_test.go +++ b/internal/config/config_test.go @@ -203,7 +203,7 @@ func TestValidate_SampleRatesIgnoredWhenObservabilityDisabled(t *testing.T) { Logs: OTelLogs{SampleRate: -1}, }, } - assert.NoError(t, cfg.Validate()) + assert.NoError(t, withDefaultBackends(cfg).Validate()) } func TestValidate_SampleRatesIgnoredWhenSignalDisabled(t *testing.T) { @@ -219,7 +219,7 @@ func TestValidate_SampleRatesIgnoredWhenSignalDisabled(t *testing.T) { Logs: OTelLogs{Enabled: false, SampleRate: -1}, }, } - assert.NoError(t, cfg.Validate()) + assert.NoError(t, withDefaultBackends(cfg).Validate()) } func TestLoad_Defaults_PrometheusDisabled(t *testing.T) { @@ -336,7 +336,7 @@ func TestValidate_PrometheusV1PathAllowedOnSidecarPort(t *testing.T) { Settings: Settings{Dir: "./settings"}, Prometheus: Prometheus{Enabled: true, Path: "/v1/metrics", Port: 9091}, } - assert.NoError(t, cfg.Validate()) + assert.NoError(t, withDefaultBackends(cfg).Validate()) } func TestValidate_PrometheusOnly_NoOTel(t *testing.T) { @@ -348,7 +348,7 @@ func TestValidate_PrometheusOnly_NoOTel(t *testing.T) { Settings: Settings{Dir: "./settings"}, Prometheus: Prometheus{Enabled: true, Path: "/metrics", Port: 0}, } - assert.NoError(t, cfg.Validate()) + assert.NoError(t, withDefaultBackends(cfg).Validate()) } func TestValidate_PrometheusIgnoredWhenDisabled(t *testing.T) { @@ -365,7 +365,7 @@ func TestValidate_PrometheusIgnoredWhenDisabled(t *testing.T) { Port: 8080, }, } - assert.NoError(t, cfg.Validate()) + assert.NoError(t, withDefaultBackends(cfg).Validate()) } // TestEnvSettingsDir_MatchesStructTag pins the exported constant to the diff --git a/tests/integration/setup_test.go b/tests/integration/setup_test.go index a01064a2..ade560f7 100644 --- a/tests/integration/setup_test.go +++ b/tests/integration/setup_test.go @@ -161,7 +161,10 @@ func setup() (int, func()) { DataDir: dataDir, Server: config.Server{ShutdownTimeout: 10}, ClickHouse: config.ClickHouse{Password: testCHPassword}, - Cache: config.Cache{L1MaxCost: 1 << 30}, // 1 GB + MQ: config.MQ{Backend: config.MQEmbedded}, + Cache: config.Cache{Backend: config.CacheLocal, L1MaxCost: 1 << 30}, // 1 GB + Dedupe: config.Dedupe{Backend: config.DedupePebble}, + Coord: config.Coord{Backend: config.CoordLocal}, Settings: config.Settings{Dir: settingsDir}, } a, err := app.New(ctx, app.Options{Config: cfg, Listener: ln}) diff --git a/tests/integration/tenants_test.go b/tests/integration/tenants_test.go index 40c5a700..ca00f42f 100644 --- a/tests/integration/tenants_test.go +++ b/tests/integration/tenants_test.go @@ -62,7 +62,10 @@ func TestNestedDirectory_PerTenantPoolsAndDiscovery(t *testing.T) { Server: config.Server{ShutdownTimeout: 10}, ClickHouse: config.ClickHouse{Password: testCHPassword}, Auth: config.Auth{OperatorKey: operatorKey}, - Cache: config.Cache{L1MaxCost: 1 << 20}, + MQ: config.MQ{Backend: config.MQEmbedded}, + Cache: config.Cache{Backend: config.CacheLocal, L1MaxCost: 1 << 20}, + Dedupe: config.Dedupe{Backend: config.DedupePebble}, + Coord: config.Coord{Backend: config.CoordLocal}, Settings: config.Settings{Dir: root}, } a, err := app.New(ctx, app.Options{Config: cfg, Listener: ln}) From fde17ba4578a2a03521e085a77fd2d3326a75c7d Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Thu, 24 Sep 2026 23:30:43 -0400 Subject: [PATCH 02/27] docs(config): say coord.backend is reserved; sync the boot-config lists Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- config.yaml | 2 +- docs/src/content/docs/architecture.md | 7 ++++--- docs/src/content/docs/configuration.mdx | 4 ++-- internal/app/wire.go | 2 +- internal/config/backends.go | 8 ++++---- internal/settings/settings.go | 8 +++++--- 7 files changed, 18 insertions(+), 15 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 9f2e821b..6564775a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block, which until that backend lands is an unknown key and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name every backend: the zero value is not the default, and `app.New` refuses it. +- **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block, which until that backend lands is an unknown key and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name the `mq`, `cache` and `dedupe` backends: the zero value is not the default, and `app.New` refuses it. `coord.backend` is reserved: nothing reads it until the lease layer lands, and the sweeper still runs in every process. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/config.yaml b/config.yaml index 65c384ea..5519b78c 100644 --- a/config.yaml +++ b/config.yaml @@ -50,7 +50,7 @@ mq: dedupe: backend: pebble # Pebble under /pebble coord: - backend: local + backend: local # reserved: nothing is elected yet # In-process L1 cache size. The query time-bucket # (query.timestamp_bucket_seconds) is a settings key. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 938751ff..42591807 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -89,7 +89,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring -- **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability, ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. +- **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability (after which each `Config.Warnings` line is logged at `WARN`), ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. - **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out @@ -115,8 +115,9 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ ### `config/` — Configuration -- **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). -- **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load`, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. +- **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — the implementation each layer runs on, resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). +- **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load` when `Config.NeedsDataDir` says a selected backend keeps state there, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. +- **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` at the end of `Validate`), which refuses a value not on the list and names the ones that are. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns the valid combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), which `app.New` logs at `WARN`. - **strict.go** — `rejectUnknownKeys`, the YAML half: re-reads the file as a generic tree and walks it against the struct's `yaml` tags, listing every key the struct doesn't declare. cleanenv itself is lenient by design, which is exactly wrong for boot config once keys have moved to the settings directory. - **persistence.go** — `WarnIfFreshDataDir` logs the startup `WARN` when `data_dir` is missing or empty (on a redeploy, the sign that the volume didn't persist); `LogStorageInitError` attaches the UID-65532 `permissionHint` to a NATS or Pebble open failure that looks like a permission denial — the same hint string `CheckDataDir` uses. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 8d311adc..ee264bf4 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -46,7 +46,7 @@ Each layer's implementation is chosen once, at boot. Today every layer has one b | `mq.backend` | `WH_MQ_BACKEND` | `embedded` | The message queue. `embedded`: NATS JetStream inside this process, under `/nats`. It listens on no port, so no other process can reach its queue. | | `cache.backend` | `WH_CACHE_BACKEND` | `local` | The query-result cache. `local`: in this process, sized by `cache.l1_max_cost`. | | `dedupe.backend` | `WH_DEDUPE_BACKEND` | `pebble` | Where ingest dedupe keeps the event ids it has seen. `pebble`: in this process, under `/pebble`, open while any tenant has dedupe on. | -| `coord.backend` | `WH_COORD_BACKEND` | `local` | Where the leases for work only one process may do at a time are held. `local`: in this process. | +| `coord.backend` | `WH_COORD_BACKEND` | `local` | Reserved for the leases that will elect work only one process may do at a time, such as the sweeper. Nothing is elected yet: every process runs its own sweeper, and `local`, the only value, changes nothing. | Settings for one backend go in a sub-block named after it, `.`, read only when that backend is selected. A sub-block for a backend this build does not have is an unknown key and refuses boot. `mq` and `dedupe` also appear in the settings directory's `config.json`, with different keys (`mq.max_bytes_gb`, `dedupe.enabled`, …): those are per-tenant tunables and stay there, and one written in `config.yaml` refuses boot as an unknown key. @@ -217,7 +217,7 @@ dedupe: backend: pebble # in-process Pebble under /pebble coord: - backend: local + backend: local # reserved: nothing is elected yet auth: jwt_secret: change-me-in-production # jwks_url and role_claim are settings (config.json) diff --git a/internal/app/wire.go b/internal/app/wire.go index d5372025..3cfcad4b 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -633,7 +633,7 @@ func (a *App) wireCache() error { // refuses a backend with no case, so reaching it means a Config built by hand // without one (the zero value is not the default), or a case missing here. func unreachableBackend[T ~string](key string, got T) error { - return fmt.Errorf("%s %q has no wiring: a Config built without config.Load must name every backend", key, got) + return fmt.Errorf("%s %q has no wiring: a Config built without config.Load must name the backend of every layer it wires", key, got) } // wireSweeper adds the active sweeper — purges messages that are both diff --git a/internal/config/backends.go b/internal/config/backends.go index 8328cea6..f2ab9330 100644 --- a/internal/config/backends.go +++ b/internal/config/backends.go @@ -71,12 +71,12 @@ func (d Dedupe) validate() error { return checkBackend("dedupe.backend", "WH_DEDUPE_BACKEND", d.Backend, dedupeBackends) } -// CoordBackend names the lease implementation singleton work (the sweeper) -// is elected through. +// CoordBackend names where leases for singleton work (the sweeper) are held. +// Nothing reads it yet: the lease layer (#613) wires it. type CoordBackend string -// CoordLocal holds leases in this process: correct while no other process -// shares its queue. +// CoordLocal holds leases in this process, which is enough while no other +// process shares its queue. const CoordLocal CoordBackend = "local" var coordBackends = []CoordBackend{CoordLocal} diff --git a/internal/settings/settings.go b/internal/settings/settings.go index 55ec089d..8db981b3 100644 --- a/internal/settings/settings.go +++ b/internal/settings/settings.go @@ -59,9 +59,11 @@ type PipesFile struct { // TenantConfig is the shape of config.json: the behavioral tunables that // migrate out of boot config. Boot config (config.yaml/env) keeps only what -// cannot change under a running process — resource sizing (`data_dir`, -// `cache.l1_max_cost`), listeners, the observability -// exporters — and the secrets (`clickhouse.password`, `auth.jwt_secret`, +// cannot change under a running process — the implementation each layer +// runs on (`mq.backend`, `cache.backend`, `dedupe.backend`, +// `coord.backend`), resource sizing (`data_dir`, `cache.l1_max_cost`, +// `clickhouse.max_total_conns`), listeners, the observability exporters — +// and the secrets (`clickhouse.password`, `auth.jwt_secret`, // `auth.operator_key`), which never belong in a tracked JSON file. Every // block and every top-level key inside it is REQUIRED: the binary carries no // compiled defaults, so the adopted snapshot is exactly what the files say. From f129d5775f2ba700c4997d37e3999e5b7f062849 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Thu, 24 Sep 2026 23:36:56 -0400 Subject: [PATCH 03/27] docs(config): no backend has a sub-block yet; index backends.go in AGENTS.md Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- AGENTS.md | 2 +- CHANGELOG.md | 2 +- docs/src/content/docs/configuration.mdx | 2 +- 3 files changed, 3 insertions(+), 3 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 16595721..a7e4a28e 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -34,7 +34,7 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, `...
.` for a namespace — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version that leads every namespace key of one tenant, orphaning every cached query keyed by its tables in one step (a pipe result names no table and keeps its TTL, [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the table-keyed cache (the structured-query results) of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) -- **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability — boot is the validator, there is no dry run +- **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability when a selected backend keeps state there (`NeedsDataDir`); `backends.go` holds each layer's `.backend` (only the in-process value today; `coord.backend` reserved) — boot is the validator, there is no dry run - **`dedupe/`** — `Deduplicator` interface → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) diff --git a/CHANGELOG.md b/CHANGELOG.md index 6564775a..0aac1259 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block, which until that backend lands is an unknown key and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name the `mq`, `cache` and `dedupe` backends: the zero value is not the default, and `app.New` refuses it. `coord.backend` is reserved: nothing reads it until the lease layer lands, and the sweeper still runs in every process. +- **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block; no backend has settings yet, so every such sub-block is an unknown key for now and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name the `mq`, `cache` and `dedupe` backends: the zero value is not the default, and `app.New` refuses it. `coord.backend` is reserved: nothing reads it until the lease layer lands, and the sweeper still runs in every process. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index ee264bf4..193a6c21 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -48,7 +48,7 @@ Each layer's implementation is chosen once, at boot. Today every layer has one b | `dedupe.backend` | `WH_DEDUPE_BACKEND` | `pebble` | Where ingest dedupe keeps the event ids it has seen. `pebble`: in this process, under `/pebble`, open while any tenant has dedupe on. | | `coord.backend` | `WH_COORD_BACKEND` | `local` | Reserved for the leases that will elect work only one process may do at a time, such as the sweeper. Nothing is elected yet: every process runs its own sweeper, and `local`, the only value, changes nothing. | -Settings for one backend go in a sub-block named after it, `.`, read only when that backend is selected. A sub-block for a backend this build does not have is an unknown key and refuses boot. `mq` and `dedupe` also appear in the settings directory's `config.json`, with different keys (`mq.max_bytes_gb`, `dedupe.enabled`, …): those are per-tenant tunables and stay there, and one written in `config.yaml` refuses boot as an unknown key. +Settings for one backend will go in a sub-block named after it, `.`, read only when that backend is selected. No backend has settings yet, so today any such sub-block, `mq.embedded` included, is an unknown key and refuses boot. `mq` and `dedupe` also appear in the settings directory's `config.json`, with different keys (`mq.max_bytes_gb`, `dedupe.enabled`, …): those are per-tenant tunables and stay there, and one written in `config.yaml` refuses boot as an unknown key. ### Server From 67df3173993f30d18cf92c00810fa1e886346ad6 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Thu, 24 Sep 2026 23:56:52 -0400 Subject: [PATCH 04/27] perf(cache): flat version index, pruned per tenant The local version index nested each table under its tenant's version and each scope under its table's, and was never pruned: every InvalidateTenant left the tenant's whole index behind. It now holds one version per tenant, (tenant, table) and (tenant, table, scope), bumped in place. A tenant bump drops the tenant's index and its next key gets a process-unique generation; a table bump drops the table's scopes. After each reload the wiring prunes the index to the tenants served. Fixes #262 for the local backend. Part of #613. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- AGENTS.md | 4 +- CHANGELOG.md | 1 + docs/src/content/docs/architecture.md | 2 +- internal/app/app_test.go | 42 +++++ internal/app/wire.go | 16 +- internal/cache/local.go | 19 +- internal/cache/local_test.go | 33 ++++ internal/cache/version_manager.go | 232 +++++++++++++++++-------- internal/cache/version_manager_test.go | 192 ++++++++++++++------ 9 files changed, 406 insertions(+), 135 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index d3e34545..6207f60a 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -29,9 +29,9 @@ One binary: Eighteen internal packages under `internal/` (plus `internal/testutil/` for shared test helpers): - **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers -- **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` and the `mq.max_bytes_gb` reconcile handing the MQ each served tenant's own gap window and byte budget, and `defaultSetting`/`onDefaultAdopt` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, and the same hook's `Hub.Prune` ends the open streams of a tenant no longer served), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it +- **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` and the `mq.max_bytes_gb` reconcile handing the MQ each served tenant's own gap window and byte budget, and `defaultSetting`/`onDefaultAdopt` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, the same hook's `Hub.Prune` ends the open streams of a tenant no longer served, and `LocalCache.Prune` drops its cache version index), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it - **`auth/`** — JWT auth middleware: HMAC **or** JWKS verification with `alg` pinned to the active verifier, role extraction from a configurable claim path; always runs, never rejects (bad token → empty role + stashed reason). One verifier per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 9): `Authenticator` keys them by `tenant.ID` — the request store's `settings.Store.Tenant()`, through an injected `TenantSource`; `tenant.Default` on the tenant-exempt routes — built from each tenant's `auth` block by `Reconfigure`, dropped by `Prune` once the tenant stops being served (rejected or removed), released by `Close`; the secrets (`Config`) are boot-level and shared. A JWKS key set is fetched off the boot and reload paths: until one has been stored the verifier is pending and a token-bearing request gets `503` + `Retry-After` from `api.refuseUnverifiable` (`auth.ErrVerifierPending`), never a `default_role` evaluation; refresh is library-managed (Eric, 2026-09-22), response capped at 1 MiB; the operator key's admin role is the request tenant's -- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, `..
.
.` for a namespace — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version folded into every namespace and query key of one tenant, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read and `Set` files the fill under it, so a write landing mid-query orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) +- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, a version per tenant, per (tenant, table) and per (tenant, table, scope), keyed by name and bumped in place (one entry per live namespace however often it is bumped, [#262](https://github.com/Wave-RF/WaveHouse/issues/262)) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` drops the tenant's index so its next key gets a process-unique generation, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read and `Set` files the fill under it, so a write landing mid-query orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) - **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability — boot is the validator, there is no dry run diff --git a/CHANGELOG.md b/CHANGELOG.md index 83f2fd83..50ea8f2e 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -79,6 +79,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed - **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/app/wire.go`, `docs/src/content/docs/{architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` and pipe execution rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL. The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. +- **The cache's version index no longer grows with every bump, and forgets a tenant no longer served** (`internal/cache/{local,version_manager}.go` (+ tests), `internal/app/wire.go` (+ tests), `docs/src/content/docs/architecture.md`, `AGENTS.md`): fixes [#262](https://github.com/Wave-RF/WaveHouse/issues/262) for the in-process cache, part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The index nested each table under its tenant's version and each scope under its table's, and never pruned, so every tenant invalidation left the tenant's whole index behind, and it grew with every tenant ever served. It now holds one version per tenant, per (tenant, table) and per (tenant, table, scope), bumped in place. A tenant invalidation drops the tenant's index and hands its next key a generation unique within the process, so nothing cached before it can match again, and a table bump drops the table's scope versions. After each settings reload the index of every tenant no longer served, removed or rejected, is dropped the same way; its cached results are orphaned with it, as they already were when such a tenant came back on a pool. No change to what is cached or served. The Redis-compatible backend (#613) bounds its versions with a TTL instead. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 3ba9a9eb..de4f6dfc 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -111,7 +111,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on, each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope: one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)). `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. -- **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: a namespace key is `..
.
.`, and a query key is folded with the tenant's version and each dependency's namespace key and namespace version, so bumping a table (a scopeless write) or one scope — scope is reserved and empty today, so every write is the whole-table bump — orphans every dependent entry without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results, and the flat directory's single tenant simply carries the `0` prefix. `BumpTenant` (behind `InvalidateTenant`) advances the tenant version that leads every namespace key of one tenant, orphaning its every namespace and every cached query in one step (the tenant version is folded into every query key, so a pipe result with no dependencies is orphaned too) — a table no bump ever keyed included, which an enumeration of the index would miss — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. +- **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: one version per tenant, per (tenant, table) and per (tenant, table, scope), each keyed by its name alone and bumped in place, so the index holds one entry per live tenant, table and scope however often each is bumped ([#262](https://github.com/Wave-RF/WaveHouse/issues/262)). A query key folds the tenant's version and, for each dependency, its tenant's, table's and scope's, so bumping a table (a scopeless write) orphans every scope of it, and bumping one scope orphans that scope and the whole-table view — scope is reserved and empty today, so every write is the whole-table bump — all without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results. A tenant's version is a *generation*, unique within the process and handed out by the first key built for the tenant; `BumpTenant` (behind `InvalidateTenant`) drops the tenant's whole index, so the next key gets a fresh generation no cached entry folds, orphaning every cached result of the tenant in one step — a pipe result with no dependencies, and a table no bump ever keyed, included — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). `LocalCache.Prune` does the same for every tenant no longer served, which `internal/app` runs after each settings reload, so a tenant removed or rejected stops holding its index. A table bump drops the table's scope versions with it, since every key they were folded into also folds the old table version; and a bump of a tenant with no index is a no-op, since no key folds its next generation yet. The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. ### `config/` — Configuration diff --git a/internal/app/app_test.go b/internal/app/app_test.go index 2ecd77d9..1f665b6a 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -16,6 +16,7 @@ import ( "os" "path/filepath" "strings" + "sync" "sync/atomic" "syscall" "testing" @@ -635,6 +636,47 @@ func TestReload_ReadmittedTenantCacheIsOrphaned(t *testing.T) { assert.Equal(t, []tenant.ID{"globex", "acme"}, mock.GetTenants(), "restored: the same") } +// pruneRecorder is a cache that records, at each Prune, which of the tenants +// it is asked about are still served. +type pruneRecorder struct { + testutil.MockCache + mu sync.Mutex + served []map[tenant.ID]bool +} + +func (p *pruneRecorder) Prune(served func(tenant.ID) bool) { + p.mu.Lock() + defer p.mu.Unlock() + p.served = append(p.served, map[tenant.ID]bool{"acme": served("acme"), "globex": served("globex")}) +} + +func (p *pruneRecorder) last() map[tenant.ID]bool { + p.mu.Lock() + defer p.mu.Unlock() + if len(p.served) == 0 { + return nil + } + return p.served[len(p.served)-1] +} + +// Every reload prunes the cache's version index down to the tenants served, +// so a tenant rejected or removed stops holding it (#262). +func TestReload_PrunesCacheIndexToServedTenants(t *testing.T) { + root := writeNestedSettings(t, map[string]map[string]any{"acme": nil, "globex": nil}) + a := newApp(t, testConfig(t, root), Options{}) + rec := &pruneRecorder{} + a.cache = rec + + rewriteSettings(t, filepath.Join(root, "globex"), invalidQuery) + a.tenants.Reload("test") + assert.Equal(t, map[tenant.ID]bool{"acme": true, "globex": false}, rec.last(), "rejected") + + rewriteSettings(t, filepath.Join(root, "globex"), nil) + require.NoError(t, os.RemoveAll(filepath.Join(root, "acme"))) + a.tenants.Reload("test") + assert.Equal(t, map[tenant.ID]bool{"acme": false, "globex": true}, rec.last(), "removed; the repaired one served again") +} + // keepalive is a config.json patch setting the stream block's keepalive pair. func keepalive(interval, buckets int) map[string]any { return map[string]any{"stream": map[string]any{"keepalive_interval": interval, "keepalive_buckets": buckets, "gap_window_minutes": 15}} diff --git a/internal/app/wire.go b/internal/app/wire.go index 4c9f1da2..8811ba63 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -591,7 +591,16 @@ func (a *App) wireMQ() error { return nil } -// wireCache opens the L1 cache — the only tier in standalone mode. +// pruner is a cache whose version index lives in the process and would +// otherwise keep a tenant that stopped being served (cache.LocalCache). +type pruner interface { + Prune(served func(tenant.ID) bool) +} + +// wireCache opens the L1 cache — the only tier in standalone mode. After +// every reload a tenant no longer served, removed or rejected alike, has its +// version index dropped (#262); its cache is orphaned with it, as it would +// be anyway when it came back (wireClickHouse). func (a *App) wireCache() error { l1, err := cache.NewLocal(a.cfg.Cache.L1MaxCost) if err != nil { @@ -600,6 +609,11 @@ func (a *App) wireCache() error { // TODO: eventually this is where we can switch between ristretto, redis, tiered (both), etc a.cache = l1 a.add(component{name: "cache", close: withoutContext(l1.Close)}) + a.tenants.AfterAdopt(func([]tenant.ID) { + if p, ok := a.cache.(pruner); ok { + p.Prune(a.served) + } + }) return nil } diff --git a/internal/cache/local.go b/internal/cache/local.go index 1a755e05..3e0e14b5 100644 --- a/internal/cache/local.go +++ b/internal/cache/local.go @@ -69,10 +69,11 @@ func (l *LocalCache) Set(_ context.Context, snap Snapshot, value []byte, ttl tim // view. Returns the number of namespaces processed. // // This bumps exactly what it's given. A whole-table bump already subsumes every -// per-scope bump for the same table (the table version is embedded in every -// namespace key), so a caller that knows a whole-table bump is coming should drop -// the now-redundant scope entries itself — the ingest worker does this as it -// builds the batch, where it already loops once and knows it's a single table. +// per-scope bump for the same table (every key that folds a scope version +// folds the table version too), so a caller that knows a whole-table bump is +// coming should drop the now-redundant scope entries itself — the ingest +// worker does this as it builds the batch, where it already loops once and +// knows it's a single table. func (l *LocalCache) Invalidate(_ context.Context, namespaces []Namespace) (uint64, error) { for _, ns := range namespaces { if ns.Scope == "" { @@ -85,13 +86,21 @@ func (l *LocalCache) Invalidate(_ context.Context, namespaces []Namespace) (uint } // InvalidateTenant orphans every cached result of tenant id, pipe results -// included: one version bump, nothing enumerated (see +// included: its version index is dropped, nothing enumerated (see // VersionManager.BumpTenant). func (l *LocalCache) InvalidateTenant(_ context.Context, id tenant.ID) error { l.versionManager.BumpTenant(id) return nil } +// Prune drops the version index of every tenant served rejects, orphaning +// its entries as InvalidateTenant would, so a tenant removed or rejected at +// a reload stops holding memory (#262). The entries themselves go with +// their TTL or Ristretto's eviction. +func (l *LocalCache) Prune(served func(tenant.ID) bool) { + l.versionManager.Prune(served) +} + // Wait blocks until all buffered writes have been applied. // Exposed for testing; production callers rarely need this. func (l *LocalCache) Wait() { diff --git a/internal/cache/local_test.go b/internal/cache/local_test.go index 2fdb3237..ad77a40d 100644 --- a/internal/cache/local_test.go +++ b/internal/cache/local_test.go @@ -2,10 +2,13 @@ package cache_test import ( "testing" + "time" + "github.com/stretchr/testify/assert" "github.com/stretchr/testify/require" "github.com/Wave-RF/WaveHouse/internal/cache" + "github.com/Wave-RF/WaveHouse/internal/tenant" "github.com/Wave-RF/WaveHouse/internal/testutil/cachetest" ) @@ -23,3 +26,33 @@ func TestLocalCache_Conformance(t *testing.T) { t.Parallel() cachetest.Run(t, newLocal, cachetest.Options{MaxValueBytes: localMaxCost}) } + +// A tenant that stops being served has its index dropped: what it cached is +// orphaned — it misses when served again — and a tenant still served keeps +// its entries. +func TestLocalCache_Prune(t *testing.T) { + t.Parallel() + c, err := cache.NewLocal(localMaxCost) + require.NoError(t, err) + t.Cleanup(func() { _ = c.Close() }) + ctx := t.Context() + fill := func(id tenant.ID) { + _, snap, err := c.Lookup(ctx, id, "q", []cache.Namespace{{Tenant: id, Table: "events"}}) + require.NoError(t, err) + require.NoError(t, c.Set(ctx, snap, []byte("rows"), time.Minute)) + } + get := func(id tenant.ID) []byte { + e, _, err := c.Lookup(ctx, id, "q", []cache.Namespace{{Tenant: id, Table: "events"}}) + require.NoError(t, err) + return e.Value + } + fill("acme") + fill("globex") + c.Wait() + require.NotNil(t, get("acme")) + require.NotNil(t, get("globex")) + + c.Prune(func(id tenant.ID) bool { return id == "acme" }) + assert.NotNil(t, get("acme"), "still served") + assert.Nil(t, get("globex"), "pruned: orphaned, never revived") +} diff --git a/internal/cache/version_manager.go b/internal/cache/version_manager.go index d97bc70a..dfe18b8c 100644 --- a/internal/cache/version_manager.go +++ b/internal/cache/version_manager.go @@ -9,27 +9,44 @@ import ( "github.com/Wave-RF/WaveHouse/internal/tenant" ) -// VersionManager handles the safe tracking of table + scope versioning. -// It uses a standard map because versions must NEVER be evicted under memory pressure. -// TODO: this potentially could be bad/dangerous with a low amount of RAM available/high memory pressure AND a TON of tables/scopes per table... will need to work out eventually +// VersionManager is the invalidation index: one version per tenant, per +// (tenant, table) and per (tenant, table, scope), in maps keyed by name +// alone, never by another version (#262). A bump overwrites a version in +// place, so the index holds one entry per live tenant, table and scope +// however often each is bumped, and forgetting a tenant releases all of it. +// +// A query key folds all three versions of each dependency, which gives the +// lattice: a table bump orphans every scope, a scope bump that scope and the +// whole-table view, and a tenant bump everything of the tenant's. type VersionManager struct { mu sync.RWMutex - // tenantVersions leads every key of a tenant, so BumpTenant orphans the - // tenant's every namespace and query in one step — the ones no bump ever - // keyed included, which is what an enumeration of the maps would miss. - tenantVersions map[tenant.ID]uint64 // -> tenant_version - tableVersions map[string]uint64 // ..
-> table_version - namespaceVersions map[string]uint64 // ..
.. -> namespace_version + // tenants holds each tenant's index from the first query key built for + // it until the tenant is bumped or pruned. + tenants map[tenant.ID]*tenantVersions + + // lastGen is the last generation handed to a tenant; see tenantVersions.gen. + lastGen uint64 } -// NewVersionManager initializes the thread-safe version store. -func NewVersionManager() *VersionManager { - return &VersionManager{ - tenantVersions: make(map[tenant.ID]uint64), - tableVersions: make(map[string]uint64), - namespaceVersions: make(map[string]uint64), - } +// tenantVersions is one tenant's slice of the index. +type tenantVersions struct { + // gen is the tenant's version: unique within the process, so a tenant + // forgotten and recreated can never fold a generation an entry was + // cached under. That is what makes dropping the tenant's whole index a + // safe bump. + gen uint64 + tables map[string]*tableVersions +} + +// tableVersions is one table's version and its scopes'. A missing table or +// scope reads as 0: an entry is only ever removed together with a bump of +// the version above it (a table bump clears the scopes, a tenant bump +// drops the tables), so a 0 read after a removal never matches an entry +// cached before it. +type tableVersions struct { + version uint64 + scopes map[string]uint64 } // Namespace is one (tenant, table, scope) a cached result depends on. The @@ -41,76 +58,101 @@ type Namespace struct { Scope string } -// tableKeyLocked renders the table-versions key, -// "..
"; caller must hold vm.mu. A tenant id -// cannot contain a dot and callers encode the table dot-free, so the tokens -// can never run together. -func (vm *VersionManager) tableKeyLocked(id tenant.ID, table string) string { - return fmt.Sprintf("%s.%d.%s", id, vm.tenantVersions[id], table) -} - -// namespaceKeyLocked builds the namespace-table key; caller must hold vm.mu. -func (vm *VersionManager) namespaceKeyLocked(ns Namespace) string { - tk := vm.tableKeyLocked(ns.Tenant, ns.Table) - return fmt.Sprintf("%s.%d.%s", tk, vm.tableVersions[tk], ns.Scope) -} - -// NamespaceKey renders the namespace-table key for ns at its tenant's and -// table's current versions: -// "..
.." (scopeless -// scope is "", so e.g. ".0.
.."). -func (vm *VersionManager) NamespaceKey(ns Namespace) string { - vm.mu.RLock() - defer vm.mu.RUnlock() - return vm.namespaceKeyLocked(ns) +// NewVersionManager initializes the thread-safe version store. +func NewVersionManager() *VersionManager { + return &VersionManager{tenants: make(map[tenant.ID]*tenantVersions)} } // QueryKey builds the queries-table key for tenant id's result that depends // on deps: the query's sha (hash of SQL+params) folded with the tenant's -// version and every dependency's namespace key AND its namespace version, so -// a bump of the tenant or of any dependency misses the key — a result with no -// deps (a pipe) is orphaned by BumpTenant too. A structured query passes one -// Namespace; a pipe passes several. Deps are sorted so their order never -// changes the key. +// version and, for every dependency, its tenant's, table's and scope's +// versions, so a bump of the tenant or of any dependency misses the key — a +// result with no deps (a pipe) is orphaned by BumpTenant too. A structured +// query passes one Namespace, a pipe none yet (#343). Deps are sorted so +// their order never changes the key. Every version is read under one lock, +// so the key is one consistent snapshot. +// +// The first key built for a tenant creates its index at a fresh generation. func (vm *VersionManager) QueryKey(id tenant.ID, sha string, deps []Namespace) string { + vm.mu.RLock() + key, ok := vm.queryKeyLocked(id, sha, deps, false) + vm.mu.RUnlock() + if ok { + return key + } + vm.mu.Lock() + defer vm.mu.Unlock() + key, _ = vm.queryKeyLocked(id, sha, deps, true) + return key +} + +// queryKeyLocked renders QueryKey with vm.mu held — for writing when create +// is set, which creates the index of each tenant the key names that has +// none; otherwise such a tenant reports !ok. +func (vm *VersionManager) queryKeyLocked(id tenant.ID, sha string, deps []Namespace, create bool) (string, bool) { + index := func(id tenant.ID) (*tenantVersions, bool) { + tv := vm.tenants[id] + if tv == nil && create { + tv = vm.newTenantLocked(id) + } + return tv, tv != nil + } + own, ok := index(id) + if !ok { + return "", false + } segs := make([]string, len(deps)) - // Lock per dependency rather than across the whole loop: each dep's table + - // namespace versions are read together (consistent for that dep), but we don't - // hold the lock across all deps. A concurrent bump can land between deps; the - // caller files its fill under this key (a Snapshot), so a bump that lands - // anywhere after the read of a version orphans it. The sort/join run with - // no lock held. for i, d := range deps { - vm.mu.RLock() - nsKey := vm.namespaceKeyLocked(d) - segs[i] = fmt.Sprintf("%s.%d", nsKey, vm.namespaceVersions[nsKey]) - vm.mu.RUnlock() + tv, ok := index(d.Tenant) + if !ok { + return "", false + } + var table, scope uint64 + if t := tv.tables[d.Table]; t != nil { + table, scope = t.version, t.scopes[d.Scope] + } + segs[i] = fmt.Sprintf("%s.%d.%s.%d.%s.%d", d.Tenant, tv.gen, d.Table, table, d.Scope, scope) } - vm.mu.RLock() - tv := vm.tenantVersions[id] - vm.mu.RUnlock() sort.Strings(segs) - return fmt.Sprintf("%s|%s.%d|%s", sha, id, tv, strings.Join(segs, "|")) + return fmt.Sprintf("%s|%s.%d|%s", sha, id, own.gen, strings.Join(segs, "|")), true +} + +func (vm *VersionManager) newTenantLocked(id tenant.ID) *tenantVersions { + vm.lastGen++ + tv := &tenantVersions{gen: vm.lastGen, tables: make(map[string]*tableVersions)} + vm.tenants[id] = tv + return tv +} + +// tableLocked is the entry for a tenant's table, created at version 0, or +// nil when the tenant has no index: no key folds its current generation +// yet, so there is nothing a bump could orphan. Caller holds vm.mu for +// writing. +func (vm *VersionManager) tableLocked(id tenant.ID, table string) *tableVersions { + tv := vm.tenants[id] + if tv == nil { + return nil + } + t := tv.tables[table] + if t == nil { + t = &tableVersions{} + tv.tables[table] = t + } + return t } // BumpTable advances a tenant's table version, orphaning every namespace — and // every cached query — that depends on the table, in one step (the whole-table -// nuke). The same table under another tenant is untouched. +// nuke). The table's scope versions are dropped with it: every key they were +// folded into also folds the old table version. The same table under another +// tenant is untouched. func (vm *VersionManager) BumpTable(id tenant.ID, table string) { vm.mu.Lock() defer vm.mu.Unlock() - vm.tableVersions[vm.tableKeyLocked(id, table)]++ -} - -// BumpTenant advances a tenant's version, orphaning its every namespace — -// and every cached query, whatever its deps — in one step (the whole-tenant -// nuke): every namespace and query key of the tenant carries the version, so -// nothing has to be enumerated, and a table no bump ever keyed is orphaned -// like the rest. Other tenants are untouched. -func (vm *VersionManager) BumpTenant(id tenant.ID) { - vm.mu.Lock() - defer vm.mu.Unlock() - vm.tenantVersions[id]++ + if t := vm.tableLocked(id, table); t != nil { + t.version++ + t.scopes = nil + } } // BumpNamespace advances one (tenant, table, scope) namespace plus the table's @@ -119,8 +161,54 @@ func (vm *VersionManager) BumpTenant(id tenant.ID) { func (vm *VersionManager) BumpNamespace(ns Namespace) { vm.mu.Lock() defer vm.mu.Unlock() - vm.namespaceVersions[vm.namespaceKeyLocked(ns)]++ + t := vm.tableLocked(ns.Tenant, ns.Table) + if t == nil { + return + } + if t.scopes == nil { + t.scopes = make(map[string]uint64) + } + t.scopes[ns.Scope]++ if ns.Scope != "" { - vm.namespaceVersions[vm.namespaceKeyLocked(Namespace{Tenant: ns.Tenant, Table: ns.Table})]++ + t.scopes[""]++ + } +} + +// BumpTenant orphans every cached query of a tenant, whatever its deps, in +// one step (the whole-tenant nuke), by dropping the tenant's index: the next +// key built for it gets a fresh generation, which no cached entry folds. +// Nothing has to be enumerated, a table no bump ever keyed is orphaned like +// the rest, and the index the tenant held is released. Other tenants are +// untouched. +func (vm *VersionManager) BumpTenant(id tenant.ID) { + vm.mu.Lock() + defer vm.mu.Unlock() + delete(vm.tenants, id) +} + +// Prune drops the index of every tenant keep rejects, as BumpTenant would, +// so a tenant that stops being served stops holding memory; one served again +// starts over at a fresh generation. +func (vm *VersionManager) Prune(keep func(tenant.ID) bool) { + vm.mu.Lock() + defer vm.mu.Unlock() + for id := range vm.tenants { + if !keep(id) { + delete(vm.tenants, id) + } + } +} + +// size is the number of versions the index holds, for tests. +func (vm *VersionManager) size() int { + vm.mu.RLock() + defer vm.mu.RUnlock() + n := len(vm.tenants) + for _, tv := range vm.tenants { + n += len(tv.tables) + for _, t := range tv.tables { + n += len(t.scopes) + } } + return n } diff --git a/internal/cache/version_manager_test.go b/internal/cache/version_manager_test.go index e1a1c330..c9278577 100644 --- a/internal/cache/version_manager_test.go +++ b/internal/cache/version_manager_test.go @@ -8,45 +8,17 @@ import ( "github.com/Wave-RF/WaveHouse/internal/tenant" ) -func TestVersionManager_NamespaceKey(t *testing.T) { - t.Parallel() - vm := NewVersionManager() - - // The tenant leads at its default version (0), then the table at its - // default version (0); a scopeless namespace renders a trailing dot. The - // flat directory's tenant is "0". - assert.Equal(t, "acme.0.users.0.", vm.NamespaceKey(Namespace{Tenant: "acme", Table: "users"})) - assert.Equal(t, "acme.0.users.0.org_1", vm.NamespaceKey(Namespace{Tenant: "acme", Table: "users", Scope: "org_1"})) - assert.Equal(t, "0.0.users.0.", vm.NamespaceKey(Namespace{Tenant: tenant.Default, Table: "users"})) - - // The table version is embedded in every namespace key for that tenant's - // table, so a BumpTable is reflected across all its scopes at once — and - // nowhere else: the same table under another tenant keeps its version. - vm.BumpTable("acme", "users") - assert.Equal(t, "acme.0.users.1.", vm.NamespaceKey(Namespace{Tenant: "acme", Table: "users"})) - assert.Equal(t, "acme.0.users.1.org_1", vm.NamespaceKey(Namespace{Tenant: "acme", Table: "users", Scope: "org_1"})) - assert.Equal(t, "globex.0.users.0.", vm.NamespaceKey(Namespace{Tenant: "globex", Table: "users"})) - - // The tenant version leads every key of the tenant, so a BumpTenant moves - // every table of acme's — the never-bumped orders table included — to a - // fresh key space, at table version 0 again, and no other tenant's. - vm.BumpTenant("acme") - assert.Equal(t, "acme.1.users.0.", vm.NamespaceKey(Namespace{Tenant: "acme", Table: "users"})) - assert.Equal(t, "acme.1.orders.0.", vm.NamespaceKey(Namespace{Tenant: "acme", Table: "orders"})) - assert.Equal(t, "globex.0.users.0.", vm.NamespaceKey(Namespace{Tenant: "globex", Table: "users"})) -} - func TestVersionManager_QueryKey(t *testing.T) { t.Parallel() vm := NewVersionManager() - // One dependency at default versions: - // sha | . | ..
.... + // sha | . | ..
...; + // acme's index is created by its first key, at generation 1. key := vm.QueryKey("acme", "hash123", []Namespace{{Tenant: "acme", Table: "users", Scope: "org_1"}}) - assert.Equal(t, "hash123|acme.0|acme.0.users.0.org_1.0", key) + assert.Equal(t, "hash123|acme.1|acme.1.users.0.org_1.0", key) // No deps (a pipe) still folds the tenant version. - assert.Equal(t, "hash123|acme.0|", vm.QueryKey("acme", "hash123", nil)) + assert.Equal(t, "hash123|acme.1|", vm.QueryKey("acme", "hash123", nil)) // Dependency order must not change the key (segments are sorted). deps1 := []Namespace{{Tenant: "acme", Table: "a"}, {Tenant: "acme", Table: "b"}} @@ -59,6 +31,9 @@ func TestVersionManager_QueryKey(t *testing.T) { vm.QueryKey("acme", "h", []Namespace{{Tenant: "acme", Table: "users"}}), vm.QueryKey("globex", "h", []Namespace{{Tenant: "globex", Table: "users"}})) assert.NotEqual(t, vm.QueryKey("acme", "h", nil), vm.QueryKey("globex", "h", nil)) + + // Reading keys is stable: nothing but a bump moves a version. + assert.Equal(t, key, vm.QueryKey("acme", "hash123", []Namespace{{Tenant: "acme", Table: "users", Scope: "org_1"}})) } func TestVersionManager_BumpTable(t *testing.T) { @@ -69,16 +44,16 @@ func TestVersionManager_BumpTable(t *testing.T) { orders := []Namespace{{Tenant: "acme", Table: "orders", Scope: "org_1"}} globexUsers := []Namespace{{Tenant: "globex", Table: "users", Scope: "org_1"}} - usersBefore := vm.QueryKey(users[0].Tenant, "h", users) - ordersBefore := vm.QueryKey(orders[0].Tenant, "h", orders) - globexBefore := vm.QueryKey(globexUsers[0].Tenant, "h", globexUsers) + usersBefore := vm.QueryKey("acme", "h", users) + ordersBefore := vm.QueryKey("acme", "h", orders) + globexBefore := vm.QueryKey("globex", "h", globexUsers) // Bumping a table changes the key for that tenant's table but leaves other // tables — and the same table under another tenant — alone. vm.BumpTable("acme", "users") - assert.NotEqual(t, usersBefore, vm.QueryKey(users[0].Tenant, "h", users)) - assert.Equal(t, ordersBefore, vm.QueryKey(orders[0].Tenant, "h", orders)) - assert.Equal(t, globexBefore, vm.QueryKey(globexUsers[0].Tenant, "h", globexUsers)) + assert.NotEqual(t, usersBefore, vm.QueryKey("acme", "h", users)) + assert.Equal(t, ordersBefore, vm.QueryKey("acme", "h", orders)) + assert.Equal(t, globexBefore, vm.QueryKey("globex", "h", globexUsers)) } func TestVersionManager_BumpNamespace(t *testing.T) { @@ -90,24 +65,44 @@ func TestVersionManager_BumpNamespace(t *testing.T) { otherScope := []Namespace{{Tenant: "acme", Table: "users", Scope: "org_2"}} otherTenant := []Namespace{{Tenant: "globex", Table: "users", Scope: "org_1"}} - scopedBefore := vm.QueryKey(scoped[0].Tenant, "h", scoped) - wholeBefore := vm.QueryKey(wholeTable[0].Tenant, "h", wholeTable) - otherBefore := vm.QueryKey(otherScope[0].Tenant, "h", otherScope) - otherTenantBefore := vm.QueryKey(otherTenant[0].Tenant, "h", otherTenant) + scopedBefore := vm.QueryKey("acme", "h", scoped) + wholeBefore := vm.QueryKey("acme", "h", wholeTable) + otherBefore := vm.QueryKey("acme", "h", otherScope) + otherTenantBefore := vm.QueryKey("globex", "h", otherTenant) // Bumping (acme, users, org_1) changes that scope AND the whole-table view, // but leaves every other scope — and the same scope under another tenant — // valid. vm.BumpNamespace(Namespace{Tenant: "acme", Table: "users", Scope: "org_1"}) - assert.NotEqual(t, scopedBefore, vm.QueryKey(scoped[0].Tenant, "h", scoped)) - assert.NotEqual(t, wholeBefore, vm.QueryKey(wholeTable[0].Tenant, "h", wholeTable)) - assert.Equal(t, otherBefore, vm.QueryKey(otherScope[0].Tenant, "h", otherScope)) - assert.Equal(t, otherTenantBefore, vm.QueryKey(otherTenant[0].Tenant, "h", otherTenant)) + assert.NotEqual(t, scopedBefore, vm.QueryKey("acme", "h", scoped)) + assert.NotEqual(t, wholeBefore, vm.QueryKey("acme", "h", wholeTable)) + assert.Equal(t, otherBefore, vm.QueryKey("acme", "h", otherScope)) + assert.Equal(t, otherTenantBefore, vm.QueryKey("globex", "h", otherTenant)) +} + +// A table bump drops the table's scope versions, which then read as 0 again +// — safe only because every key a scope version was folded into also folds +// the table version the bump moved. Pinned so a table bump that forgot to +// advance the table version would revive the scoped entry. +func TestVersionManager_BumpTableDropsScopes(t *testing.T) { + t.Parallel() + vm := NewVersionManager() + scoped := []Namespace{{Tenant: "acme", Table: "users", Scope: "org_1"}} + + fresh := vm.QueryKey("acme", "h", scoped) + vm.BumpNamespace(scoped[0]) + bumped := vm.QueryKey("acme", "h", scoped) + vm.BumpTable("acme", "users") + after := vm.QueryKey("acme", "h", scoped) + + assert.NotEqual(t, fresh, after) + assert.NotEqual(t, bumped, after) + assert.Equal(t, 2, vm.size(), "the tenant and its table; the scopes went with the table bump") } // TestVersionManager_BumpTenant: a tenant's every namespace is orphaned in -// one step — a table that was never bumped (so has no key of its own to bump) -// included — and no other tenant's is touched. +// one step — a table that was never bumped included — and no other tenant's +// is touched. func TestVersionManager_BumpTenant(t *testing.T) { t.Parallel() vm := NewVersionManager() @@ -115,18 +110,107 @@ func TestVersionManager_BumpTenant(t *testing.T) { users := []Namespace{{Tenant: "acme", Table: "users", Scope: "org_1"}} orders := []Namespace{{Tenant: "acme", Table: "orders"}} globexUsers := []Namespace{{Tenant: "globex", Table: "users", Scope: "org_1"}} + vm.QueryKey("acme", "h", users) vm.BumpTable("acme", "users") - usersBefore := vm.QueryKey(users[0].Tenant, "h", users) - ordersBefore := vm.QueryKey(orders[0].Tenant, "h", orders) - globexBefore := vm.QueryKey(globexUsers[0].Tenant, "h", globexUsers) + usersBefore := vm.QueryKey("acme", "h", users) + ordersBefore := vm.QueryKey("acme", "h", orders) + globexBefore := vm.QueryKey("globex", "h", globexUsers) vm.BumpTenant("acme") - assert.NotEqual(t, usersBefore, vm.QueryKey(users[0].Tenant, "h", users)) - assert.NotEqual(t, ordersBefore, vm.QueryKey(orders[0].Tenant, "h", orders), "a table no bump ever keyed is orphaned too") - assert.Equal(t, globexBefore, vm.QueryKey(globexUsers[0].Tenant, "h", globexUsers)) + assert.NotEqual(t, usersBefore, vm.QueryKey("acme", "h", users)) + assert.NotEqual(t, ordersBefore, vm.QueryKey("acme", "h", orders), "a table no bump ever keyed is orphaned too") + assert.Equal(t, globexBefore, vm.QueryKey("globex", "h", globexUsers)) pipeBefore := vm.QueryKey("acme", "h", nil) vm.BumpTenant("acme") assert.NotEqual(t, pipeBefore, vm.QueryKey("acme", "h", nil), "a result with no deps is orphaned too") } + +// Dropping a tenant's index is a bump only because the index it gets back +// never repeats a generation: every key built before any of these drops must +// differ from every key built after it. A counter per tenant restarting at 0 +// fails this, reviving the first entry. +func TestVersionManager_GenerationsNeverRepeat(t *testing.T) { + t.Parallel() + vm := NewVersionManager() + deps := []Namespace{{Tenant: "acme", Table: "users"}} + seen := map[string]bool{} + for i := range 100 { + key := vm.QueryKey("acme", "h", deps) + assert.False(t, seen[key], "round %d revived %s", i, key) + seen[key] = true + if i%2 == 0 { + vm.BumpTenant("acme") + } else { + vm.Prune(func(tenant.ID) bool { return false }) + } + } +} + +// A bump of a tenant with no index is a no-op: no key folds its next +// generation yet, so nothing needs orphaning — and an insert still in flight +// for a tenant just pruned does not bring its index back. +func TestVersionManager_BumpWithoutIndex(t *testing.T) { + t.Parallel() + vm := NewVersionManager() + vm.BumpTable("acme", "users") + vm.BumpNamespace(Namespace{Tenant: "acme", Table: "users", Scope: "org_1"}) + vm.BumpTenant("acme") + assert.Zero(t, vm.size()) +} + +func TestVersionManager_Prune(t *testing.T) { + t.Parallel() + vm := NewVersionManager() + acme := []Namespace{{Tenant: "acme", Table: "users"}} + globex := []Namespace{{Tenant: "globex", Table: "users"}} + acmeBefore := vm.QueryKey("acme", "h", acme) + globexBefore := vm.QueryKey("globex", "h", globex) + vm.BumpTable("acme", "users") + vm.BumpTable("globex", "users") + acmeBumped := vm.QueryKey("acme", "h", acme) + globexBumped := vm.QueryKey("globex", "h", globex) + + vm.Prune(func(id tenant.ID) bool { return id == "globex" }) + assert.Equal(t, 2, vm.size(), "globex and its table; acme released") + assert.Equal(t, globexBumped, vm.QueryKey("globex", "h", globex), "a kept tenant is untouched") + + back := vm.QueryKey("acme", "h", acme) + assert.NotEqual(t, acmeBefore, back, "a pruned tenant never revives what it cached") + assert.NotEqual(t, acmeBumped, back) + assert.NotEqual(t, globexBefore, globexBumped) +} + +// The index holds one version per live tenant, table and scope, however often +// each is bumped (#262): the nested index this replaced kept every table and +// scope under every tenant version it had seen. +func TestVersionManager_SizeDoesNotGrowWithBumps(t *testing.T) { + t.Parallel() + vm := NewVersionManager() + touch := func() { + for _, id := range []tenant.ID{"acme", "globex"} { + for _, table := range []string{"users", "orders"} { + for _, scope := range []string{"", "org_1", "org_2"} { + vm.QueryKey(id, "h", []Namespace{{Tenant: id, Table: table, Scope: scope}}) + vm.BumpNamespace(Namespace{Tenant: id, Table: table, Scope: scope}) + } + } + } + } + touch() + settled := vm.size() + assert.Equal(t, 2+2*2+2*2*3, settled, "two tenants, two tables each, three scopes each") + + for i := range 10_000 { + switch i % 3 { + case 0: + vm.BumpTable("acme", []string{"users", "orders"}[i%2]) + case 1: + vm.BumpTenant("globex") + } + touch() + assert.LessOrEqual(t, vm.size(), settled) + } + assert.Equal(t, settled, vm.size()) +} From ed8022bf34127e5cdd751d916f7d7dde162971fc Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 00:06:14 -0400 Subject: [PATCH 05/27] docs(cache): name the cache prune hook; pin LocalCache as a pruner Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- AGENTS.md | 2 +- docs/src/content/docs/architecture.md | 2 +- internal/app/wire.go | 9 +++++++-- 3 files changed, 9 insertions(+), 4 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 6207f60a..adbec1d8 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -29,7 +29,7 @@ One binary: Eighteen internal packages under `internal/` (plus `internal/testutil/` for shared test helpers): - **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers -- **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` and the `mq.max_bytes_gb` reconcile handing the MQ each served tenant's own gap window and byte budget, and `defaultSetting`/`onDefaultAdopt` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, the same hook's `Hub.Prune` ends the open streams of a tenant no longer served, and `LocalCache.Prune` drops its cache version index), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it +- **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` and the `mq.max_bytes_gb` reconcile handing the MQ each served tenant's own gap window and byte budget, and `defaultSetting`/`onDefaultAdopt` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, the same hook's `Hub.Prune` ends the open streams of a tenant no longer served, and `wireCache`'s hook drops, through `LocalCache.Prune`, the cache version index of a tenant no longer served ([#262](https://github.com/Wave-RF/WaveHouse/issues/262))), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it - **`auth/`** — JWT auth middleware: HMAC **or** JWKS verification with `alg` pinned to the active verifier, role extraction from a configurable claim path; always runs, never rejects (bad token → empty role + stashed reason). One verifier per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 9): `Authenticator` keys them by `tenant.ID` — the request store's `settings.Store.Tenant()`, through an injected `TenantSource`; `tenant.Default` on the tenant-exempt routes — built from each tenant's `auth` block by `Reconfigure`, dropped by `Prune` once the tenant stops being served (rejected or removed), released by `Close`; the secrets (`Config`) are boot-level and shared. A JWKS key set is fetched off the boot and reload paths: until one has been stored the verifier is pending and a token-bearing request gets `503` + `Retry-After` from `api.refuseUnverifiable` (`auth.ErrVerifierPending`), never a `default_role` evaluation; refresh is library-managed (Eric, 2026-09-22), response capped at 1 MiB; the operator key's admin role is the request tenant's - **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, a version per tenant, per (tenant, table) and per (tenant, table, scope), keyed by name and bumped in place (one entry per live namespace however often it is bumped, [#262](https://github.com/Wave-RF/WaveHouse/issues/262)) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` drops the tenant's index so its next key gets a process-unique generation, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read and `Set` files the fill under it, so a write landing mid-query orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index de4f6dfc..368c9368 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -90,7 +90,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability, ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the whole cache — structured-query and pipe results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe` builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ` hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the whole cache — structured-query and pipe results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth, dedupe and cache hooks use too), ending the open streams of a tenant no longer served, and `wireCache`'s hook prunes the cache's version index the same way (`LocalCache.Prune`, [#262](https://github.com/Wave-RF/WaveHouse/issues/262)), so a tenant no longer served stops holding it. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe` builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ` hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out diff --git a/internal/app/wire.go b/internal/app/wire.go index 8811ba63..900c1999 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -143,8 +143,9 @@ func gapWindows(tenants *settings.Registry) map[tenant.ID]time.Duration { } // served reports whether the registry is serving tenant id: what the -// per-tenant resources — verifiers, dedupe stores, open streams — are pruned -// by once a reload removes or rejects their tenant. +// per-tenant resources — verifiers, dedupe stores, open streams, the cache +// version index — are pruned by once a reload removes or rejects their +// tenant. func (a *App) served(id tenant.ID) bool { _, ok := a.tenants.For(id) return ok @@ -597,6 +598,10 @@ type pruner interface { Prune(served func(tenant.ID) bool) } +// The hook below asserts pruner at run time; this keeps LocalCache from +// silently dropping out of it. +var _ pruner = (*cache.LocalCache)(nil) + // wireCache opens the L1 cache — the only tier in standalone mode. After // every reload a tenant no longer served, removed or rejected alike, has its // version index dropped (#262); its cache is orphaned with it, as it would From c06808400e5a5eb24fdf65b26ac3531d31ed9884 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 02:53:43 -0400 Subject: [PATCH 06/27] feat(config): cache.backend=redis selects the shared cache The boot config's cache.backend now takes redis, configured by a new cache.redis block (WH_CACHE_REDIS_*), and wireCache builds a RedisCache from it, reading the TLS files. A malformed block refuses boot; an unreachable server or a rejected password boots bypassed and keeps reconnecting. An integration test boots two instances over one Redis and one ClickHouse: an ingest through one is served fresh by the other inside the stale entry's TTL, and a paused Redis leaves queries succeeding. The e2e suite now runs against Redis, so its coverage exclude is gone. Part of #613 (E4). Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- .testcoverage.yml | 6 - AGENTS.md | 4 +- CHANGELOG.md | 3 +- README.md | 2 +- config.yaml | 13 +- deployments/compose/dependencies.yaml | 13 + docs/src/content/docs/api.md | 2 +- docs/src/content/docs/architecture.md | 13 +- docs/src/content/docs/configuration.mdx | 78 ++++- docs/src/content/docs/deployment.md | 20 ++ docs/src/content/docs/development.md | 4 +- docs/src/content/docs/getting-started.md | 2 +- docs/src/content/docs/pipes.mdx | 2 +- docs/src/content/docs/settings-directory.mdx | 2 +- internal/app/app_test.go | 71 +++++ internal/app/wire.go | 41 ++- internal/config/backends.go | 40 ++- internal/config/backends_test.go | 2 +- internal/config/cache_redis.go | 153 ++++++++++ internal/config/cache_redis_test.go | 298 +++++++++++++++++++ scripts/orchestrator/main.go | 51 +++- tests/e2e/fixtures/config.yaml | 16 +- tests/integration/shared_cache_test.go | 271 +++++++++++++++++ 23 files changed, 1056 insertions(+), 51 deletions(-) create mode 100644 internal/config/cache_redis.go create mode 100644 internal/config/cache_redis_test.go create mode 100644 tests/integration/shared_cache_test.go diff --git a/.testcoverage.yml b/.testcoverage.yml index def851f0..aff1a694 100644 --- a/.testcoverage.yml +++ b/.testcoverage.yml @@ -73,9 +73,3 @@ exclude: - ^internal/settings/ - ^cmd/wavehouse/validate\.go$ - ^cmd/wavehouse/bootstrap\.go$ - # The Redis-compatible cache backend: the e2e stack runs LocalCache - # (no Redis server, and no config selects the backend yet — #613 E4), - # so the binary carries these files but e2e can never reach them; they - # pulled the e2e gate to 56.4%. The integration suite runs them against - # real servers, and the merged total still counts them. - - ^internal/cache/(redis|redis_codec|breaker|pending|metrics)\.go$ diff --git a/AGENTS.md b/AGENTS.md index ebdd877c..68c8f8df 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -31,10 +31,10 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers - **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` and the `mq.max_bytes_gb` reconcile handing the MQ each served tenant's own gap window and byte budget, and `defaultSetting`/`onDefaultAdopt` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, the same hook's `Hub.Prune` ends the open streams of a tenant no longer served, and `wireCache`'s hook drops, through `LocalCache.Prune`, the cache version index of a tenant no longer served ([#262](https://github.com/Wave-RF/WaveHouse/issues/262))), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it - **`auth/`** — JWT auth middleware: HMAC **or** JWKS verification with `alg` pinned to the active verifier, role extraction from a configurable claim path; always runs, never rejects (bad token → empty role + stashed reason). One verifier per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 9): `Authenticator` keys them by `tenant.ID` — the request store's `settings.Store.Tenant()`, through an injected `TenantSource`; `tenant.Default` on the tenant-exempt routes — built from each tenant's `auth` block by `Reconfigure`, dropped by `Prune` once the tenant stops being served (rejected or removed), released by `Close`; the secrets (`Config`) are boot-level and shared. A JWKS key set is fetched off the boot and reload paths: until one has been stored the verifier is pending and a token-bearing request gets `503` + `Retry-After` from `api.refuseUnverifiable` (`auth.ErrVerifierPending`), never a `default_role` evaluation; refresh is library-managed (Eric, 2026-09-22), response capped at 1 MiB; the operator key's admin role is the request tenant's -- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index), and `RedisCache`, the Redis-compatible shared backend (random version tokens under the tenant's hash tag, one-round-trip lookups, bypass on failure behind a circuit breaker, deferred invalidations retried; built and tested, not yet selectable by config — [#613](https://github.com/Wave-RF/WaveHouse/issues/613) E4). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, a version per tenant, per (tenant, table) and per (tenant, table, scope), keyed by name and bumped in place (one entry per live namespace however often it is bumped, [#262](https://github.com/Wave-RF/WaveHouse/issues/262)) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` drops the tenant's index so its next key gets a process-unique generation, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read and `Set` files the fill under it, so a write landing mid-query orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) +- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index), and `RedisCache`, the Redis-compatible shared backend (random version tokens under the tenant's hash tag, one-round-trip lookups, bypass on failure behind a circuit breaker, deferred invalidations retried; selected by `cache.backend: redis`, configured by the boot config's `cache.redis` block — [#613](https://github.com/Wave-RF/WaveHouse/issues/613)). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, a version per tenant, per (tenant, table) and per (tenant, table, scope), keyed by name and bumped in place (one entry per live namespace however often it is bumped, [#262](https://github.com/Wave-RF/WaveHouse/issues/262)) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` drops the tenant's index so its next key gets a process-unique generation, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read and `Set` files the fill under it, so a write landing mid-query orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) -- **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability when a selected backend keeps state there (`NeedsDataDir`); `backends.go` holds each layer's `.backend` (only the in-process value today; `coord.backend` reserved) — boot is the validator, there is no dry run +- **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability when a selected backend keeps state there (`NeedsDataDir`); `backends.go` holds each layer's `.backend` (the in-process value by default; `cache.backend` also takes `redis`, whose sub-block is `cache_redis.go`; `coord.backend` reserved) — boot is the validator, there is no dry run - **`dedupe/`** — `Deduplicator` interface → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) diff --git a/CHANGELOG.md b/CHANGELOG.md index 5bbe3053..ab628fdf 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,8 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A Redis-compatible shared cache backend, built but not yet selectable** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row open a circuit breaker that skips the server until a probe succeeds, and deferred invalidations are retried until they land, collapsing to one tenant-wide bump per tenant past 100,000 keys. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. Nothing selects it yet: the `cache.backend` switch and the `cache.redis.*` settings arrive with the wiring (E4), so every deployment still runs `LocalCache`. Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression and server-stops-answering cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. +- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone`, `cluster`, `sentinel`), `sentinel_master`, `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `-1` never compresses, since the loader reads a `0` in the file as unset) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. A malformed block — no address, an address without a port, an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. +- **A Redis-compatible shared cache backend** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row open a circuit breaker that skips the server until a probe succeeds, and deferred invalidations are retried until they land, collapsing to one tenant-wide bump per tenant past 100,000 keys. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. `cache.backend: redis` selects it (the entry above). Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression and server-stops-answering cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block; no backend has settings yet, so every such sub-block is an unknown key for now and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name the `mq`, `cache` and `dedupe` backends: the zero value is not the default, and `app.New` refuses it. `coord.backend` is reserved: nothing reads it until the lease layer lands, and the sweeper still runs in every process. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. diff --git a/README.md b/README.md index fce01a0a..65f6da14 100644 --- a/README.md +++ b/README.md @@ -72,7 +72,7 @@ ClickHouse is a phenomenal OLAP database, but pointing a frontend right at it le If you're building user-facing analytics, WaveHouse is like **Supabase for ClickHouse**. Or an **open-source Tinybird** that pushes data to the frontend in real time over SSE, not just pull-based REST. - **Ingest** — async durable WAL (embedded NATS JetStream), `200 OK` instantly, background batch-flush; schema-validated against `system.columns`; optional ID-based dedup (idempotent ingest); dead-letter queue for failed inserts. -- **Query** — in-process Ristretto cache + `singleflight` coalescing; type-safe structured query AST; Tinybird-style named pipes (parameterized SQL endpoints). +- **Query** — result cache (in-process Ristretto, or a Redis shared by every instance) + `singleflight` coalescing; type-safe structured query AST; Tinybird-style named pipes (parameterized SQL endpoints). - **Real-time** — native SSE push, broadcast *before* the ClickHouse flush, with JetStream gap-fill for late/reconnecting clients. - **Security** — Hasura-style per-table, per-role column + row policies with JWT claim templating, defined in the hot-reloadable settings directory. - **Client** — `@wavehouse/sdk`: TypeScript client with query builder, live queries, streaming, and schema codegen; one runtime dependency (an SSE frame parser, ~1.4 KB gzipped). diff --git a/config.yaml b/config.yaml index 5519b78c..aa9171e4 100644 --- a/config.yaml +++ b/config.yaml @@ -43,8 +43,8 @@ clickhouse: password: "" max_total_conns: 0 # ceiling on open native connections across pools; 0 = none -# Each layer's implementation, chosen at boot. Only the in-process backend -# exists for each today, and it is the default. +# Each layer's implementation, chosen at boot. The in-process backend is the +# default for each, and the only one for these three. mq: backend: embedded # NATS JetStream under /nats dedupe: @@ -52,11 +52,18 @@ dedupe: coord: backend: local # reserved: nothing is elected yet -# In-process L1 cache size. The query time-bucket +# The query-result cache: local (in-process, sized by l1_max_cost) or redis +# (one Redis-compatible server shared by every instance; see the redis block +# below and the Configuration page for every key). The query time-bucket # (query.timestamp_bucket_seconds) is a settings key. cache: backend: local l1_max_cost: 67108864 + # redis: # read only with backend: redis + # addrs: ["localhost:6379"] # `docker compose -f deployments/compose/dependencies.yaml --profile redis up -d` + # key_prefix: wh + # timeout: 100ms # per operation; slower is a miss, never a failed query + # The password is a secret: WH_CACHE_REDIS_PASSWORD, not this file. # Auth has no on/off switch — the JWT middleware always runs. A request with no # token, or an invalid/expired one, falls back to the policy default_role; diff --git a/deployments/compose/dependencies.yaml b/deployments/compose/dependencies.yaml index de147c72..4cc4f2fd 100644 --- a/deployments/compose/dependencies.yaml +++ b/deployments/compose/dependencies.yaml @@ -33,5 +33,18 @@ services: timeout: 2s retries: 15 + # Optional shared cache for trying cache.backend=redis locally, e.g. two + # host-side instances on different ports: `--profile redis`. No + # persistence, and /data on tmpfs, so it leaves no volume behind. + redis: + profiles: [redis] + # Pinned to match internal/cache's integration suite. + image: redis:8.10.2-alpine + command: ["redis-server", "--save", "", "--appendonly", "no", "--maxmemory", "256mb", "--maxmemory-policy", "allkeys-lru"] + ports: + - "6379:6379" + tmpfs: + - /data + volumes: clickhouse-data: diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index 1634aaab..c82739c0 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -560,7 +560,7 @@ The inbound request body is capped at 1 MiB; a body over the cap is rejected wit ### `GET/POST /v1/pipes/{name}` — Execute Named Pipe -Executes a pre-defined named query (pipe) with parameter binding. Parameters can be supplied via query string and/or JSON body. Results are cached in the shared L1 (Ristretto) with singleflight coalescing — same machinery as the structured query endpoint, keyed by [tenant](/deployment#multi-tenant-deployments) like it, and again, unlike `/v1/ops/query`. +Executes a pre-defined named query (pipe) with parameter binding. Parameters can be supplied via query string and/or JSON body. Results are cached in the query cache ([`cache.backend`](/configuration#backends): in-process, or a Redis shared by every instance) with singleflight coalescing — same machinery as the structured query endpoint, keyed by [tenant](/deployment#multi-tenant-deployments) like it, and again, unlike `/v1/ops/query`. **Query Parameters:** Any key matching a pipe parameter name. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 2a06d9ff..6c59e5f1 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -28,7 +28,7 @@ flowchart TD MQ --> BC["Buffer Consumer
(batch flush)"] BC -.->|failed inserts| DLQ["DLQ"]:::fail - QH["Query Handler"] --> Cache["Cache
(Ristretto + singleflight)"] + QH["Query Handler"] --> Cache["Cache
(local or Redis + singleflight)"] SSH["SSE Handler"] --> Hub["Stream Hub
(project once per role)"] @@ -90,7 +90,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability (after which each `Config.Warnings` line is logged at `WARN`), ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the whole cache — structured-query and pipe results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth, dedupe and cache hooks use too), ending the open streams of a tenant no longer served, and `wireCache`'s hook prunes the cache's version index the same way (`LocalCache.Prune`, [#262](https://github.com/Wave-RF/WaveHouse/issues/262)), so a tenant no longer served stops holding it. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. `wireCache` has two: `local`, the in-process `LocalCache`, and `redis`, the shared `RedisCache` built from the `cache.redis` block, which boots bypassed rather than failing when its server is unreachable. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the whole cache — structured-query and pipe results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth, dedupe and cache hooks use too), ending the open streams of a tenant no longer served, and `wireCache`'s hook prunes the cache's version index the same way (`LocalCache.Prune`, [#262](https://github.com/Wave-RF/WaveHouse/issues/262)), so a tenant no longer served stops holding it. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out @@ -112,7 +112,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on, each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope: one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)). `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. - **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: one version per tenant, per (tenant, table) and per (tenant, table, scope), each keyed by its name alone and bumped in place, so the index holds one entry per live tenant, table and scope however often each is bumped ([#262](https://github.com/Wave-RF/WaveHouse/issues/262)). A query key folds the tenant's version and, for each dependency, its tenant's, table's and scope's, so bumping a table (a scopeless write) orphans every scope of it, and bumping one scope orphans that scope and the whole-table view — scope is reserved and empty today, so every write is the whole-table bump — all without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results. A tenant's version is a *generation*, unique within the process and handed out by the first key built for the tenant; `BumpTenant` (behind `InvalidateTenant`) drops the tenant's whole index, so the next key gets a fresh generation no cached entry folds, orphaning every cached result of the tenant in one step — a pipe result with no dependencies, and a table no bump ever keyed, included — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). `LocalCache.Prune` does the same for every tenant no longer served, which `internal/app` runs after each settings reload, so a tenant removed or rejected stops holding its index. A table bump drops the table's scope versions with it, since every key they were folded into also folds the old table version; and a bump of a tenant with no index is a no-op, since no key folds its next generation yet. The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. -- **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET`, `MGET` and `PING`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token (a string of another length, a list, a hash) is replaced, so a token lost to eviction, expiry or a restart is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing. Nothing selects this backend yet — the boot config and wiring arrive with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s E4. +- **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET`, `MGET` and `PING`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token (a string of another length, a list, a hash) is replaced, so a token lost to eviction, expiry or a restart is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing. `cache.backend: redis` selects it: `internal/app`'s `wireCache` maps the boot config's `cache.redis` block onto `RedisConfig`, reading the TLS files, and releases it with the other components. - **redis_codec.go** — the key schema (`tokenKeys`, `bumpKeys`, `valueKey`) and the value frame: format, flags, expiry, the token list, the payload, zstd-compressed from `CompressMinBytes` when that is smaller. A stored value is capped at `MaxValueBytes` and a decoded one at eight times that, which refuses a zip bomb planted in a shared server; an unknown format, as a newer process writes during a rolling upgrade, is a miss. - **breaker.go** — the circuit breaker: `BreakerThreshold` (5) consecutive failures open it, and while open every operation skips the server at once; after `BreakerOpenFor` (5 s) one background `PING` decides whether it closes. Only a transport failure or a timeout counts against the server: any reply, an error reply or one the backend cannot use included, counts as a success, for operations and the probe alike, and a caller that gave up first counts as nothing. - **pending.go** — invalidations the server did not take, kept per token key (repeats coalesce) and retried by a background loop with backoff from 100 ms to 10 s until they land. Landing late is still correct: a fresh token orphans the pre-write entries and any fill made meanwhile. Past `PendingMax` (100,000) keys the set collapses to one tenant bump per affected tenant. `Close` makes one last attempt at the bumps still owed, past the breaker (an open one is why they are owed); what that attempt cannot deliver is lost, and the entries they would orphan are served until their TTL — the same failure as a worker stopping between an insert and its invalidation. @@ -120,9 +120,10 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ ### `config/` — Configuration -- **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — the implementation each layer runs on, resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). +- **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — the implementation each layer runs on, resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `cache.redis.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). - **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load` when `Config.NeedsDataDir` says a selected backend keeps state there, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. -- **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` at the end of `Validate`), which refuses a value not on the list and names the ones that are. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns the valid combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), which `app.New` logs at `WARN`. +- **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` at the end of `Validate`), which refuses a value not on the list and names the ones that are. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns what a valid configuration is still likely to get wrong — the combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), a `cache.redis` block that is not read, certificate verification turned off — which `app.New` logs at `WARN`. +- **cache_redis.go** — `CacheRedisConfig`, the `cache.redis` sub-block, and its checks: an address, each `host:port`, a known mode, a sentinel's master name, `db` 0 in cluster mode, positive timeouts and sizes, a `version_ttl` of at least 2 s, and `compress_min_bytes` positive or `-1` (never: cleanenv reads a `0` in the file as unset and applies the default, so `0` cannot mean off). `CacheRedisTLS.Config` builds the `tls.Config`, reading the files; `Validate` calls it so an unreadable file refuses boot, and `internal/app` calls it again to build the connection. A TLS key set while `tls.enabled` is off is an error rather than a plaintext connection. - **strict.go** — `rejectUnknownKeys`, the YAML half: re-reads the file as a generic tree and walks it against the struct's `yaml` tags, listing every key the struct doesn't declare. cleanenv itself is lenient by design, which is exactly wrong for boot config once keys have moved to the settings directory. - **persistence.go** — `WarnIfFreshDataDir` logs the startup `WARN` when `data_dir` is missing or empty (on a redeploy, the sign that the volume didn't persist); `LogStorageInitError` attaches the UID-65532 `permissionHint` to a NATS or Pebble open failure that looks like a permission denial — the same hint string `CheckDataDir` uses. @@ -366,7 +367,7 @@ Client GET /v1/stream | Analytics DB | ClickHouse | Primary data store + schema source of truth | | Message Queue | NATS + JetStream | Durable event streaming | | L1 Cache | Ristretto v2 | In-process memory cache | -| Shared cache | [rueidis](https://github.com/redis/rueidis) | Redis-compatible client for the shared backend (not yet selectable) | +| Shared cache | [rueidis](https://github.com/redis/rueidis) | Redis-compatible client for the shared backend (`cache.backend: redis`) | | Embedded KV | Pebble | Optional deduplication | | Config | cleanenv | YAML + env var config loading | | Release | GoReleaser | Cross-platform binary builds | diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 193a6c21..82744a35 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -39,16 +39,16 @@ This page is boot config only — what the platform operator owns (wiring, lifec ### Backends -Each layer's implementation is chosen once, at boot. Today every layer has one backend, the in-process one, and it is the default, so a config that sets none of these keys runs as it always has. A value this build has no backend for refuses boot and names the valid ones. +Each layer's implementation is chosen once, at boot. Every layer defaults to its in-process backend, so a config that sets none of these keys runs as it always has; the cache also has a shared one, `redis`. A value this build has no backend for refuses boot and names the valid ones. | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | | `mq.backend` | `WH_MQ_BACKEND` | `embedded` | The message queue. `embedded`: NATS JetStream inside this process, under `/nats`. It listens on no port, so no other process can reach its queue. | -| `cache.backend` | `WH_CACHE_BACKEND` | `local` | The query-result cache. `local`: in this process, sized by `cache.l1_max_cost`. | +| `cache.backend` | `WH_CACHE_BACKEND` | `local` | The query-result cache. `local`: in this process, sized by `cache.l1_max_cost`. `redis`: one Redis-compatible server shared by every process, configured by [`cache.redis`](#cache), so an insert one process makes invalidates what every process has cached. | | `dedupe.backend` | `WH_DEDUPE_BACKEND` | `pebble` | Where ingest dedupe keeps the event ids it has seen. `pebble`: in this process, under `/pebble`, open while any tenant has dedupe on. | | `coord.backend` | `WH_COORD_BACKEND` | `local` | Reserved for the leases that will elect work only one process may do at a time, such as the sweeper. Nothing is elected yet: every process runs its own sweeper, and `local`, the only value, changes nothing. | -Settings for one backend will go in a sub-block named after it, `.`, read only when that backend is selected. No backend has settings yet, so today any such sub-block, `mq.embedded` included, is an unknown key and refuses boot. `mq` and `dedupe` also appear in the settings directory's `config.json`, with different keys (`mq.max_bytes_gb`, `dedupe.enabled`, …): those are per-tenant tunables and stay there, and one written in `config.yaml` refuses boot as an unknown key. +Settings for one backend go in a sub-block named after it, `.`, read only when that backend is selected. `cache.redis` is the only one so far; any other, `mq.embedded` included, is an unknown key and refuses boot. A `cache.redis.addrs` set while `cache.backend` is `local` is logged at `WARN` at boot, since the block is not read. `mq` and `dedupe` also appear in the settings directory's `config.json`, with different keys (`mq.max_bytes_gb`, `dedupe.enabled`, …): those are per-tenant tunables and stay there, and one written in `config.yaml` refuses boot as an unknown key. ### Server @@ -120,7 +120,34 @@ Each tenant's queue has its own disk budget, `mq.max_bytes_gb`, a hot-reloadable | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | -| `cache.l1_max_cost` | `WH_CACHE_L1_MAX_COST` | `67108864` | Maximum L1 cache size in bytes (~64 MB). The time-range bucket structured queries normalize to is `query.timestamp_bucket_seconds` in the [Settings Directory](/settings-directory#configjson-keys). | +| `cache.l1_max_cost` | `WH_CACHE_L1_MAX_COST` | `67108864` | Maximum size in bytes (~64 MB) of the `local` backend's in-process cache. The time-range bucket structured queries normalize to is `query.timestamp_bucket_seconds` in the [Settings Directory](/settings-directory#configjson-keys). | + +The `redis` backend's settings, read only when `cache.backend` is `redis`. It runs on Redis, Valkey, Dragonfly, ElastiCache (including Serverless) and MemoryDB: it sends only `GET`, `SET`, `MGET` and `PING`. [Deployment](/deployment#multiple-instances-and-the-shared-cache) covers sizing, `maxmemory-policy` and what a reader on another instance can see. + +| YAML Key | Env Var | Default | Description | +| --- | --- | ------- | ----------- | +| `cache.redis.addrs` | `WH_CACHE_REDIS_ADDRS` | — | **Required** with `backend: redis.` `host:port` of the server; several are a cluster's seed nodes or the sentinels. Comma-separated in the env var. | +| `cache.redis.mode` | `WH_CACHE_REDIS_MODE` | `standalone` | `standalone`, `cluster` or `sentinel`. | +| `cache.redis.sentinel_master` | `WH_CACHE_REDIS_SENTINEL_MASTER` | — | The master set name. Required with `mode: sentinel`. | +| `cache.redis.username` | `WH_CACHE_REDIS_USERNAME` | — | ACL user. Empty uses the server's `default` user. | +| `cache.redis.password` | `WH_CACHE_REDIS_PASSWORD` | — | A secret: set it through the environment (or your secret store's env injection), not in a tracked `config.yaml`. | +| `cache.redis.db` | `WH_CACHE_REDIS_DB` | `0` | Database number (`SELECT`). Standalone and sentinel only: a cluster has only database `0`, and any other value refuses boot. | +| `cache.redis.tls.enabled` | `WH_CACHE_REDIS_TLS_ENABLED` | `false` | Connect over TLS, verifying the server against the system roots or `ca_file`. Any other `tls` key set while this is off refuses boot, rather than connecting in plaintext. | +| `cache.redis.tls.ca_file` | `WH_CACHE_REDIS_TLS_CA_FILE` | — | PEM file of the authorities to trust instead of the system roots. | +| `cache.redis.tls.cert_file` | `WH_CACHE_REDIS_TLS_CERT_FILE` | — | Client certificate (PEM) for mutual TLS. Set together with `key_file`. | +| `cache.redis.tls.key_file` | `WH_CACHE_REDIS_TLS_KEY_FILE` | — | The client certificate's private key (PEM). | +| `cache.redis.tls.server_name` | `WH_CACHE_REDIS_TLS_SERVER_NAME` | — | Name to verify the server's certificate against, when it differs from the address. | +| `cache.redis.tls.insecure_skip_verify` | `WH_CACHE_REDIS_TLS_INSECURE_SKIP_VERIFY` | `false` | Accept any server certificate. Logged at `WARN` at boot: whoever can intercept the connection can read and replace cached results. | +| `cache.redis.key_prefix` | `WH_CACHE_REDIS_KEY_PREFIX` | `wh` | Leads every key, so several deployments can share one server. No `{` or `}`. | +| `cache.redis.timeout` | `WH_CACHE_REDIS_TIMEOUT` | `100ms` | Per operation. A lookup or fill that takes longer is a miss or a skipped fill, never a failed query. | +| `cache.redis.dial_timeout` | `WH_CACHE_REDIS_DIAL_TIMEOUT` | `1s` | Per connection attempt. | +| `cache.redis.max_value_bytes` | `WH_CACHE_REDIS_MAX_VALUE_BYTES` | `1048576` | Largest result stored, after compression (1 MiB). A larger one is returned to the caller but not cached. | +| `cache.redis.compress_min_bytes` | `WH_CACHE_REDIS_COMPRESS_MIN_BYTES` | `1024` | Results at least this large are zstd-compressed when that makes them smaller. `-1` never compresses. `0` is not "off": in the YAML file it reads as unset and takes the default, and in the env var it refuses boot. | +| `cache.redis.version_ttl` | `WH_CACHE_REDIS_VERSION_TTL` | `168h` | How long a table's or tenant's version token outlives its last write, so the tokens of dropped tables and removed tenants eventually expire. At least `2s`. An expired token only causes misses. | + +**When the server is unreachable or misbehaves, the cache is bypassed; queries are not.** A failure or a timeout makes the lookup a miss and the fill a no-op, and five in a row open a circuit breaker that skips the server entirely until a probe, every 5 s, gets an answer. Queries then go straight to ClickHouse, still coalesced per instance by `singleflight`. An invalidation the server did not take is kept and retried until it lands. `/readyz` does not depend on the cache. + +**Boot does not wait for the server.** A malformed block (an address without a port, `mode: cluster` with `db` other than `0`, an unreadable or unparsable TLS file) refuses boot. A server that cannot be reached, or that refuses the credentials, does not: the process boots with the cache bypassed and keeps reconnecting, with backoff up to 30 s. A rejected password (`WRONGPASS`, `NOAUTH`) is logged at `ERROR` on every attempt; any other failure at `WARN`. This is deliberate: a rotated Redis password must not crash-loop every instance at once. Watch `wavehouse_cache_breaker_open`, which reads `1` while the cache is bypassed, including before the first connection. ### Authentication @@ -210,8 +237,28 @@ mq: backend: embedded # in-process NATS JetStream under /nats cache: - backend: local - l1_max_cost: 67108864 + backend: local # local | redis + l1_max_cost: 67108864 # the local backend's size + redis: # read only with backend: redis + addrs: [] # required with backend: redis, e.g. ["redis:6379"] + mode: standalone # standalone | cluster | sentinel + sentinel_master: "" + username: "" + password: "" # a secret: set WH_CACHE_REDIS_PASSWORD instead + db: 0 + tls: + enabled: false + ca_file: "" + cert_file: "" + key_file: "" + server_name: "" + insecure_skip_verify: false + key_prefix: wh + timeout: 100ms + dial_timeout: 1s + max_value_bytes: 1048576 + compress_min_bytes: 1024 # -1 = never + version_ttl: 168h dedupe: backend: pebble # in-process Pebble under /pebble @@ -267,6 +314,25 @@ WH_CH_MAX_TOTAL_CONNS=0 WH_MQ_BACKEND=embedded WH_CACHE_BACKEND=local WH_CACHE_L1_MAX_COST=67108864 +# Read only with WH_CACHE_BACKEND=redis; WH_CACHE_REDIS_ADDRS is then required. +WH_CACHE_REDIS_ADDRS= +WH_CACHE_REDIS_MODE=standalone +WH_CACHE_REDIS_SENTINEL_MASTER= +WH_CACHE_REDIS_USERNAME= +WH_CACHE_REDIS_PASSWORD= +WH_CACHE_REDIS_DB=0 +WH_CACHE_REDIS_TLS_ENABLED=false +WH_CACHE_REDIS_TLS_CA_FILE= +WH_CACHE_REDIS_TLS_CERT_FILE= +WH_CACHE_REDIS_TLS_KEY_FILE= +WH_CACHE_REDIS_TLS_SERVER_NAME= +WH_CACHE_REDIS_TLS_INSECURE_SKIP_VERIFY=false +WH_CACHE_REDIS_KEY_PREFIX=wh +WH_CACHE_REDIS_TIMEOUT=100ms +WH_CACHE_REDIS_DIAL_TIMEOUT=1s +WH_CACHE_REDIS_MAX_VALUE_BYTES=1048576 +WH_CACHE_REDIS_COMPRESS_MIN_BYTES=1024 +WH_CACHE_REDIS_VERSION_TTL=168h WH_DEDUPE_BACKEND=pebble WH_COORD_BACKEND=local diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 79dae7d6..63f1b1a9 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -397,6 +397,26 @@ The folder name is the tenant id, and each folder is a complete settings directo `X-Tenant-ID` is a generic name, and some gateways and service meshes stamp one on every request. WaveHouse used to ignore it; now, over a settings directory that holds the four files, any value other than `0` names an unknown tenant, so **every `/v1` route outside `/v1/ops/*` answers `404 unknown tenant: `** (a `400` when the value is not a tenant id at all, a dotted hostname, say) — the SDK's `/v1/health` reachability ping included, while the bare probes and the admin surface stay green. Strip the inbound header at the edge ([header forwarding](/reverse-proxy#header-and-auth-forwarding)) unless you are using it deliberately. +## Multiple instances and the shared cache + +Several WaveHouse instances can serve one ClickHouse behind a load balancer, but most of what each one holds is its own. The message queue is embedded, so an event is inserted by the instance that took its `POST /v1/ingest`, and reaches only that instance's SSE subscribers. The dedupe store is per instance too, so an id one instance has seen is new to another. + +The query-result cache is the layer that can be shared today. With the default `cache.backend: local`, each instance caches in its own memory, and an insert invalidates only the cache of the instance that made it. Every other instance keeps serving its cached results for the rows before the insert until each entry's TTL runs out, between 10 s and 1 h depending on how long the query took. With [`cache.backend: redis`](/configuration#cache), every instance reads and fills one Redis-compatible server, and an insert on any instance invalidates the cached results of every instance. + +**What another instance can see.** Ingest is already asynchronous: `/v1/ingest` answers before the batch is inserted. Once the inserting instance's worker has written the batch to ClickHouse, it replaces the table's version token in Redis, and from then on a lookup on any instance misses and reads the new rows. The cache adds no delay of its own beyond that single write. The exceptions: + +- **The server is unreachable from the inserting instance.** The invalidation is kept and retried until it lands (`wavehouse_cache_invalidations_pending` counts what is owed). Meanwhile other instances that can still reach the server keep serving the older results, for as long as the outage lasts and at most until each entry's TTL. An instance that stops while invalidations are still owed loses them, with the same bound. The same thing happens today when a process stops between an insert and its invalidation. +- **A failover to a replica that had not yet received the latest token writes** can bring back entries filed under the older tokens, bounded by the replication lag at the moment of failover and those entries' TTL. WaveHouse never reads from replicas. +- **Admin writes through `POST /v1/ops/query`** do not invalidate the cache ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). With a shared cache, the stale results they leave are served by every instance, not only one. + +**Sizing the server.** Every key WaveHouse writes has a TTL, and a version token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses, never bring back an entry it had invalidated. So set `maxmemory` and let the server evict: `maxmemory-policy allkeys-lru` (or `allkeys-lfu`, `volatile-lru`, `volatile-lfu`). Under `noeviction`, a full server refuses the writes. Fills then fail (counted by `wavehouse_cache_set_failures_total{reason="oom"}`), lookups keep working, and invalidations are kept and retried. A stored result is capped at `cache.redis.max_value_bytes` (1 MiB compressed). A tenant's version tokens share one hash tag, so each lookup reads them in one `MGET` in cluster mode as well. The results themselves carry no hash tag and spread across shards. Persistence is not needed: an empty server after a restart is a cold cache, not a wrong one. + +**Coalescing stays per instance.** `singleflight` collapses identical concurrent queries within each instance, so a cold hot query costs at most one ClickHouse query per instance, not one per request. + +**Metrics** (meter `wavehouse-cache`, every series labeled `backend="redis"`, no tenant label): `wavehouse_cache_lookups_total{result}` (`hit`, `miss`, `stale`, `bypass`, `error`), `wavehouse_cache_op_duration_seconds{op}` (`lookup`, `set`, `invalidate`), `wavehouse_cache_breaker_open` (1 while the cache is bypassed), `wavehouse_cache_invalidations_total{result}` (`ok`, `deferred`), `wavehouse_cache_invalidations_pending`, `wavehouse_cache_value_bytes`, `wavehouse_cache_oversize_total` and `wavehouse_cache_set_failures_total{reason}` (`oom`, `timeout`, `other`). Two signals are worth alerting on: `wavehouse_cache_breaker_open` at 1, or `wavehouse_cache_invalidations_pending` above 0, for more than a few minutes. + +For local development, `docker compose -f deployments/compose/dependencies.yaml --profile redis up -d` starts a Redis on `localhost:6379` with no persistence. + ## ClickHouse Schema WaveHouse uses a **Bring Your Own Schema** model. You create your tables in ClickHouse with whatever columns and engines you need. WaveHouse discovers the schemas automatically via `system.columns` and validates ingest data against them — see [Schema Validation](/api#post-v1ingesttabletable--ingest-data) for the rules a record must satisfy. diff --git a/docs/src/content/docs/development.md b/docs/src/content/docs/development.md index 8e6a8503..2a9ea7d5 100644 --- a/docs/src/content/docs/development.md +++ b/docs/src/content/docs/development.md @@ -82,7 +82,7 @@ make dev WaveHouse is now running at `http://localhost:8080` in standalone mode with: - **Embedded NATS** (JetStream) — no external MQ needed -- **L1 cache only** (Ristretto) — no external cache needed +- **In-process cache** (Ristretto, `cache.backend: local`) — no external cache needed; to try the shared one, start Redis with `docker compose -f deployments/compose/dependencies.yaml --profile redis up -d` and set `WH_CACHE_BACKEND=redis WH_CACHE_REDIS_ADDRS=localhost:6379` - **Trial policy** — the dev settings directory `./settings` is seeded on first run with the compose stack's permissive `public` policy, so tokenless requests to the demo tables work (see [Test the API](#test-the-api)) - **Dedup disabled** by default — no Pebble needed - **Schema discovery** — automatically finds your ClickHouse tables @@ -454,7 +454,7 @@ WaveHouse/ │ ├── api/ # HTTP handlers, router, middleware │ ├── app/ # Process wiring (build every component, run under one errgroup, release in reverse) │ ├── auth/ # JWT/JWKS authentication middleware -│ ├── cache/ # L1 (Ristretto) + L2 caching +│ ├── cache/ # Query cache: in-process (Ristretto) or shared (Redis-compatible) │ ├── chconn/ # ClickHouse pools, one per connection tuple (reconciled on settings reload) │ ├── chsql/ # Shared ClickHouse SQL helpers (quoting + bind-safety) │ ├── config/ # YAML + env var configuration diff --git a/docs/src/content/docs/getting-started.md b/docs/src/content/docs/getting-started.md index f24c3ee6..667137d1 100644 --- a/docs/src/content/docs/getting-started.md +++ b/docs/src/content/docs/getting-started.md @@ -75,7 +75,7 @@ curl -s -X POST "http://localhost:8080/v1/query?table=clicks" \ -d '{"columns": ["page", "button", "score"], "limit": 10}' ``` -`POST /v1/query?table={table}` and `GET/POST /v1/pipes/{name}` are cached in-process (L1 Ristretto) with singleflight coalescing — duplicate concurrent queries hit ClickHouse once. For raw SQL there's `POST /v1/ops/query` (an admin escape hatch that never caches, emitting `Cache-Control: no-store`), but it's **admin-only** — the trial `public` role can't reach it. To use it, swap the public default for real auth: configure a JWT secret and present a token whose role is the policy [`admin_role`](/access-control#admin_role--the-privileged-role). +`POST /v1/query?table={table}` and `GET/POST /v1/pipes/{name}` are cached — in-process by default, or in a Redis shared by every instance with [`cache.backend: redis`](/configuration#cache) — with singleflight coalescing, so duplicate concurrent queries hit ClickHouse once. For raw SQL there's `POST /v1/ops/query` (an admin escape hatch that never caches, emitting `Cache-Control: no-store`), but it's **admin-only** — the trial `public` role can't reach it. To use it, swap the public default for real auth: configure a JWT secret and present a token whose role is the policy [`admin_role`](/access-control#admin_role--the-privileged-role). :::tip[Prefer a type-safe client?] The [TypeScript SDK](/sdk) wraps this endpoint in a chainable query builder with autocomplete on your table names and row types — plus live queries and streaming. The raw shapes are in the [structured query reference](/api#post-v1querytabletable--structured-query). diff --git a/docs/src/content/docs/pipes.mdx b/docs/src/content/docs/pipes.mdx index b8c7efa0..b6412922 100644 --- a/docs/src/content/docs/pipes.mdx +++ b/docs/src/content/docs/pipes.mdx @@ -169,7 +169,7 @@ curl -X POST http://localhost:8080/v1/pipes/top_pages \ -d '{"start_date": "2024-01-01", "limit": 20}' ``` -The response is a JSON array of rows. Results flow through the shared in-process L1 cache (Ristretto) with singleflight coalescing, so concurrent identical calls hit ClickHouse once; an `X-Cache: HIT` or `X-Cache: MISS` header tells you which path served the response. +The response is a JSON array of rows. Results flow through the query cache (in-process by default, or a Redis shared by every instance with [`cache.backend: redis`](/configuration#cache)) with singleflight coalescing, so concurrent identical calls hit ClickHouse once; an `X-Cache: HIT` or `X-Cache: MISS` header tells you which path served the response. | Status | Body | Cause | | ------ | ---- | ----- | diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index 56135462..922f3a3f 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -179,7 +179,7 @@ The tenant tunables. Every key is required (a missing one is a validation error) } ``` -What stays in boot config is only what cannot change under a running process — the implementation each layer runs on (`mq.backend`, `cache.backend`, `dedupe.backend`, `coord.backend`), resource sizing (`data_dir`, `cache.l1_max_cost`, `clickhouse.max_total_conns`), the listeners, the observability exporters — and the **secrets**: `clickhouse.password`, `auth.jwt_secret`, `auth.operator_key`. Secrets never belong in a tracked JSON file, so they stay in the environment and are combined with the wiring here on every (re)connect; rotating one is a restart. See [Configuration](/configuration). Everything else lives here and reloads. +What stays in boot config is only what cannot change under a running process — the implementation each layer runs on (`mq.backend`, `cache.backend`, `dedupe.backend`, `coord.backend`) and a shared backend's connection (`cache.redis`), resource sizing (`data_dir`, `cache.l1_max_cost`, `clickhouse.max_total_conns`), the listeners, the observability exporters — and the **secrets**: `clickhouse.password`, `cache.redis.password`, `auth.jwt_secret`, `auth.operator_key`. Secrets never belong in a tracked JSON file, so they stay in the environment and are combined with the wiring here on every (re)connect; rotating one is a restart. See [Configuration](/configuration). Everything else lives here and reloads. ## Deduplication diff --git a/internal/app/app_test.go b/internal/app/app_test.go index cf727b6d..2ce9ca5b 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -702,6 +702,77 @@ func TestReload_PrunesCacheIndexToServedTenants(t *testing.T) { assert.Equal(t, map[tenant.ID]bool{"acme": false, "globex": true}, rec.last(), "removed; the repaired one served again") } +// redisTestConfig is testConfig with cache.backend=redis at addr, carrying +// the defaults Load would apply. +func redisTestConfig(t *testing.T, settingsDir, addr string) *config.Config { + t.Helper() + cfg := testConfig(t, settingsDir) + cfg.Cache = config.Cache{Backend: config.CacheRedis, Redis: config.CacheRedisConfig{ + Addrs: []string{addr}, Mode: config.RedisStandalone, KeyPrefix: "wh", + Timeout: 100 * time.Millisecond, DialTimeout: 200 * time.Millisecond, + MaxValueBytes: 1 << 20, CompressMinBytes: 1 << 10, VersionTTL: time.Hour, + }} + require.NoError(t, cfg.Validate()) + return cfg +} + +// cache.backend=redis wires the shared backend. A server that cannot be +// reached does not refuse boot: the cache starts bypassed, and the reload +// hook that prunes an in-process index leaves it alone. +func TestNew_RedisCacheBootsBypassedWhenUnreachable(t *testing.T) { + root := writeNestedSettings(t, map[string]map[string]any{"acme": nil, "globex": nil}) + a := newApp(t, redisTestConfig(t, root, closedAddr(t)), Options{}) + _, ok := a.cache.(*cache.RedisCache) + require.True(t, ok, "cache is %T", a.cache) + assert.Contains(t, componentNames(a), "cache") + + require.NoError(t, os.RemoveAll(filepath.Join(root, "acme"))) + a.tenants.Reload("test") // the prune hook must not trip on a non-pruner + + entry, snap, err := a.cache.Lookup(t.Context(), "globex", "sha", nil) + require.NoError(t, err) + assert.Nil(t, entry.Value, "bypassed: a miss") + assert.NoError(t, a.cache.Set(t.Context(), snap, []byte("v"), time.Minute), "and the fill a no-op") +} + +// A TLS file that went missing between validation and wiring refuses boot, +// naming the key. +func TestNew_RedisCacheRefusesAnUnreadableTLSFile(t *testing.T) { + guardGlobals(t) + cfg := redisTestConfig(t, writeSettings(t, nil), closedAddr(t)) + cfg.Cache.Redis.TLS = config.CacheRedisTLS{Enabled: true, CAFile: filepath.Join(t.TempDir(), "gone.pem")} + _, err := New(t.Context(), Options{Config: cfg}) + require.ErrorContains(t, err, "cache init: cache.redis.tls.ca_file") +} + +// The boot config's defaults are the backend's, and -1 is the backend's +// "never compress". Driven from Load, not a literal, so a default changed on +// one side only fails here. +func TestRedisConfig_FromLoadedDefaults(t *testing.T) { + t.Setenv("WH_SETTINGS_DIR", t.TempDir()) + t.Setenv("WH_CACHE_BACKEND", "redis") + t.Setenv("WH_CACHE_REDIS_ADDRS", "a:6379,b:6379") + t.Setenv("WH_CACHE_REDIS_PASSWORD", "pw") + loaded, err := config.Load(filepath.Join(t.TempDir(), "none.yaml")) + require.NoError(t, err) + got, err := redisConfig(loaded.Cache.Redis) + require.NoError(t, err) + assert.Equal(t, cache.RedisConfig{ + Addrs: []string{"a:6379", "b:6379"}, Mode: cache.RedisStandalone, Password: "pw", + KeyPrefix: cache.DefaultRedisKeyPrefix, Timeout: cache.DefaultRedisTimeout, + DialTimeout: cache.DefaultRedisDialTimeout, MaxValueBytes: cache.DefaultRedisMaxValueBytes, + CompressMinBytes: cache.DefaultRedisCompressMinBytes, VersionTTL: cache.DefaultRedisVersionTTL, + }, got) + + loaded.Cache.Redis.CompressMinBytes = -1 + loaded.Cache.Redis.Mode = config.RedisCluster + got, err = redisConfig(loaded.Cache.Redis) + require.NoError(t, err) + assert.Zero(t, got.CompressMinBytes) + assert.Equal(t, cache.RedisCluster, got.Mode) + assert.Equal(t, cache.RedisSentinel, config.RedisSentinel) +} + // keepalive is a config.json patch setting the stream block's keepalive pair. func keepalive(interval, buckets int) map[string]any { return map[string]any{"stream": map[string]any{"keepalive_interval": interval, "keepalive_buckets": buckets, "gap_window_minutes": 15}} diff --git a/internal/app/wire.go b/internal/app/wire.go index 3cbe71a8..e1fe828d 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -627,7 +627,7 @@ var _ pruner = (*cache.LocalCache)(nil) // is chosen. After every reload a tenant no longer served, removed or // rejected alike, has its in-process version index dropped (#262); its cache // is orphaned with it, as it would be anyway when it came back -// (wireClickHouse). +// (wireClickHouse). A shared backend keeps no such index and is skipped. func (a *App) wireCache() error { var c cache.Cache switch b := a.cfg.Cache.Backend; b { @@ -637,6 +637,16 @@ func (a *App) wireCache() error { return fmt.Errorf("cache init: %w", err) } c = l1 + case config.CacheRedis: + rc, err := redisConfig(a.cfg.Cache.Redis) + if err != nil { + return fmt.Errorf("cache init: %w", err) + } + r, err := cache.NewRedis(rc) + if err != nil { + return fmt.Errorf("cache init: %w", err) + } + c = r default: return unreachableBackend("cache.backend", b) } @@ -650,6 +660,35 @@ func (a *App) wireCache() error { return nil } +// redisConfig maps the boot config's cache.redis block onto the backend's +// config. Load has applied every default and validated the block; the TLS +// files are read again here, so the connection uses what is on disk now. +func redisConfig(r config.CacheRedisConfig) (cache.RedisConfig, error) { + t, err := r.TLS.Config() + if err != nil { + return cache.RedisConfig{}, err + } + compressMin := r.CompressMinBytes + if compressMin < 0 { + compressMin = 0 // the backend's "never" + } + return cache.RedisConfig{ + Addrs: r.Addrs, + Mode: r.Mode, + SentinelMaster: r.SentinelMaster, + Username: r.Username, + Password: r.Password, + DB: r.DB, + TLS: t, + KeyPrefix: r.KeyPrefix, + Timeout: r.Timeout, + DialTimeout: r.DialTimeout, + MaxValueBytes: r.MaxValueBytes, + CompressMinBytes: compressMin, + VersionTTL: r.VersionTTL, + }, nil +} + // unreachableBackend is each layer switch's default case. config.Validate // refuses a backend with no case, so reaching it means a Config built by hand // without one (the zero value is not the default), or a case missing here. diff --git a/internal/config/backends.go b/internal/config/backends.go index f2ab9330..893d2304 100644 --- a/internal/config/backends.go +++ b/internal/config/backends.go @@ -35,21 +35,34 @@ func (m MQ) validate() error { // CacheBackend names the query-result cache implementation. type CacheBackend string -// CacheLocal is the in-process Ristretto cache, sized by cache.l1_max_cost. -const CacheLocal CacheBackend = "local" +const ( + // CacheLocal is the in-process Ristretto cache, sized by + // cache.l1_max_cost. + CacheLocal CacheBackend = "local" + // CacheRedis is one Redis-compatible server shared by every process, + // configured by cache.redis. + CacheRedis CacheBackend = "redis" +) -var cacheBackends = []CacheBackend{CacheLocal} +var cacheBackends = []CacheBackend{CacheLocal, CacheRedis} // Cache selects and sizes the query-result cache. The time-range bucket // structured queries normalize to is a settings-directory key // (query.timestamp_bucket_seconds) — query shaping, not process memory. type Cache struct { - Backend CacheBackend `yaml:"backend" env:"WH_CACHE_BACKEND" env-default:"local"` - L1MaxCost int64 `yaml:"l1_max_cost" env:"WH_CACHE_L1_MAX_COST" env-default:"67108864"` + Backend CacheBackend `yaml:"backend" env:"WH_CACHE_BACKEND" env-default:"local"` + L1MaxCost int64 `yaml:"l1_max_cost" env:"WH_CACHE_L1_MAX_COST" env-default:"67108864"` + Redis CacheRedisConfig `yaml:"redis"` } func (c Cache) validate() error { - return checkBackend("cache.backend", "WH_CACHE_BACKEND", c.Backend, cacheBackends) + if err := checkBackend("cache.backend", "WH_CACHE_BACKEND", c.Backend, cacheBackends); err != nil { + return err + } + if c.Backend == CacheRedis { + return c.Redis.validate() + } + return nil } // DedupeBackend names where ingest dedupe keeps the ids it has seen. @@ -126,13 +139,20 @@ func (c *Config) NeedsDataDir() bool { } // Warnings returns what a valid configuration is still likely to get wrong, -// one line each, for boot to log at WARN. They are not errors because each is -// correct for a single replica, and one process cannot count its replicas. +// one line each, for boot to log at WARN. The shared-queue ones are not +// errors because each is correct for a single replica, and one process +// cannot count its replicas. func (c *Config) Warnings() []string { + var out []string + if c.Cache.Backend == CacheRedis && c.Cache.Redis.TLS.InsecureSkipVerify { + out = append(out, "cache.redis.tls.insecure_skip_verify is on: the cache accepts any certificate, so whoever can intercept the connection can read and replace cached query results") + } + if c.Cache.Backend != CacheRedis && c.Cache.Redis.hasAddrs() { + out = append(out, "cache.redis.addrs is set but cache.backend is "+string(c.Cache.Backend)+": the redis block is not read; set cache.backend=redis to share the cache") + } if !c.Distributed() { - return nil + return out } - var out []string if c.Cache.Backend == CacheLocal { out = append(out, "cache.backend=local with a shared mq.backend is correct for one replica only: an event ingested on another replica never invalidates this one's cache, so its reads stay stale until the cached entry expires") } diff --git a/internal/config/backends_test.go b/internal/config/backends_test.go index 0e70dc7a..77833697 100644 --- a/internal/config/backends_test.go +++ b/internal/config/backends_test.go @@ -111,7 +111,7 @@ func TestValidate_UnknownBackend(t *testing.T) { want string }{ {"mq", func(c *Config) { c.MQ.Backend = "kafka" }, `mq.backend (WH_MQ_BACKEND) "kafka" is not a backend this build has; valid: embedded`}, - {"cache", func(c *Config) { c.Cache.Backend = "redis" }, `cache.backend (WH_CACHE_BACKEND) "redis" is not a backend this build has; valid: local`}, + {"cache", func(c *Config) { c.Cache.Backend = "memcached" }, `cache.backend (WH_CACHE_BACKEND) "memcached" is not a backend this build has; valid: local, redis`}, {"dedupe", func(c *Config) { c.Dedupe.Backend = "dynamodb" }, `dedupe.backend (WH_DEDUPE_BACKEND) "dynamodb" is not a backend this build has; valid: pebble`}, {"coord", func(c *Config) { c.Coord.Backend = "nats" }, `coord.backend (WH_COORD_BACKEND) "nats" is not a backend this build has; valid: local`}, // The zero value, which a Config built without Load carries. diff --git a/internal/config/cache_redis.go b/internal/config/cache_redis.go new file mode 100644 index 00000000..f0af6713 --- /dev/null +++ b/internal/config/cache_redis.go @@ -0,0 +1,153 @@ +package config + +import ( + "crypto/tls" + "crypto/x509" + "errors" + "fmt" + "net" + "os" + "strings" + "time" +) + +// Redis deployment modes for cache.redis.mode. +const ( + RedisStandalone = "standalone" + RedisCluster = "cluster" + RedisSentinel = "sentinel" +) + +// CacheRedisConfig configures cache.backend=redis: one Redis-compatible server +// (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB) shared by every process. +// Read only when that backend is selected. +type CacheRedisConfig struct { + // Addrs are host:port pairs: the server, or seeds for a cluster, or the + // sentinels. + Addrs []string `yaml:"addrs" env:"WH_CACHE_REDIS_ADDRS"` + Mode string `yaml:"mode" env:"WH_CACHE_REDIS_MODE" env-default:"standalone"` + SentinelMaster string `yaml:"sentinel_master" env:"WH_CACHE_REDIS_SENTINEL_MASTER"` + Username string `yaml:"username" env:"WH_CACHE_REDIS_USERNAME"` + Password string `yaml:"password" env:"WH_CACHE_REDIS_PASSWORD"` + DB int `yaml:"db" env:"WH_CACHE_REDIS_DB"` + TLS CacheRedisTLS `yaml:"tls"` + // KeyPrefix leads every key, so deployments can share one server. + KeyPrefix string `yaml:"key_prefix" env:"WH_CACHE_REDIS_KEY_PREFIX" env-default:"wh"` + Timeout time.Duration `yaml:"timeout" env:"WH_CACHE_REDIS_TIMEOUT" env-default:"100ms"` + DialTimeout time.Duration `yaml:"dial_timeout" env:"WH_CACHE_REDIS_DIAL_TIMEOUT" env-default:"1s"` + // MaxValueBytes is the largest value stored, after compression. + MaxValueBytes int `yaml:"max_value_bytes" env:"WH_CACHE_REDIS_MAX_VALUE_BYTES" env-default:"1048576"` + // CompressMinBytes is the smallest value zstd-compressed; -1 never + // compresses. Not 0: the loader reads a 0 in the file as unset and + // applies the default, so 0 cannot mean off. + CompressMinBytes int `yaml:"compress_min_bytes" env:"WH_CACHE_REDIS_COMPRESS_MIN_BYTES" env-default:"1024"` + // VersionTTL is how long a version token outlives its last bump. + VersionTTL time.Duration `yaml:"version_ttl" env:"WH_CACHE_REDIS_VERSION_TTL" env-default:"168h"` +} + +// CacheRedisTLS is cache.redis.tls. The files are paths, read at boot. +type CacheRedisTLS struct { + Enabled bool `yaml:"enabled" env:"WH_CACHE_REDIS_TLS_ENABLED"` + CAFile string `yaml:"ca_file" env:"WH_CACHE_REDIS_TLS_CA_FILE"` + CertFile string `yaml:"cert_file" env:"WH_CACHE_REDIS_TLS_CERT_FILE"` + KeyFile string `yaml:"key_file" env:"WH_CACHE_REDIS_TLS_KEY_FILE"` + ServerName string `yaml:"server_name" env:"WH_CACHE_REDIS_TLS_SERVER_NAME"` + InsecureSkipVerify bool `yaml:"insecure_skip_verify" env:"WH_CACHE_REDIS_TLS_INSECURE_SKIP_VERIFY"` +} + +// hasAddrs reports whether any address is set. An env file's blank +// `WH_CACHE_REDIS_ADDRS=` loads as one empty address, which is none. +func (r CacheRedisConfig) hasAddrs() bool { + return len(r.Addrs) > 1 || len(r.Addrs) == 1 && r.Addrs[0] != "" +} + +func (r CacheRedisConfig) validate() error { + if !r.hasAddrs() { + return errors.New("cache.backend=redis needs cache.redis.addrs (WH_CACHE_REDIS_ADDRS): the server's host:port, or a cluster's seeds, or the sentinels") + } + for _, a := range r.Addrs { + if _, _, err := net.SplitHostPort(a); err != nil { + return fmt.Errorf("cache.redis.addrs (WH_CACHE_REDIS_ADDRS) %q: want host:port: %w", a, err) + } + } + switch r.Mode { + case RedisStandalone, RedisCluster, RedisSentinel: + default: + return fmt.Errorf("cache.redis.mode (WH_CACHE_REDIS_MODE) %q: valid: %s, %s, %s", r.Mode, RedisStandalone, RedisCluster, RedisSentinel) + } + if r.Mode == RedisSentinel && r.SentinelMaster == "" { + return errors.New("cache.redis.mode=sentinel needs cache.redis.sentinel_master (WH_CACHE_REDIS_SENTINEL_MASTER), the master set name") + } + if r.DB < 0 { + return fmt.Errorf("cache.redis.db (WH_CACHE_REDIS_DB) %d is negative", r.DB) + } + if r.Mode == RedisCluster && r.DB != 0 { + return fmt.Errorf("cache.redis.db (WH_CACHE_REDIS_DB) %d: a Redis cluster has only database 0", r.DB) + } + if r.KeyPrefix == "" || strings.ContainsAny(r.KeyPrefix, "{}") { + return fmt.Errorf("cache.redis.key_prefix (WH_CACHE_REDIS_KEY_PREFIX) %q: want a non-empty prefix without a hash-tag brace", r.KeyPrefix) + } + for _, d := range []struct { + key string + v time.Duration + }{ + {"cache.redis.timeout (WH_CACHE_REDIS_TIMEOUT)", r.Timeout}, + {"cache.redis.dial_timeout (WH_CACHE_REDIS_DIAL_TIMEOUT)", r.DialTimeout}, + } { + if d.v <= 0 { + return fmt.Errorf("%s %s must be positive", d.key, d.v) + } + } + if r.VersionTTL < 2*time.Second { + return fmt.Errorf("cache.redis.version_ttl (WH_CACHE_REDIS_VERSION_TTL) %s is under 2s", r.VersionTTL) + } + if r.MaxValueBytes <= 0 { + return fmt.Errorf("cache.redis.max_value_bytes (WH_CACHE_REDIS_MAX_VALUE_BYTES) %d must be positive", r.MaxValueBytes) + } + if r.CompressMinBytes == 0 || r.CompressMinBytes < -1 { + return fmt.Errorf("cache.redis.compress_min_bytes (WH_CACHE_REDIS_COMPRESS_MIN_BYTES) %d: want a positive size, or -1 to never compress", r.CompressMinBytes) + } + if _, err := r.TLS.Config(); err != nil { + return err + } + return nil +} + +// Config builds the tls.Config the block describes, reading its files, or +// nil when TLS is off. A file set while TLS is off is an error rather than +// a silently plaintext connection. +func (t CacheRedisTLS) Config() (*tls.Config, error) { + if !t.Enabled { + if t != (CacheRedisTLS{}) { + return nil, errors.New("cache.redis.tls: files, server_name or insecure_skip_verify are set but cache.redis.tls.enabled (WH_CACHE_REDIS_TLS_ENABLED) is off") + } + return nil, nil + } + if (t.CertFile == "") != (t.KeyFile == "") { + return nil, errors.New("cache.redis.tls: cert_file and key_file must be set together") + } + cfg := &tls.Config{ + MinVersion: tls.VersionTLS12, + ServerName: t.ServerName, + InsecureSkipVerify: t.InsecureSkipVerify, //nolint:gosec // G402: the operator's cache.redis.tls.insecure_skip_verify, warned about at boot + } + if t.CAFile != "" { + pemBytes, err := os.ReadFile(t.CAFile) + if err != nil { + return nil, fmt.Errorf("cache.redis.tls.ca_file: %w", err) + } + pool := x509.NewCertPool() + if !pool.AppendCertsFromPEM(pemBytes) { + return nil, fmt.Errorf("cache.redis.tls.ca_file: no certificates in %s", t.CAFile) + } + cfg.RootCAs = pool + } + if t.CertFile != "" { + cert, err := tls.LoadX509KeyPair(t.CertFile, t.KeyFile) + if err != nil { + return nil, fmt.Errorf("cache.redis.tls.cert_file: %w", err) + } + cfg.Certificates = []tls.Certificate{cert} + } + return cfg, nil +} diff --git a/internal/config/cache_redis_test.go b/internal/config/cache_redis_test.go new file mode 100644 index 00000000..6d989c73 --- /dev/null +++ b/internal/config/cache_redis_test.go @@ -0,0 +1,298 @@ +package config + +import ( + "crypto/ecdsa" + "crypto/elliptic" + "crypto/rand" + "crypto/x509" + "crypto/x509/pkix" + "encoding/pem" + "math/big" + "os" + "path/filepath" + "testing" + "time" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" +) + +// redisBackend is what Load produces for cache.backend=redis with only the +// address set. +func redisBackend() Config { + c := defaultBackends() + c.Cache.Backend = CacheRedis + c.Cache.Redis = CacheRedisConfig{ + Addrs: []string{"redis:6379"}, Mode: RedisStandalone, KeyPrefix: "wh", + Timeout: 100 * time.Millisecond, DialTimeout: time.Second, + MaxValueBytes: 1 << 20, CompressMinBytes: 1 << 10, VersionTTL: 168 * time.Hour, + } + return c +} + +func TestLoad_CacheRedisDefaults(t *testing.T) { + t.Setenv("WH_CACHE_BACKEND", "redis") + t.Setenv("WH_CACHE_REDIS_ADDRS", "redis:6379") + cfg, err := Load("nonexistent.yaml") + require.NoError(t, err) + want := redisBackend() + assert.Equal(t, want.Cache.Redis, cfg.Cache.Redis) + assert.Empty(t, cfg.Warnings()) +} + +func TestLoad_CacheRedisFromEnv(t *testing.T) { + dir := t.TempDir() + caFile, certFile, keyFile := writeTestPKI(t, dir) + for k, v := range map[string]string{ + "WH_CACHE_BACKEND": "redis", + "WH_CACHE_REDIS_ADDRS": "s1:26379,s2:26379", + "WH_CACHE_REDIS_MODE": "sentinel", + "WH_CACHE_REDIS_SENTINEL_MASTER": "mymaster", + "WH_CACHE_REDIS_USERNAME": "wavehouse", + "WH_CACHE_REDIS_PASSWORD": "s3cret", + "WH_CACHE_REDIS_DB": "2", + "WH_CACHE_REDIS_TLS_ENABLED": "true", + "WH_CACHE_REDIS_TLS_CA_FILE": caFile, + "WH_CACHE_REDIS_TLS_CERT_FILE": certFile, + "WH_CACHE_REDIS_TLS_KEY_FILE": keyFile, + "WH_CACHE_REDIS_TLS_SERVER_NAME": "redis.internal", + "WH_CACHE_REDIS_KEY_PREFIX": "staging", + "WH_CACHE_REDIS_TIMEOUT": "250ms", + "WH_CACHE_REDIS_DIAL_TIMEOUT": "3s", + "WH_CACHE_REDIS_MAX_VALUE_BYTES": "2048", + "WH_CACHE_REDIS_COMPRESS_MIN_BYTES": "-1", + "WH_CACHE_REDIS_VERSION_TTL": "24h", + "WH_CACHE_REDIS_TLS_INSECURE_SKIP_VERIFY": "false", + } { + t.Setenv(k, v) + } + cfg, err := Load("nonexistent.yaml") + require.NoError(t, err) + assert.Equal(t, CacheRedisConfig{ + Addrs: []string{"s1:26379", "s2:26379"}, Mode: RedisSentinel, SentinelMaster: "mymaster", + Username: "wavehouse", Password: "s3cret", DB: 2, + TLS: CacheRedisTLS{ + Enabled: true, CAFile: caFile, CertFile: certFile, KeyFile: keyFile, ServerName: "redis.internal", + }, + KeyPrefix: "staging", Timeout: 250 * time.Millisecond, DialTimeout: 3 * time.Second, + MaxValueBytes: 2048, CompressMinBytes: -1, VersionTTL: 24 * time.Hour, + }, cfg.Cache.Redis) + tc, err := cfg.Cache.Redis.TLS.Config() + require.NoError(t, err) + assert.Equal(t, "redis.internal", tc.ServerName) + assert.NotNil(t, tc.RootCAs) + assert.Len(t, tc.Certificates, 1) +} + +func TestLoad_CacheRedisFromYAML(t *testing.T) { + t.Parallel() + path := filepath.Join(t.TempDir(), "config.yaml") + require.NoError(t, os.WriteFile(path, []byte(` +settings: + dir: ./settings +cache: + backend: redis + redis: + addrs: ["n1:6379", "n2:6379"] + mode: cluster + key_prefix: prod + timeout: 50ms + version_ttl: 72h +`), 0o600)) + cfg, err := Load(path) + require.NoError(t, err) + r := cfg.Cache.Redis + assert.Equal(t, CacheRedis, cfg.Cache.Backend) + assert.Equal(t, []string{"n1:6379", "n2:6379"}, r.Addrs) + assert.Equal(t, RedisCluster, r.Mode) + assert.Equal(t, "prod", r.KeyPrefix) + assert.Equal(t, 50*time.Millisecond, r.Timeout) + assert.Equal(t, 72*time.Hour, r.VersionTTL) + assert.Equal(t, time.Second, r.DialTimeout, "an unset key takes its default") + assert.Equal(t, 1024, r.CompressMinBytes) +} + +// A 0 in the file is read as unset, so it takes the default instead of +// switching compression off: why "never" is -1. Pinned so a loader that +// starts honoring the 0 is noticed. +func TestLoad_CacheRedisCompressZeroInYAMLIsTheDefault(t *testing.T) { + t.Parallel() + path := filepath.Join(t.TempDir(), "config.yaml") + require.NoError(t, os.WriteFile(path, []byte(` +settings: + dir: ./settings +cache: + backend: redis + redis: + addrs: ["r:6379"] + compress_min_bytes: 0 +`), 0o600)) + cfg, err := Load(path) + require.NoError(t, err) + assert.Equal(t, 1024, cfg.Cache.Redis.CompressMinBytes) +} + +func TestLoad_CacheRedisRefusesUnknownKeys(t *testing.T) { + t.Parallel() + path := filepath.Join(t.TempDir(), "config.yaml") + require.NoError(t, os.WriteFile(path, []byte(` +settings: + dir: ./settings +cache: + backend: redis + redis: + addr: r:6379 + near_cache: + max_cost: 1 + tls: + ca: /x + memcached: + addrs: ["m:11211"] +`), 0o600)) + _, err := Load(path) + require.Error(t, err) + assert.Contains(t, err.Error(), "cache.memcached, cache.redis.addr, cache.redis.near_cache, cache.redis.tls.ca") +} + +// The documented env file lists WH_CACHE_REDIS_ADDRS blank: that is no +// address, not an unread block to warn about, and not a valid redis one. +func TestLoad_CacheRedisBlankAddrs(t *testing.T) { + t.Setenv("WH_CACHE_REDIS_ADDRS", "") + cfg, err := Load("nonexistent.yaml") + require.NoError(t, err) + assert.Empty(t, cfg.Warnings()) + + t.Setenv("WH_CACHE_BACKEND", "redis") + _, err = Load("nonexistent.yaml") + require.ErrorContains(t, err, "cache.backend=redis needs cache.redis.addrs") +} + +func TestUnboundEnv_KnowsTheCacheRedisVariables(t *testing.T) { + t.Parallel() + assert.Empty(t, unboundEnv([]string{ + "WH_CACHE_REDIS_ADDRS=r:6379", "WH_CACHE_REDIS_PASSWORD=x", "WH_CACHE_REDIS_TLS_CA_FILE=/ca.pem", + "WH_CACHE_REDIS_VERSION_TTL=1h", "WH_CACHE_REDIS_COMPRESS_MIN_BYTES=-1", + })) + assert.Equal(t, []string{"WH_CACHE_REDIS_ADDR"}, unboundEnv([]string{"WH_CACHE_REDIS_ADDR=r:6379"})) +} + +func TestValidate_CacheRedis(t *testing.T) { + t.Parallel() + dir := t.TempDir() + caFile, certFile, keyFile := writeTestPKI(t, dir) + notPEM := filepath.Join(dir, "not.pem") + require.NoError(t, os.WriteFile(notPEM, []byte("hello"), 0o600)) + cases := []struct { + name string + set func(*CacheRedisConfig) + want string // "" = valid + }{ + {"defaults", func(*CacheRedisConfig) {}, ""}, + {"no addrs", func(r *CacheRedisConfig) { r.Addrs = nil }, "cache.backend=redis needs cache.redis.addrs (WH_CACHE_REDIS_ADDRS)"}, + {"addr without port", func(r *CacheRedisConfig) { r.Addrs = []string{"redis"} }, `cache.redis.addrs (WH_CACHE_REDIS_ADDRS) "redis": want host:port`}, + {"mode", func(r *CacheRedisConfig) { r.Mode = "replica" }, `cache.redis.mode (WH_CACHE_REDIS_MODE) "replica": valid: standalone, cluster, sentinel`}, + {"sentinel without master", func(r *CacheRedisConfig) { r.Mode = RedisSentinel }, "needs cache.redis.sentinel_master"}, + {"sentinel", func(r *CacheRedisConfig) { r.Mode, r.SentinelMaster = RedisSentinel, "m" }, ""}, + {"cluster db", func(r *CacheRedisConfig) { r.Mode, r.DB = RedisCluster, 1 }, "a Redis cluster has only database 0"}, + {"standalone db", func(r *CacheRedisConfig) { r.DB = 3 }, ""}, + {"negative db", func(r *CacheRedisConfig) { r.DB = -1 }, "is negative"}, + {"empty prefix", func(r *CacheRedisConfig) { r.KeyPrefix = "" }, "cache.redis.key_prefix"}, + {"brace prefix", func(r *CacheRedisConfig) { r.KeyPrefix = "a{b}" }, "hash-tag brace"}, + {"zero timeout", func(r *CacheRedisConfig) { r.Timeout = 0 }, "cache.redis.timeout (WH_CACHE_REDIS_TIMEOUT) 0s must be positive"}, + {"negative dial timeout", func(r *CacheRedisConfig) { r.DialTimeout = -time.Second }, "cache.redis.dial_timeout"}, + {"short version ttl", func(r *CacheRedisConfig) { r.VersionTTL = time.Second }, "cache.redis.version_ttl (WH_CACHE_REDIS_VERSION_TTL) 1s is under 2s"}, + {"zero max value", func(r *CacheRedisConfig) { r.MaxValueBytes = 0 }, "cache.redis.max_value_bytes"}, + {"compress 0", func(r *CacheRedisConfig) { r.CompressMinBytes = 0 }, "or -1 to never compress"}, + {"compress -2", func(r *CacheRedisConfig) { r.CompressMinBytes = -2 }, "or -1 to never compress"}, + {"compress never", func(r *CacheRedisConfig) { r.CompressMinBytes = -1 }, ""}, + {"tls files while off", func(r *CacheRedisConfig) { r.TLS.CAFile = caFile }, "cache.redis.tls.enabled (WH_CACHE_REDIS_TLS_ENABLED) is off"}, + {"tls system roots", func(r *CacheRedisConfig) { r.TLS.Enabled = true }, ""}, + {"tls full", func(r *CacheRedisConfig) { + r.TLS = CacheRedisTLS{Enabled: true, CAFile: caFile, CertFile: certFile, KeyFile: keyFile} + }, ""}, + {"tls cert without key", func(r *CacheRedisConfig) { r.TLS = CacheRedisTLS{Enabled: true, CertFile: certFile} }, "cert_file and key_file must be set together"}, + {"tls missing ca", func(r *CacheRedisConfig) { + r.TLS = CacheRedisTLS{Enabled: true, CAFile: filepath.Join(dir, "missing.pem")} + }, "cache.redis.tls.ca_file"}, + {"tls ca not pem", func(r *CacheRedisConfig) { r.TLS = CacheRedisTLS{Enabled: true, CAFile: notPEM} }, "no certificates in"}, + {"tls bad pair", func(r *CacheRedisConfig) { + r.TLS = CacheRedisTLS{Enabled: true, CertFile: certFile, KeyFile: notPEM} + }, "cache.redis.tls.cert_file"}, + } + for _, tc := range cases { + t.Run(tc.name, func(t *testing.T) { + t.Parallel() + cfg := redisBackend() + tc.set(&cfg.Cache.Redis) + err := cfg.Validate() + if tc.want == "" { + require.NoError(t, err) + return + } + require.Error(t, err) + assert.Contains(t, err.Error(), tc.want) + }) + } +} + +// The redis block is read only when selected: an invalid one under +// backend=local does not refuse boot, it warns that it is ignored. +func TestValidate_CacheRedisIgnoredUnlessSelected(t *testing.T) { + t.Parallel() + cfg := defaultBackends() + cfg.Cache.Redis.Addrs = []string{"no-port"} + require.NoError(t, cfg.Validate()) + got := cfg.Warnings() + require.Len(t, got, 1) + assert.Contains(t, got[0], "cache.redis.addrs is set but cache.backend is local") +} + +func TestWarnings_CacheRedis(t *testing.T) { + t.Parallel() + cfg := redisBackend() + assert.Empty(t, cfg.Warnings()) + cfg.Cache.Redis.TLS = CacheRedisTLS{Enabled: true, InsecureSkipVerify: true} + require.NoError(t, cfg.Validate()) + got := cfg.Warnings() + require.Len(t, got, 1) + assert.Contains(t, got[0], "cache.redis.tls.insecure_skip_verify is on") + + // A shared cache clears the shared-queue warning about a local one. + cfg = redisBackend() + cfg.MQ.Backend = "shared" + got = cfg.Warnings() + require.Len(t, got, 1) + assert.Contains(t, got[0], "dedupe.backend=pebble") +} + +// writeTestPKI writes a self-signed authority and a client certificate it +// signed, returning the three paths cache.redis.tls names. +func writeTestPKI(t *testing.T, dir string) (caFile, certFile, keyFile string) { + t.Helper() + caKey, err := ecdsa.GenerateKey(elliptic.P256(), rand.Reader) + require.NoError(t, err) + ca := &x509.Certificate{ + SerialNumber: big.NewInt(1), Subject: pkix.Name{CommonName: "test ca"}, + NotBefore: time.Now().Add(-time.Hour), NotAfter: time.Now().Add(time.Hour), + IsCA: true, BasicConstraintsValid: true, KeyUsage: x509.KeyUsageCertSign, + } + caDER, err := x509.CreateCertificate(rand.Reader, ca, ca, &caKey.PublicKey, caKey) + require.NoError(t, err) + leafKey, err := ecdsa.GenerateKey(elliptic.P256(), rand.Reader) + require.NoError(t, err) + leaf := &x509.Certificate{ + SerialNumber: big.NewInt(2), Subject: pkix.Name{CommonName: "wavehouse"}, + NotBefore: time.Now().Add(-time.Hour), NotAfter: time.Now().Add(time.Hour), + KeyUsage: x509.KeyUsageDigitalSignature, ExtKeyUsage: []x509.ExtKeyUsage{x509.ExtKeyUsageClientAuth}, + } + leafDER, err := x509.CreateCertificate(rand.Reader, leaf, ca, &leafKey.PublicKey, caKey) + require.NoError(t, err) + keyDER, err := x509.MarshalECPrivateKey(leafKey) + require.NoError(t, err) + write := func(name, typ string, der []byte) string { + path := filepath.Join(dir, name) + require.NoError(t, os.WriteFile(path, pem.EncodeToMemory(&pem.Block{Type: typ, Bytes: der}), 0o600)) + return path + } + return write("ca.pem", "CERTIFICATE", caDER), write("client.pem", "CERTIFICATE", leafDER), write("client.key", "EC PRIVATE KEY", keyDER) +} diff --git a/scripts/orchestrator/main.go b/scripts/orchestrator/main.go index a128d402..10eacba8 100644 --- a/scripts/orchestrator/main.go +++ b/scripts/orchestrator/main.go @@ -1,10 +1,12 @@ // E2E orchestrator — drives a clean, isolated E2E test session against -// one ClickHouse + one WaveHouse, then runs the vitest suite once. +// one ClickHouse + one Redis + one WaveHouse, then runs the vitest suite +// once. // // Lifecycle: // -// 1. Start ClickHouse via testcontainers-go (random host ports — no -// conflict with `make dev` or other compose stacks). +// 1. Start ClickHouse and Redis (the fixture's shared cache) via +// testcontainers-go (random host ports — no conflict with `make dev` or +// other compose stacks). // 2. Pick a random free TCP port on 127.0.0.1 and start bin/wavehouse-cov // bound to it (WH_SERVER_PORT) with auth enabled. Random port avoids // conflicts with `make dev`, dev servers, and previous runs that may @@ -44,6 +46,7 @@ import ( "syscall" "time" + "github.com/moby/moby/api/types/container" "github.com/testcontainers/testcontainers-go" "github.com/testcontainers/testcontainers-go/wait" ) @@ -181,6 +184,43 @@ func run() error { return fmt.Errorf("settings dir: %w", err) } + // The shared cache the fixture's cache.backend=redis names: the suite + // runs the backend a multi-instance deployment runs. No persistence, and + // /data on tmpfs so the image's VOLUME leaves no anonymous volume behind. + log.Println("→ starting Redis testcontainer...") + redis, err := testcontainers.GenericContainer(ctx, testcontainers.GenericContainerRequest{ + ContainerRequest: testcontainers.ContainerRequest{ + // Pinned to match internal/cache's integration suite. + Image: "redis:8.10.2-alpine", + Cmd: []string{"redis-server", "--save", "", "--appendonly", "no"}, + ExposedPorts: []string{"6379/tcp"}, + HostConfigModifier: func(hc *container.HostConfig) { + hc.Tmpfs = map[string]string{"/data": ""} + }, + WaitingFor: wait.ForLog("Ready to accept connections").WithStartupTimeout(60 * time.Second), + }, + Started: true, + }) + if err != nil { + return fmt.Errorf("redis start: %w", err) + } + defer func() { + log.Println("→ terminating Redis testcontainer...") + if err := redis.Terminate(context.Background()); err != nil { + log.Printf(" redis terminate: %v", err) + } + }() + redisHost, err := redis.Host(ctx) + if err != nil { + return fmt.Errorf("redis host: %w", err) + } + redisPort, err := redis.MappedPort(ctx, "6379") + if err != nil { + return fmt.Errorf("redis port: %w", err) + } + redisAddr := net.JoinHostPort(redisHost, redisPort.Port()) + log.Printf("✓ Redis ready: %s", redisAddr) + whPort, err := pickFreePort(ctx) if err != nil { return fmt.Errorf("pick free port: %w", err) @@ -198,14 +238,15 @@ func run() error { // tests/e2e/fixtures/config.yaml and the tunables, policy, roles, and // pipes in tests/e2e/fixtures/settings — edit them there, not here. The // vars below are the per-run dynamic overrides (port, scratch paths, the - // patched settings copy) plus GOCOVERDIR and WH_CONFIG, which can't live - // in YAML. + // patched settings copy, the Redis address) plus GOCOVERDIR and + // WH_CONFIG, which can't live in YAML. whCmd.Env = append(os.Environ(), "GOCOVERDIR="+coverDir, "WH_CONFIG="+filepath.Join(repoRoot, "tests", "e2e", "fixtures", "config.yaml"), "WH_SERVER_PORT="+strconv.Itoa(whPort), "WH_SETTINGS_DIR="+settingsDir, "WH_DATA_DIR="+dataDir, + "WH_CACHE_REDIS_ADDRS="+redisAddr, ) if os.Getenv("OTEL_EXPORTER_OTLP_ENDPOINT") == "" { diff --git a/tests/e2e/fixtures/config.yaml b/tests/e2e/fixtures/config.yaml index 4ce30280..69e8cb33 100644 --- a/tests/e2e/fixtures/config.yaml +++ b/tests/e2e/fixtures/config.yaml @@ -1,8 +1,9 @@ # WaveHouse config for the e2e harness (scripts/orchestrator + tests/e2e). # -# Dynamic values — ClickHouse addr/HTTP port (testcontainer), WaveHouse -# server port (free port), WH_DATA_DIR (per-run scratch) — are injected by -# the orchestrator via env vars. Everything below is pinned here so the +# Dynamic values — ClickHouse addr/HTTP port (testcontainer), the Redis +# address (testcontainer, WH_CACHE_REDIS_ADDRS), WaveHouse server port (free +# port), WH_DATA_DIR (per-run scratch) — are injected by the orchestrator via +# env vars. Everything below is pinned here so the # rig config is visible and editable without recompiling Go. # JWT validation. The suite signs test tokens with this fixed dev secret; @@ -13,6 +14,15 @@ auth: # operator_key: "" # non-JWT full-access operator credential (Authorization: Operator, or X-Operator-Key); unset in the rig +# The shared cache, as a multi-instance deployment runs it; the in-process +# backend is covered by the unit and integration suites. The timeout is ten +# times the default so a loaded runner's slow round trip is not a bypassed +# lookup that a HIT assertion reads as a failure. +cache: + backend: redis + redis: + timeout: 1s + # Tenant tunables (schema refresh_interval 5s so schema-discovery tests don't # wait a minute; dedupe enabled + id_field; DLQ on; CORS "*"; a 1 GiB NATS # stream budget so the testcontainer stays tiny), the policy, and the pipes diff --git a/tests/integration/shared_cache_test.go b/tests/integration/shared_cache_test.go new file mode 100644 index 00000000..b665ccb6 --- /dev/null +++ b/tests/integration/shared_cache_test.go @@ -0,0 +1,271 @@ +//go:build integration + +package tests + +import ( + "context" + "fmt" + "io" + "net" + "net/http" + "net/url" + "strings" + "sync/atomic" + "testing" + "time" + + "github.com/moby/moby/api/types/container" + "github.com/moby/moby/client" + "github.com/redis/rueidis" + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + "github.com/testcontainers/testcontainers-go" + "github.com/testcontainers/testcontainers-go/wait" + + "github.com/Wave-RF/WaveHouse/internal/app" + "github.com/Wave-RF/WaveHouse/internal/config" +) + +// Pinned, as internal/cache's integration suite pins it. +const redisImage = "redis:8.10.2-alpine" + +// minCacheTTL is cache.QueryTimeToTTL's floor: a fill made less than this +// long ago cannot have expired, so a miss inside it is an invalidation. +const minCacheTTL = 10 * time.Second + +var cachePrefixes atomic.Uint64 + +// startRedis runs a throwaway Redis with no persistence, its /data on tmpfs +// so the image's VOLUME leaves no anonymous volume behind. +func startRedis(t *testing.T) (testcontainers.Container, string) { + t.Helper() + ctx := context.Background() + ctr, err := testcontainers.GenericContainer(ctx, testcontainers.GenericContainerRequest{ + ContainerRequest: testcontainers.ContainerRequest{ + Image: redisImage, + Cmd: []string{"redis-server", "--save", "", "--appendonly", "no"}, + ExposedPorts: []string{"6379/tcp"}, + HostConfigModifier: func(hc *container.HostConfig) { + hc.Tmpfs = map[string]string{"/data": ""} + }, + WaitingFor: wait.ForLog("Ready to accept connections").WithStartupTimeout(90 * time.Second), + }, + Started: true, + }) + testcontainers.CleanupContainer(t, ctr) + require.NoError(t, err) + host, err := ctr.Host(ctx) + require.NoError(t, err) + port, err := ctr.MappedPort(ctx, "6379/tcp") + require.NoError(t, err) + return ctr, net.JoinHostPort(host, port.Port()) +} + +// bootRedisApp runs a second, independent WaveHouse — its own embedded NATS, +// ingest worker and data_dir — against the suite's ClickHouse, with +// cache.backend=redis. It returns the instance's base URL. +func bootRedisApp(t *testing.T, redisAddr, prefix string, timeout time.Duration) string { + t.Helper() + e := env(t) + ctx := context.Background() + settingsDir, err := writeTestSettings(e.ch) + require.NoError(t, err) + var lc net.ListenConfig + ln, err := lc.Listen(ctx, "tcp", "127.0.0.1:0") + require.NoError(t, err) + cfg := &config.Config{ + DataDir: t.TempDir(), + Server: config.Server{ShutdownTimeout: 10}, + ClickHouse: config.ClickHouse{Password: testCHPassword}, + MQ: config.MQ{Backend: config.MQEmbedded}, + Cache: config.Cache{Backend: config.CacheRedis, Redis: config.CacheRedisConfig{ + Addrs: []string{redisAddr}, Mode: config.RedisStandalone, KeyPrefix: prefix, + Timeout: timeout, DialTimeout: 2 * time.Second, + MaxValueBytes: 1 << 20, CompressMinBytes: 1 << 10, VersionTTL: time.Hour, + }}, + Dedupe: config.Dedupe{Backend: config.DedupePebble}, + Coord: config.Coord{Backend: config.CoordLocal}, + Settings: config.Settings{Dir: settingsDir}, + } + a, err := app.New(ctx, app.Options{Config: cfg, Listener: ln}) + require.NoError(t, err) + runCtx, stop := context.WithCancel(ctx) + runDone := make(chan error, 1) + go func() { runDone <- a.Run(runCtx) }() + t.Cleanup(func() { + stop() + assert.NoError(t, <-runDone) + closeCtx, cancel := context.WithTimeout(context.Background(), 10*time.Second) + defer cancel() + assert.NoError(t, a.Close(closeCtx)) + }) + baseURL := "http://" + ln.Addr().String() + require.NoError(t, waitForLive(ctx, baseURL, 30*time.Second)) + return baseURL +} + +// structuredQuery posts a select-all structured query and returns the +// status, the X-Cache header and the body. +func structuredQuery(t *testing.T, baseURL, table string) (int, string, string) { + t.Helper() + status, xc, body, err := tryStructuredQuery(baseURL, table) + require.NoError(t, err) + return status, xc, body +} + +// tryStructuredQuery is structuredQuery for an Eventually condition, which +// runs off the test goroutine and so must not call require. +func tryStructuredQuery(baseURL, table string) (int, string, string, error) { + req, err := http.NewRequestWithContext(context.Background(), http.MethodPost, + baseURL+"/v1/query?table="+url.QueryEscape(table), strings.NewReader(`{"select_all":true}`)) + if err != nil { + return 0, "", "", err + } + req.Header.Set("Content-Type", "application/json") + resp, err := http.DefaultClient.Do(req) + if err != nil { + return 0, "", "", err + } + defer func() { _ = resp.Body.Close() }() + body, err := io.ReadAll(resp.Body) + return resp.StatusCode, resp.Header.Get("X-Cache"), string(body), err +} + +func ingestRow(t *testing.T, baseURL, table, user string) { + t.Helper() + resp, err := http.Post(baseURL+"/v1/ingest?table="+url.QueryEscape(table), "application/json", + strings.NewReader(fmt.Sprintf(`{"user_id":%q,"value":1}`, user))) + require.NoError(t, err) + defer func() { _ = resp.Body.Close() }() + require.Equal(t, http.StatusOK, resp.StatusCode) +} + +// Two WaveHouse processes share one Redis and one ClickHouse. A result one +// fills is a hit for the other, and an insert one process's worker makes +// invalidates what the other cached: the other's next query is a miss that +// returns the new row, well inside the TTL the stale entry was filed with. +func TestSharedCache_IngestOnOneInstanceInvalidatesAnother(t *testing.T) { + table := createTable(t, "user_id String, value Float64", "ORDER BY user_id") + _, redisAddr := startRedis(t) + prefix := fmt.Sprintf("it%d", cachePrefixes.Add(1)) + a := bootRedisApp(t, redisAddr, prefix, 2*time.Second) + b := bootRedisApp(t, redisAddr, prefix, 2*time.Second) + + rc, err := rueidis.NewClient(rueidis.ClientOption{InitAddress: []string{redisAddr}, DisableCache: true, ForceSingleClient: true}) + require.NoError(t, err) + t.Cleanup(rc.Close) + tableToken := func() string { + v, err := rc.Do(context.Background(), rc.B().Get().Key(prefix+":{0}:B:"+table).Build()).ToString() + if rueidis.IsRedisNil(err) { + return "" + } + require.NoError(t, err) + return v + } + + status, xc, body := structuredQuery(t, b, table) + require.Equal(t, http.StatusOK, status, body) + require.Equal(t, "MISS", xc) + status, xc, _ = structuredQuery(t, a, table) + require.Equal(t, http.StatusOK, status) + require.Equal(t, "HIT", xc, "a fills, b hits: one cache") + + // An insert's worker and its invalidation run on whichever process took + // the ingest, so the discriminating window is the stale entry's TTL: if + // the batch window and load push the new row past it, the round proves + // nothing and runs again with a fresh fill. + for round := 1; ; round++ { + user := fmt.Sprintf("user-%d", round) + status, xc, body = structuredQuery(t, b, table) + require.Equal(t, http.StatusOK, status, body) + require.NotContains(t, body, user) + filled := time.Now() + if xc == "HIT" { + // The previous round's fill: refresh it so the TTL window starts now. + _, err := rc.Do(context.Background(), rc.B().Flushdb().Build()).ToString() + require.NoError(t, err) + status, xc, body = structuredQuery(t, b, table) + require.Equal(t, http.StatusOK, status, body) + require.Equal(t, "MISS", xc) + filled = time.Now() + } + before := tableToken() + + ingestRow(t, a, table, user) + var seenAt time.Time + require.Eventually(t, func() bool { + var err error + status, xc, body, err = tryStructuredQuery(b, table) + if err == nil && status == http.StatusOK && strings.Contains(body, user) { + seenAt = time.Now() + return true + } + return false + }, 30*time.Second, 100*time.Millisecond, "b never served the row ingested through a") + assert.NotEqual(t, before, tableToken(), "a's worker bumped the table token in the shared server") + if seenAt.Sub(filled) < minCacheTTL-time.Second { + assert.Equal(t, "MISS", xc, "the first answer carrying the new row is b's refill") + status, xc, _ = structuredQuery(t, b, table) + require.Equal(t, http.StatusOK, status) + assert.Equal(t, "HIT", xc, "b's refill is cached again") + return + } + require.Less(t, round, 3, "b served the new row only once its stale entry could have expired, in every round: the invalidation never reached it, or ingest is too slow here to tell") + t.Logf("round %d: row landed %s after the fill, past the TTL floor; retrying", round, seenAt.Sub(filled)) + } +} + +// A Redis that stops answering costs queries nothing but the cache: they +// keep succeeding, straight from ClickHouse, each a miss; an ingest made +// meanwhile is visible at once. Once it answers again, the cache serves hits. +func TestSharedCache_RedisDownQueriesBypass(t *testing.T) { + ctx := context.Background() + table := createTable(t, "user_id String, value Float64", "ORDER BY user_id") + ctr, redisAddr := startRedis(t) + const timeout = 200 * time.Millisecond + a := bootRedisApp(t, redisAddr, fmt.Sprintf("it%d", cachePrefixes.Add(1)), timeout) + + status, xc, _ := structuredQuery(t, a, table) + require.Equal(t, http.StatusOK, status) + require.Equal(t, "MISS", xc) + status, xc, _ = structuredQuery(t, a, table) + require.Equal(t, http.StatusOK, status) + require.Equal(t, "HIT", xc) + + d, err := testcontainers.NewDockerClientWithOpts(ctx) + require.NoError(t, err) + t.Cleanup(func() { _ = d.Close() }) + _, err = d.ContainerPause(ctx, ctr.GetContainerID(), client.ContainerPauseOptions{}) + require.NoError(t, err) + paused := true + unpause := func() { + if paused { + paused = false + _, err := d.ContainerUnpause(ctx, ctr.GetContainerID(), client.ContainerUnpauseOptions{}) + require.NoError(t, err) + } + } + t.Cleanup(unpause) + + for range 8 { + start := time.Now() + status, xc, body := structuredQuery(t, a, table) + require.Equal(t, http.StatusOK, status, body) + assert.Equal(t, "MISS", xc) + // A lookup and a fill each wait at most the timeout; the query itself + // is a few ms. Generous for -race under load. + assert.Less(t, time.Since(start), 5*timeout+2*time.Second) + } + + ingestRow(t, a, table, "while-down") + require.Eventually(t, func() bool { + status, _, body, err := tryStructuredQuery(a, table) + return err == nil && status == http.StatusOK && strings.Contains(body, "while-down") + }, 30*time.Second, 200*time.Millisecond, "a row ingested while the cache is down is served") + + unpause() + require.Eventually(t, func() bool { + _, xc, body, err := tryStructuredQuery(a, table) + return err == nil && xc == "HIT" && strings.Contains(body, "while-down") + }, 30*time.Second, 200*time.Millisecond, "the cache serves hits again once the server answers") +} From 13c6cb3014d211ed9f41f47c98ad49746599d0b4 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 03:44:18 -0400 Subject: [PATCH 07/27] =?UTF-8?q?fix(cache):=20review=20round=20=E2=80=94?= =?UTF-8?q?=20e2e=20gate,=20trust=20boundary,=20#386,=20addr=20spaces?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Exclude LocalCache from the e2e gate now that e2e runs on Redis; document the shared server as a trust boundary, the noeviction and mutation-pipe (#386) staleness cases, and the connection commands an ACL user needs; refuse addresses with surrounding spaces. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- .testcoverage.yml | 5 +++++ docs/src/content/docs/api.md | 6 +++--- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/configuration.mdx | 6 +++--- docs/src/content/docs/deployment.md | 12 +++++++++++- internal/config/cache_redis.go | 6 ++++-- internal/config/cache_redis_test.go | 1 + 7 files changed, 28 insertions(+), 10 deletions(-) diff --git a/.testcoverage.yml b/.testcoverage.yml index aff1a694..59ff6d0b 100644 --- a/.testcoverage.yml +++ b/.testcoverage.yml @@ -73,3 +73,8 @@ exclude: - ^internal/settings/ - ^cmd/wavehouse/validate\.go$ - ^cmd/wavehouse/bootstrap\.go$ + # The in-process cache backend: the e2e stack runs cache.backend=redis + # (#613), so the binary carries LocalCache and its version index but e2e + # never reaches them. The unit suite and the integration suite's main + # app (cache.backend=local) cover them; the merged total still counts them. + - ^internal/cache/(local|version_manager)\.go$ diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index c82739c0..14f5d130 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -422,7 +422,7 @@ ClickHouse's inline `FORMAT` clause (e.g. `SELECT 1 FORMAT CSV` or `… FORMAT P The proxy buffers the upstream response in memory before forwarding (no row-streaming yet), so a `SELECT *` from a large table can pin RAM on the API server. To avoid an admin OOMing themselves, responses larger than 64 MiB return 502 with a `clickhouse response exceeded N bytes` error. Narrow the query with `LIMIT`, or use a streaming client outside WaveHouse that talks to ClickHouse directly (the standard escape hatch — the same admin credentials work). ::: -This endpoint **does not cache, does not singleflight, and emits `Cache-Control: no-store`** — every request goes straight to ClickHouse, mutation or read, and downstream HTTP caches are explicitly told not to store the response. Raw SQL is an admin escape hatch with infrequent, ad-hoc traffic, so the L1/singleflight machinery would only add complexity without a real hit-rate win. Use [`POST /v1/query?table={table}`](#post-v1querytabletable--structured-query) or [`GET/POST /v1/pipes/{name}`](#getpost-v1pipesname--execute-named-pipe) for the cached read paths (dashboards, high-QPS clients, etc.) — both share an in-process L1 (Ristretto) with singleflight coalescing. +This endpoint **does not cache, does not singleflight, and emits `Cache-Control: no-store`** — every request goes straight to ClickHouse, mutation or read, and downstream HTTP caches are explicitly told not to store the response. Raw SQL is an admin escape hatch with infrequent, ad-hoc traffic, so the L1/singleflight machinery would only add complexity without a real hit-rate win. Use [`POST /v1/query?table={table}`](#post-v1querytabletable--structured-query) or [`GET/POST /v1/pipes/{name}`](#getpost-v1pipesname--execute-named-pipe) for the cached read paths (dashboards, high-QPS clients, etc.) — both go through the query cache ([`cache.backend`](/configuration#backends): in-process, or a Redis shared by every instance) with singleflight coalescing. :::note[Admin only] The route is mounted under `/v1/ops/*`, behind the `RequireAdmin` gate: only a caller whose JWT role equals the policy `admin_role` (`"admin"` by default) — or who presents the non-JWT [operator key](#authentication) — may use it. A tokenless request (or a valid token without a role claim) resolves to the `default_role` (not the admin role unless `default_role` is deliberately set to it — a loudly-warned dev-only setting) and is rejected with `403`; a present-but-invalid token — expired, malformed, bad signature — keeps its stashed verification error and fails loud with `401` instead. Raw SQL has no per-statement scope check (a full SQL parser would be needed to authorize predicates), so the role gate is the entire authorization story, shared with the rest of `/v1/ops/*` (see [Admin Endpoints](#admin-endpoints)). The normal surfaces for non-admin callers are `POST /v1/ingest?table={table}` for writes, `POST /v1/query?table={table}` for structured reads, and `GET/POST /v1/pipes/{name}` for pre-defined queries — none of which expose raw SQL. @@ -538,7 +538,7 @@ Table, column, and alias names may contain any characters ClickHouse accepts — **Response:** -JSON array of result rows. Top-level `DateTime`/`DateTime64` values are returned in canonical RFC 3339 UTC (`2026-06-21T04:00:00.123Z`) — `Nullable` timestamp columns included (a SQL `NULL` renders as JSON `null`), while timestamps nested inside `Array`/`Map`/`Tuple` columns are rendered in the column's declared zone, else the ClickHouse server's, as the driver returns them — byte-identical to the [SSE stream](#get-v1stream--server-sent-events-stream) for values [canonicalized at ingest](#timestamp-canonicalization) (a fail-open pass-through that ClickHouse accepted still comes back canonical here, though it streamed in the producer's spelling). The response carries an `X-Cache: HIT` or `X-Cache: MISS` header — this endpoint shares the in-process L1 (Ristretto) + singleflight machinery (unlike `/v1/ops/query`, which always hits ClickHouse), keyed by [tenant](/deployment#multi-tenant-deployments): a request is never served from, or coalesced with, another tenant's. +JSON array of result rows. Top-level `DateTime`/`DateTime64` values are returned in canonical RFC 3339 UTC (`2026-06-21T04:00:00.123Z`) — `Nullable` timestamp columns included (a SQL `NULL` renders as JSON `null`), while timestamps nested inside `Array`/`Map`/`Tuple` columns are rendered in the column's declared zone, else the ClickHouse server's, as the driver returns them — byte-identical to the [SSE stream](#get-v1stream--server-sent-events-stream) for values [canonicalized at ingest](#timestamp-canonicalization) (a fail-open pass-through that ClickHouse accepted still comes back canonical here, though it streamed in the producer's spelling). The response carries an `X-Cache: HIT` or `X-Cache: MISS` header — this endpoint shares the query cache + singleflight machinery (unlike `/v1/ops/query`, which always hits ClickHouse), keyed by [tenant](/deployment#multi-tenant-deployments): a request is never served from, or coalesced with, another tenant's. The inbound request body is capped at 1 MiB; a body over the cap is rejected with `413`. A query AST is bounded by nature (far under 1 MiB even with a large `in`-list), and the cap blocks a single-request memory-exhaustion vector on this public endpoint. Set a tighter or higher outer limit at your [reverse proxy](/reverse-proxy#request-body-size-limits) — but it can only narrow the effective limit, not raise it past this cap. @@ -575,7 +575,7 @@ Executes a pre-defined named query (pipe) with parameter binding. Parameters can **Response:** -JSON array of result rows, with `X-Cache: HIT` or `X-Cache: MISS` indicating whether the row came from the in-process L1. +JSON array of result rows, with `X-Cache: HIT` or `X-Cache: MISS` indicating whether the rows came from the query cache. The POST parameter body is capped at 1 MiB; a body over the cap is rejected with `413` (the same 1 MiB parameter/AST-body cap as [`POST /v1/query`](#post-v1querytabletable--structured-query) — see [reverse proxy → body limits](/reverse-proxy#request-body-size-limits)). A malformed-but-within-cap body is ignored rather than rejected, since parameters may legitimately come from the query string alone. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 6c59e5f1..a16d01b9 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -112,7 +112,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on, each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope: one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)). `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. - **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: one version per tenant, per (tenant, table) and per (tenant, table, scope), each keyed by its name alone and bumped in place, so the index holds one entry per live tenant, table and scope however often each is bumped ([#262](https://github.com/Wave-RF/WaveHouse/issues/262)). A query key folds the tenant's version and, for each dependency, its tenant's, table's and scope's, so bumping a table (a scopeless write) orphans every scope of it, and bumping one scope orphans that scope and the whole-table view — scope is reserved and empty today, so every write is the whole-table bump — all without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results. A tenant's version is a *generation*, unique within the process and handed out by the first key built for the tenant; `BumpTenant` (behind `InvalidateTenant`) drops the tenant's whole index, so the next key gets a fresh generation no cached entry folds, orphaning every cached result of the tenant in one step — a pipe result with no dependencies, and a table no bump ever keyed, included — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). `LocalCache.Prune` does the same for every tenant no longer served, which `internal/app` runs after each settings reload, so a tenant removed or rejected stops holding its index. A table bump drops the table's scope versions with it, since every key they were folded into also folds the old table version; and a bump of a tenant with no index is a no-op, since no key folds its next generation yet. The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. -- **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET`, `MGET` and `PING`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token (a string of another length, a list, a hash) is replaced, so a token lost to eviction, expiry or a restart is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing. `cache.backend: redis` selects it: `internal/app`'s `wireCache` maps the boot config's `cache.redis` block onto `RedisConfig`, reading the TLS files, and releases it with the other components. +- **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — its data commands are only `GET`, `SET`, `MGET` and `PING`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token (a string of another length, a list, a hash) is replaced, so a token lost to eviction, expiry or a restart is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing. `cache.backend: redis` selects it: `internal/app`'s `wireCache` maps the boot config's `cache.redis` block onto `RedisConfig`, reading the TLS files, and releases it with the other components. - **redis_codec.go** — the key schema (`tokenKeys`, `bumpKeys`, `valueKey`) and the value frame: format, flags, expiry, the token list, the payload, zstd-compressed from `CompressMinBytes` when that is smaller. A stored value is capped at `MaxValueBytes` and a decoded one at eight times that, which refuses a zip bomb planted in a shared server; an unknown format, as a newer process writes during a rolling upgrade, is a miss. - **breaker.go** — the circuit breaker: `BreakerThreshold` (5) consecutive failures open it, and while open every operation skips the server at once; after `BreakerOpenFor` (5 s) one background `PING` decides whether it closes. Only a transport failure or a timeout counts against the server: any reply, an error reply or one the backend cannot use included, counts as a success, for operations and the probe alike, and a caller that gave up first counts as nothing. - **pending.go** — invalidations the server did not take, kept per token key (repeats coalesce) and retried by a background loop with backoff from 100 ms to 10 s until they land. Landing late is still correct: a fresh token orphans the pre-write entries and any fill made meanwhile. Past `PendingMax` (100,000) keys the set collapses to one tenant bump per affected tenant. `Close` makes one last attempt at the bumps still owed, past the breaker (an open one is why they are owed); what that attempt cannot deliver is lost, and the entries they would orphan are served until their TTL — the same failure as a worker stopping between an insert and its invalidation. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 82744a35..f03a2854 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -122,11 +122,11 @@ Each tenant's queue has its own disk budget, `mq.max_bytes_gb`, a hot-reloadable | --- | --- | ------- | ----------- | | `cache.l1_max_cost` | `WH_CACHE_L1_MAX_COST` | `67108864` | Maximum size in bytes (~64 MB) of the `local` backend's in-process cache. The time-range bucket structured queries normalize to is `query.timestamp_bucket_seconds` in the [Settings Directory](/settings-directory#configjson-keys). | -The `redis` backend's settings, read only when `cache.backend` is `redis`. It runs on Redis, Valkey, Dragonfly, ElastiCache (including Serverless) and MemoryDB: it sends only `GET`, `SET`, `MGET` and `PING`. [Deployment](/deployment#multiple-instances-and-the-shared-cache) covers sizing, `maxmemory-policy` and what a reader on another instance can see. +The `redis` backend's settings, read only when `cache.backend` is `redis`. It is tested on Redis, Valkey, Dragonfly and a Redis Cluster node, and its data commands are only `GET`, `SET`, `MGET` and `PING` (no scripts, no client tracking), which ElastiCache and MemoryDB also serve. An ACL user also needs the connection commands the client sends when it dials: `HELLO`, `CLIENT`, `SELECT`, and `CLUSTER` in cluster mode; without them it is refused (`NOPERM`) and the cache stays bypassed. Whoever can write to the server can replace cached query results, which are served after the access policy has already been applied, so treat the server as part of WaveHouse's trust boundary (see [Deployment](/deployment#multiple-instances-and-the-shared-cache)). [Deployment](/deployment#multiple-instances-and-the-shared-cache) covers sizing, `maxmemory-policy` and what a reader on another instance can see. | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | -| `cache.redis.addrs` | `WH_CACHE_REDIS_ADDRS` | — | **Required** with `backend: redis.` `host:port` of the server; several are a cluster's seed nodes or the sentinels. Comma-separated in the env var. | +| `cache.redis.addrs` | `WH_CACHE_REDIS_ADDRS` | — | **Required** with `backend: redis`. `host:port` of the server; several are a cluster's seed nodes or the sentinels. Comma-separated in the env var. | | `cache.redis.mode` | `WH_CACHE_REDIS_MODE` | `standalone` | `standalone`, `cluster` or `sentinel`. | | `cache.redis.sentinel_master` | `WH_CACHE_REDIS_SENTINEL_MASTER` | — | The master set name. Required with `mode: sentinel`. | | `cache.redis.username` | `WH_CACHE_REDIS_USERNAME` | — | ACL user. Empty uses the server's `default` user. | @@ -138,7 +138,7 @@ The `redis` backend's settings, read only when `cache.backend` is `redis`. It ru | `cache.redis.tls.key_file` | `WH_CACHE_REDIS_TLS_KEY_FILE` | — | The client certificate's private key (PEM). | | `cache.redis.tls.server_name` | `WH_CACHE_REDIS_TLS_SERVER_NAME` | — | Name to verify the server's certificate against, when it differs from the address. | | `cache.redis.tls.insecure_skip_verify` | `WH_CACHE_REDIS_TLS_INSECURE_SKIP_VERIFY` | `false` | Accept any server certificate. Logged at `WARN` at boot: whoever can intercept the connection can read and replace cached results. | -| `cache.redis.key_prefix` | `WH_CACHE_REDIS_KEY_PREFIX` | `wh` | Leads every key, so several deployments can share one server. No `{` or `}`. | +| `cache.redis.key_prefix` | `WH_CACHE_REDIS_KEY_PREFIX` | `wh` | Leads every key, so several deployments can share one server, provided you trust each as much as the others: any of them can overwrite what the rest serve. No `{` or `}`. | | `cache.redis.timeout` | `WH_CACHE_REDIS_TIMEOUT` | `100ms` | Per operation. A lookup or fill that takes longer is a miss or a skipped fill, never a failed query. | | `cache.redis.dial_timeout` | `WH_CACHE_REDIS_DIAL_TIMEOUT` | `1s` | Per connection attempt. | | `cache.redis.max_value_bytes` | `WH_CACHE_REDIS_MAX_VALUE_BYTES` | `1048576` | Largest result stored, after compression (1 MiB). A larger one is returned to the caller but not cached. | diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 63f1b1a9..b94beb18 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -151,6 +151,12 @@ WH_AUTH_JWT_SECRET= # as an admin secret — inject from your secret store, serve only over TLS. WH_AUTH_OPERATOR_KEY= +# Optional shared query cache for several instances (see Multiple instances +# and the shared cache below); the password is a secret like the ones above. +# WH_CACHE_BACKEND=redis +# WH_CACHE_REDIS_ADDRS=redis:6379 +# WH_CACHE_REDIS_PASSWORD= + # Settings directory (required): roles.json, policies.json, pipes.json, # config.json — the hot-reloadable configuration: the access-control policy # and its roles, the named pipes, and the tunables including the ClickHouse @@ -407,9 +413,13 @@ The query-result cache is the layer that can be shared today. With the default ` - **The server is unreachable from the inserting instance.** The invalidation is kept and retried until it lands (`wavehouse_cache_invalidations_pending` counts what is owed). Meanwhile other instances that can still reach the server keep serving the older results, for as long as the outage lasts and at most until each entry's TTL. An instance that stops while invalidations are still owed loses them, with the same bound. The same thing happens today when a process stops between an insert and its invalidation. - **A failover to a replica that had not yet received the latest token writes** can bring back entries filed under the older tokens, bounded by the replication lag at the moment of failover and those entries' TTL. WaveHouse never reads from replicas. +- **The server is full and `maxmemory-policy` is `noeviction`.** It refuses the token writes, so invalidations are kept and retried, and until one lands every instance serves the results from before the insert, up to their TTL. +- **A pipe that writes** (an `INSERT` in `pipes.json`) has its result cached like a read, so a repeated identical call is answered from the cache and the write does not run again ([#386](https://github.com/Wave-RF/WaveHouse/issues/386)). With a shared cache that holds on every instance, until the entry's TTL. - **Admin writes through `POST /v1/ops/query`** do not invalidate the cache ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). With a shared cache, the stale results they leave are served by every instance, not only one. -**Sizing the server.** Every key WaveHouse writes has a TTL, and a version token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses, never bring back an entry it had invalidated. So set `maxmemory` and let the server evict: `maxmemory-policy allkeys-lru` (or `allkeys-lfu`, `volatile-lru`, `volatile-lfu`). Under `noeviction`, a full server refuses the writes. Fills then fail (counted by `wavehouse_cache_set_failures_total{reason="oom"}`), lookups keep working, and invalidations are kept and retried. A stored result is capped at `cache.redis.max_value_bytes` (1 MiB compressed). A tenant's version tokens share one hash tag, so each lookup reads them in one `MGET` in cluster mode as well. The results themselves carry no hash tag and spread across shards. Persistence is not needed: an empty server after a restart is a cold cache, not a wrong one. +**Sizing the server.** Every key WaveHouse writes has a TTL, and a version token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses, never bring back an entry it had invalidated. So set `maxmemory` and let the server evict: `maxmemory-policy allkeys-lru` (or `allkeys-lfu`, `volatile-lru`, `volatile-lfu`). Under `noeviction`, a full server refuses the writes. Fills then fail (counted by `wavehouse_cache_set_failures_total{reason="oom"}`), only lookups whose version tokens already exist keep working, and invalidations are kept and retried, so the pre-insert results above stay served: avoid `noeviction`. A stored result is capped at `cache.redis.max_value_bytes` (1 MiB compressed). A tenant's version tokens share one hash tag, so each lookup reads them in one `MGET` in cluster mode as well. The results themselves carry no hash tag and spread across shards. Persistence is not needed: an empty server after a restart is a cold cache, not a wrong one. + +**The server is inside the trust boundary.** A cached result is served after the access policy has filtered it, so whoever can write to the server can change what any caller reads. Keep it on a private network, require a password or ACL user (`WH_CACHE_REDIS_PASSWORD`), use TLS across links you do not trust, and share it only with deployments you trust as much as this one. **Coalescing stays per instance.** `singleflight` collapses identical concurrent queries within each instance, so a cold hot query costs at most one ClickHouse query per instance, not one per request. diff --git a/internal/config/cache_redis.go b/internal/config/cache_redis.go index f0af6713..3ae3114e 100644 --- a/internal/config/cache_redis.go +++ b/internal/config/cache_redis.go @@ -55,8 +55,7 @@ type CacheRedisTLS struct { InsecureSkipVerify bool `yaml:"insecure_skip_verify" env:"WH_CACHE_REDIS_TLS_INSECURE_SKIP_VERIFY"` } -// hasAddrs reports whether any address is set. An env file's blank -// `WH_CACHE_REDIS_ADDRS=` loads as one empty address, which is none. +// hasAddrs reports whether any address is set; a YAML `addrs: [""]` is none. func (r CacheRedisConfig) hasAddrs() bool { return len(r.Addrs) > 1 || len(r.Addrs) == 1 && r.Addrs[0] != "" } @@ -66,6 +65,9 @@ func (r CacheRedisConfig) validate() error { return errors.New("cache.backend=redis needs cache.redis.addrs (WH_CACHE_REDIS_ADDRS): the server's host:port, or a cluster's seeds, or the sentinels") } for _, a := range r.Addrs { + if strings.TrimSpace(a) != a { + return fmt.Errorf("cache.redis.addrs (WH_CACHE_REDIS_ADDRS) %q: no spaces around an address", a) + } if _, _, err := net.SplitHostPort(a); err != nil { return fmt.Errorf("cache.redis.addrs (WH_CACHE_REDIS_ADDRS) %q: want host:port: %w", a, err) } diff --git a/internal/config/cache_redis_test.go b/internal/config/cache_redis_test.go index 6d989c73..e1b5783a 100644 --- a/internal/config/cache_redis_test.go +++ b/internal/config/cache_redis_test.go @@ -189,6 +189,7 @@ func TestValidate_CacheRedis(t *testing.T) { }{ {"defaults", func(*CacheRedisConfig) {}, ""}, {"no addrs", func(r *CacheRedisConfig) { r.Addrs = nil }, "cache.backend=redis needs cache.redis.addrs (WH_CACHE_REDIS_ADDRS)"}, + {"addr with space", func(r *CacheRedisConfig) { r.Addrs = []string{"a:6379", " b:6379"} }, "no spaces around an address"}, {"addr without port", func(r *CacheRedisConfig) { r.Addrs = []string{"redis"} }, `cache.redis.addrs (WH_CACHE_REDIS_ADDRS) "redis": want host:port`}, {"mode", func(r *CacheRedisConfig) { r.Mode = "replica" }, `cache.redis.mode (WH_CACHE_REDIS_MODE) "replica": valid: standalone, cluster, sentinel`}, {"sentinel without master", func(r *CacheRedisConfig) { r.Mode = RedisSentinel }, "needs cache.redis.sentinel_master"}, From ee51e8401cac9e4e8669a6924dcd6b1931e5445a Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 04:50:24 -0400 Subject: [PATCH 08/27] test(app): assert the wired cache is a pruner before swapping it MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit TestReload_PrunesCacheIndexToServedTenants replaced a.cache with a recorder by hand, so the run-time a.cache.(pruner) check against the cache wireCache actually builds was never exercised: a change that stopped wiring a pruning cache left the test green while pruning silently went dead. Assert the wired cache satisfies pruner before the swap. Also reword a stale assertion message ("a rejection releases; it orphans nothing yet") that no longer describes real wiring now that a rejection drops the tenant's cache index via Prune, and correct the CHANGELOG's claim that the Redis-compatible backend (#613) already bounds its versions with a TTL — that backend does not exist yet. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- CHANGELOG.md | 2 +- internal/app/app_test.go | 5 ++++- 2 files changed, 5 insertions(+), 2 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 2515caa5..674c2df0 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -79,7 +79,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed - **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/app/wire.go`, `docs/src/content/docs/{architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` and pipe execution rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL. The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. -- **The cache's version index no longer grows with every bump, and forgets a tenant no longer served** (`internal/cache/{local,version_manager}.go` (+ tests), `internal/app/wire.go` (+ tests), `docs/src/content/docs/architecture.md`, `AGENTS.md`): fixes [#262](https://github.com/Wave-RF/WaveHouse/issues/262) for the in-process cache, part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The index nested each table under its tenant's version and each scope under its table's, and never pruned, so every tenant invalidation left the tenant's whole index behind, and it grew with every tenant ever served. It now holds one version per tenant, per (tenant, table) and per (tenant, table, scope), bumped in place. A tenant invalidation drops the tenant's index and hands its next key a generation unique within the process, so nothing cached before it can match again, and a table bump drops the table's scope versions. After each settings reload the index of every tenant no longer served, removed or rejected, is dropped the same way; its cached results are orphaned with it, as they already were when such a tenant came back on a pool. No change to what is cached or served. The Redis-compatible backend (#613) bounds its versions with a TTL instead. +- **The cache's version index no longer grows with every bump, and forgets a tenant no longer served** (`internal/cache/{local,version_manager}.go` (+ tests), `internal/app/wire.go` (+ tests), `docs/src/content/docs/architecture.md`, `AGENTS.md`): fixes [#262](https://github.com/Wave-RF/WaveHouse/issues/262) for the in-process cache, part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The index nested each table under its tenant's version and each scope under its table's, and never pruned, so every tenant invalidation left the tenant's whole index behind, and it grew with every tenant ever served. It now holds one version per tenant, per (tenant, table) and per (tenant, table, scope), bumped in place. A tenant invalidation drops the tenant's index and hands its next key a generation unique within the process, so nothing cached before it can match again, and a table bump drops the table's scope versions. After each settings reload the index of every tenant no longer served, removed or rejected, is dropped the same way; its cached results are orphaned with it, as they already were when such a tenant came back on a pool. No change to what is cached or served. The Redis-compatible backend (#613) will bound its versions with a TTL instead. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/internal/app/app_test.go b/internal/app/app_test.go index 84c67070..d21f376a 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -671,7 +671,7 @@ func TestReload_ReadmittedTenantCacheIsOrphaned(t *testing.T) { rewriteSettings(t, filepath.Join(root, "globex"), invalidQuery) a.tenants.Reload("test") - assert.Empty(t, mock.GetTenants(), "a rejection releases; it orphans nothing yet") + assert.Empty(t, mock.GetTenants(), "a rejection calls no InvalidateTenant (Prune drops its index)") rewriteSettings(t, filepath.Join(root, "globex"), nil) _, adopted = a.tenants.Reload("test") require.True(t, adopted) @@ -712,6 +712,9 @@ func (p *pruneRecorder) last() map[tenant.ID]bool { func TestReload_PrunesCacheIndexToServedTenants(t *testing.T) { root := writeNestedSettings(t, map[string]map[string]any{"acme": nil, "globex": nil}) a := newApp(t, testConfig(t, root), Options{}) + _, ok := a.cache.(pruner) + require.True(t, ok, "the wired cache prunes") + rec := &pruneRecorder{} a.cache = rec From 9a4888a4959b8523f8c7907947928a2384068b66 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 04:58:48 -0400 Subject: [PATCH 09/27] fix(config): cache.redis defaults in defaults(); 0 never compresses Adopt #632's convention for the cache.redis block: its seven non-zero defaults move from env-default tags into defaults(), each gets a zero case, and the configuration.mdx rows use *(empty)* for an empty default. The doc-defaults test learns to read a duration and an empty list. With the file's zeros now kept, compress_min_bytes needs no -1: 0 means never compress, the backend's own meaning, from the file and the environment alike, and wire.go passes it through unchanged. A negative value refuses boot. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/configuration.mdx | 20 ++++++++++---------- internal/app/app_test.go | 15 +++++++++------ internal/app/wire.go | 6 +----- internal/config/cache_redis.go | 23 +++++++++++------------ internal/config/cache_redis_test.go | 19 ++++++++----------- internal/config/config.go | 14 +++++++++++--- internal/config/defaults_test.go | 14 ++++++++++++++ 9 files changed, 66 insertions(+), 49 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 6f742cb0..7d0e6ea2 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone`, `cluster`, `sentinel`), `sentinel_master`, `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `-1` never compresses, since the loader reads a `0` in the file as unset) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. A malformed block — no address, an address without a port, an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. +- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone`, `cluster`, `sentinel`), `sentinel_master`, `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. A malformed block — no address, an address without a port, an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. - **A Redis-compatible shared cache backend** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row open a circuit breaker that skips the server until a probe succeeds, and deferred invalidations are retried until they land, collapsing to one tenant-wide bump per tenant past 100,000 keys. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. `cache.backend: redis` selects it (the entry above). Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression and server-stops-answering cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index fbe6324f..2aba0526 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -126,7 +126,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — the implementation each layer runs on, the process's `roles`, resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `cache.redis.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). - **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load` when `Config.NeedsDataDir` says a selected backend keeps state there, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. - **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` in `Validate`, between `validateRoles` and `validateTopology`), which refuses a value not on the list and names the ones that are. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns what a valid configuration is still likely to get wrong — the combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), a `cache.redis` block that is not read, certificate verification turned off — which `app.New` logs at `WARN`. -- **cache_redis.go** — `CacheRedisConfig`, the `cache.redis` sub-block, and its checks: an address, each `host:port`, a known mode, a sentinel's master name, `db` 0 in cluster mode, positive timeouts and sizes, a `version_ttl` of at least 2 s, and `compress_min_bytes` positive or `-1` (never: cleanenv reads a `0` in the file as unset and applies the default, so `0` cannot mean off). `CacheRedisTLS.Config` builds the `tls.Config`, reading the files; `Validate` calls it so an unreadable file refuses boot, and `internal/app` calls it again to build the connection. A TLS key set while `tls.enabled` is off is an error rather than a plaintext connection. +- **cache_redis.go** — `CacheRedisConfig`, the `cache.redis` sub-block, and its checks: an address, each `host:port`, a known mode, a sentinel's master name, `db` 0 in cluster mode, positive timeouts and sizes, a `version_ttl` of at least 2 s, and a `compress_min_bytes` that is not negative (`0` never compresses); its defaults are in `defaults()` with the rest. `CacheRedisTLS.Config` builds the `tls.Config`, reading the files; `Validate` calls it so an unreadable file refuses boot, and `internal/app` calls it again to build the connection. A TLS key set while `tls.enabled` is off is an error rather than a plaintext connection. - **config.go**, roles — `roles` (`[]Role`: `api`, `ingest`, `sweeper`; `AllRoles` by default; `Has(Role)`) picks which components `internal/app` wires, and `instance_id` names the process (`-<8 hex>` when empty, resolved in `Load`; today only logged at boot, and a distributed coordinator will record it as a lease's holder). `validateRoles` refuses an empty list, an empty entry, an unknown or a repeated role; `validateTopology` refuses a role set the backends cannot serve: any split over the embedded MQ, and a process with exactly one of `api` and `ingest` over a local cache. `NeedsDataDir` counts Pebble only for a process running `api`, and `Warnings` is empty without `api`, since only that role opens a cache it reads or a dedupe store. - **strict.go** — `rejectUnknownKeys`, the YAML half: re-reads the file as a generic tree and walks it against the struct's `yaml` tags, listing every key the struct doesn't declare. cleanenv itself is lenient by design, which is exactly wrong for boot config once keys have moved to the settings directory. - **persistence.go** — `WarnIfFreshDataDir` logs the startup `WARN` when `data_dir` is missing or empty (on a redeploy, the sign that the volume didn't persist); `LogStorageInitError` attaches the UID-65532 `permissionHint` to a NATS or Pebble open failure that looks like a permission denial — the same hint string `CheckDataDir` uses. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 91aa72d6..0b534854 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -150,23 +150,23 @@ The `redis` backend's settings, read only when `cache.backend` is `redis`. It is | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | -| `cache.redis.addrs` | `WH_CACHE_REDIS_ADDRS` | — | **Required** with `backend: redis`. `host:port` of the server; several are a cluster's seed nodes or the sentinels. Comma-separated in the env var. | +| `cache.redis.addrs` | `WH_CACHE_REDIS_ADDRS` | *(empty)* | **Required** with `backend: redis`. `host:port` of the server; several are a cluster's seed nodes or the sentinels. Comma-separated in the env var. | | `cache.redis.mode` | `WH_CACHE_REDIS_MODE` | `standalone` | `standalone`, `cluster` or `sentinel`. | -| `cache.redis.sentinel_master` | `WH_CACHE_REDIS_SENTINEL_MASTER` | — | The master set name. Required with `mode: sentinel`. | -| `cache.redis.username` | `WH_CACHE_REDIS_USERNAME` | — | ACL user. Empty uses the server's `default` user. | -| `cache.redis.password` | `WH_CACHE_REDIS_PASSWORD` | — | A secret: set it through the environment (or your secret store's env injection), not in a tracked `config.yaml`. | +| `cache.redis.sentinel_master` | `WH_CACHE_REDIS_SENTINEL_MASTER` | *(empty)* | The master set name. Required with `mode: sentinel`. | +| `cache.redis.username` | `WH_CACHE_REDIS_USERNAME` | *(empty)* | ACL user. Empty uses the server's `default` user. | +| `cache.redis.password` | `WH_CACHE_REDIS_PASSWORD` | *(empty)* | A secret: set it through the environment (or your secret store's env injection), not in a tracked `config.yaml`. | | `cache.redis.db` | `WH_CACHE_REDIS_DB` | `0` | Database number (`SELECT`). Standalone and sentinel only: a cluster has only database `0`, and any other value refuses boot. | | `cache.redis.tls.enabled` | `WH_CACHE_REDIS_TLS_ENABLED` | `false` | Connect over TLS, verifying the server against the system roots or `ca_file`. Any other `tls` key set while this is off refuses boot, rather than connecting in plaintext. | -| `cache.redis.tls.ca_file` | `WH_CACHE_REDIS_TLS_CA_FILE` | — | PEM file of the authorities to trust instead of the system roots. | -| `cache.redis.tls.cert_file` | `WH_CACHE_REDIS_TLS_CERT_FILE` | — | Client certificate (PEM) for mutual TLS. Set together with `key_file`. | -| `cache.redis.tls.key_file` | `WH_CACHE_REDIS_TLS_KEY_FILE` | — | The client certificate's private key (PEM). | -| `cache.redis.tls.server_name` | `WH_CACHE_REDIS_TLS_SERVER_NAME` | — | Name to verify the server's certificate against, when it differs from the address. | +| `cache.redis.tls.ca_file` | `WH_CACHE_REDIS_TLS_CA_FILE` | *(empty)* | PEM file of the authorities to trust instead of the system roots. | +| `cache.redis.tls.cert_file` | `WH_CACHE_REDIS_TLS_CERT_FILE` | *(empty)* | Client certificate (PEM) for mutual TLS. Set together with `key_file`. | +| `cache.redis.tls.key_file` | `WH_CACHE_REDIS_TLS_KEY_FILE` | *(empty)* | The client certificate's private key (PEM). | +| `cache.redis.tls.server_name` | `WH_CACHE_REDIS_TLS_SERVER_NAME` | *(empty)* | Name to verify the server's certificate against, when it differs from the address. | | `cache.redis.tls.insecure_skip_verify` | `WH_CACHE_REDIS_TLS_INSECURE_SKIP_VERIFY` | `false` | Accept any server certificate. Logged at `WARN` at boot: whoever can intercept the connection can read and replace cached results. | | `cache.redis.key_prefix` | `WH_CACHE_REDIS_KEY_PREFIX` | `wh` | Leads every key, so several deployments can share one server, provided you trust each as much as the others: any of them can overwrite what the rest serve. No `{` or `}`. | | `cache.redis.timeout` | `WH_CACHE_REDIS_TIMEOUT` | `100ms` | Per operation. A lookup or fill that takes longer is a miss or a skipped fill, never a failed query. | | `cache.redis.dial_timeout` | `WH_CACHE_REDIS_DIAL_TIMEOUT` | `1s` | Per connection attempt. | | `cache.redis.max_value_bytes` | `WH_CACHE_REDIS_MAX_VALUE_BYTES` | `1048576` | Largest result stored, after compression (1 MiB). A larger one is returned to the caller but not cached. | -| `cache.redis.compress_min_bytes` | `WH_CACHE_REDIS_COMPRESS_MIN_BYTES` | `1024` | Results at least this large are zstd-compressed when that makes them smaller. `-1` never compresses. `0` is not "off": in the YAML file it reads as unset and takes the default, and in the env var it refuses boot. | +| `cache.redis.compress_min_bytes` | `WH_CACHE_REDIS_COMPRESS_MIN_BYTES` | `1024` | Results at least this large are zstd-compressed when that makes them smaller. `0` never compresses. | | `cache.redis.version_ttl` | `WH_CACHE_REDIS_VERSION_TTL` | `168h` | How long a table's or tenant's version token outlives its last write, so the tokens of dropped tables and removed tenants eventually expire. At least `2s`. An expired token only causes misses. | **When the server is unreachable or misbehaves, the cache is bypassed; queries are not.** A failure or a timeout makes the lookup a miss and the fill a no-op, and five in a row open a circuit breaker that skips the server entirely until a probe, every 5 s, gets an answer. Queries then go straight to ClickHouse, still coalesced per instance by `singleflight`. An invalidation the server did not take is kept and retried until it lands. `/readyz` does not depend on the cache. @@ -284,7 +284,7 @@ cache: timeout: 100ms dial_timeout: 1s max_value_bytes: 1048576 - compress_min_bytes: 1024 # -1 = never + compress_min_bytes: 1024 # 0 = never version_ttl: 168h dedupe: diff --git a/internal/app/app_test.go b/internal/app/app_test.go index 2ca4a205..204c1371 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -796,9 +796,10 @@ func TestNew_RedisCacheRefusesAnUnreadableTLSFile(t *testing.T) { require.ErrorContains(t, err, "cache init: cache.redis.tls.ca_file") } -// The boot config's defaults are the backend's, and -1 is the backend's -// "never compress". Driven from Load, not a literal, so a default changed on -// one side only fails here. +// The boot config's defaults are the backend's, and a compress_min_bytes of +// 0 reaches the backend as its "never compress" rather than its default. +// Driven from Load, not a literal, so a default changed on one side only +// fails here. func TestRedisConfig_FromLoadedDefaults(t *testing.T) { t.Setenv("WH_SETTINGS_DIR", t.TempDir()) t.Setenv("WH_CACHE_BACKEND", "redis") @@ -815,11 +816,13 @@ func TestRedisConfig_FromLoadedDefaults(t *testing.T) { CompressMinBytes: cache.DefaultRedisCompressMinBytes, VersionTTL: cache.DefaultRedisVersionTTL, }, got) - loaded.Cache.Redis.CompressMinBytes = -1 - loaded.Cache.Redis.Mode = config.RedisCluster + t.Setenv("WH_CACHE_REDIS_COMPRESS_MIN_BYTES", "0") + t.Setenv("WH_CACHE_REDIS_MODE", "cluster") + loaded, err = config.Load(filepath.Join(t.TempDir(), "none.yaml")) + require.NoError(t, err) got, err = redisConfig(loaded.Cache.Redis) require.NoError(t, err) - assert.Zero(t, got.CompressMinBytes) + assert.Zero(t, got.CompressMinBytes, "the backend's never, not its default") assert.Equal(t, cache.RedisCluster, got.Mode) assert.Equal(t, cache.RedisSentinel, config.RedisSentinel) } diff --git a/internal/app/wire.go b/internal/app/wire.go index 60c9e4a9..ea23606d 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -664,10 +664,6 @@ func redisConfig(r config.CacheRedisConfig) (cache.RedisConfig, error) { if err != nil { return cache.RedisConfig{}, err } - compressMin := r.CompressMinBytes - if compressMin < 0 { - compressMin = 0 // the backend's "never" - } return cache.RedisConfig{ Addrs: r.Addrs, Mode: r.Mode, @@ -680,7 +676,7 @@ func redisConfig(r config.CacheRedisConfig) (cache.RedisConfig, error) { Timeout: r.Timeout, DialTimeout: r.DialTimeout, MaxValueBytes: r.MaxValueBytes, - CompressMinBytes: compressMin, + CompressMinBytes: r.CompressMinBytes, VersionTTL: r.VersionTTL, }, nil } diff --git a/internal/config/cache_redis.go b/internal/config/cache_redis.go index 3ae3114e..997d19a3 100644 --- a/internal/config/cache_redis.go +++ b/internal/config/cache_redis.go @@ -25,24 +25,23 @@ type CacheRedisConfig struct { // Addrs are host:port pairs: the server, or seeds for a cluster, or the // sentinels. Addrs []string `yaml:"addrs" env:"WH_CACHE_REDIS_ADDRS"` - Mode string `yaml:"mode" env:"WH_CACHE_REDIS_MODE" env-default:"standalone"` + Mode string `yaml:"mode" env:"WH_CACHE_REDIS_MODE"` SentinelMaster string `yaml:"sentinel_master" env:"WH_CACHE_REDIS_SENTINEL_MASTER"` Username string `yaml:"username" env:"WH_CACHE_REDIS_USERNAME"` Password string `yaml:"password" env:"WH_CACHE_REDIS_PASSWORD"` DB int `yaml:"db" env:"WH_CACHE_REDIS_DB"` TLS CacheRedisTLS `yaml:"tls"` // KeyPrefix leads every key, so deployments can share one server. - KeyPrefix string `yaml:"key_prefix" env:"WH_CACHE_REDIS_KEY_PREFIX" env-default:"wh"` - Timeout time.Duration `yaml:"timeout" env:"WH_CACHE_REDIS_TIMEOUT" env-default:"100ms"` - DialTimeout time.Duration `yaml:"dial_timeout" env:"WH_CACHE_REDIS_DIAL_TIMEOUT" env-default:"1s"` + KeyPrefix string `yaml:"key_prefix" env:"WH_CACHE_REDIS_KEY_PREFIX"` + Timeout time.Duration `yaml:"timeout" env:"WH_CACHE_REDIS_TIMEOUT"` + DialTimeout time.Duration `yaml:"dial_timeout" env:"WH_CACHE_REDIS_DIAL_TIMEOUT"` // MaxValueBytes is the largest value stored, after compression. - MaxValueBytes int `yaml:"max_value_bytes" env:"WH_CACHE_REDIS_MAX_VALUE_BYTES" env-default:"1048576"` - // CompressMinBytes is the smallest value zstd-compressed; -1 never - // compresses. Not 0: the loader reads a 0 in the file as unset and - // applies the default, so 0 cannot mean off. - CompressMinBytes int `yaml:"compress_min_bytes" env:"WH_CACHE_REDIS_COMPRESS_MIN_BYTES" env-default:"1024"` + MaxValueBytes int `yaml:"max_value_bytes" env:"WH_CACHE_REDIS_MAX_VALUE_BYTES"` + // CompressMinBytes is the smallest value zstd-compressed; 0 never + // compresses. + CompressMinBytes int `yaml:"compress_min_bytes" env:"WH_CACHE_REDIS_COMPRESS_MIN_BYTES"` // VersionTTL is how long a version token outlives its last bump. - VersionTTL time.Duration `yaml:"version_ttl" env:"WH_CACHE_REDIS_VERSION_TTL" env-default:"168h"` + VersionTTL time.Duration `yaml:"version_ttl" env:"WH_CACHE_REDIS_VERSION_TTL"` } // CacheRedisTLS is cache.redis.tls. The files are paths, read at boot. @@ -106,8 +105,8 @@ func (r CacheRedisConfig) validate() error { if r.MaxValueBytes <= 0 { return fmt.Errorf("cache.redis.max_value_bytes (WH_CACHE_REDIS_MAX_VALUE_BYTES) %d must be positive", r.MaxValueBytes) } - if r.CompressMinBytes == 0 || r.CompressMinBytes < -1 { - return fmt.Errorf("cache.redis.compress_min_bytes (WH_CACHE_REDIS_COMPRESS_MIN_BYTES) %d: want a positive size, or -1 to never compress", r.CompressMinBytes) + if r.CompressMinBytes < 0 { + return fmt.Errorf("cache.redis.compress_min_bytes (WH_CACHE_REDIS_COMPRESS_MIN_BYTES) %d is negative: want a size, or 0 to never compress", r.CompressMinBytes) } if _, err := r.TLS.Config(); err != nil { return err diff --git a/internal/config/cache_redis_test.go b/internal/config/cache_redis_test.go index e1b5783a..5389715a 100644 --- a/internal/config/cache_redis_test.go +++ b/internal/config/cache_redis_test.go @@ -60,7 +60,7 @@ func TestLoad_CacheRedisFromEnv(t *testing.T) { "WH_CACHE_REDIS_TIMEOUT": "250ms", "WH_CACHE_REDIS_DIAL_TIMEOUT": "3s", "WH_CACHE_REDIS_MAX_VALUE_BYTES": "2048", - "WH_CACHE_REDIS_COMPRESS_MIN_BYTES": "-1", + "WH_CACHE_REDIS_COMPRESS_MIN_BYTES": "0", "WH_CACHE_REDIS_VERSION_TTL": "24h", "WH_CACHE_REDIS_TLS_INSECURE_SKIP_VERIFY": "false", } { @@ -75,7 +75,7 @@ func TestLoad_CacheRedisFromEnv(t *testing.T) { Enabled: true, CAFile: caFile, CertFile: certFile, KeyFile: keyFile, ServerName: "redis.internal", }, KeyPrefix: "staging", Timeout: 250 * time.Millisecond, DialTimeout: 3 * time.Second, - MaxValueBytes: 2048, CompressMinBytes: -1, VersionTTL: 24 * time.Hour, + MaxValueBytes: 2048, CompressMinBytes: 0, VersionTTL: 24 * time.Hour, }, cfg.Cache.Redis) tc, err := cfg.Cache.Redis.TLS.Config() require.NoError(t, err) @@ -112,10 +112,8 @@ cache: assert.Equal(t, 1024, r.CompressMinBytes) } -// A 0 in the file is read as unset, so it takes the default instead of -// switching compression off: why "never" is -1. Pinned so a loader that -// starts honoring the 0 is noticed. -func TestLoad_CacheRedisCompressZeroInYAMLIsTheDefault(t *testing.T) { +// A 0 in the file is kept, and is the backend's "never compress". +func TestLoad_CacheRedisCompressZeroInYAMLIsNever(t *testing.T) { t.Parallel() path := filepath.Join(t.TempDir(), "config.yaml") require.NoError(t, os.WriteFile(path, []byte(` @@ -129,7 +127,7 @@ cache: `), 0o600)) cfg, err := Load(path) require.NoError(t, err) - assert.Equal(t, 1024, cfg.Cache.Redis.CompressMinBytes) + assert.Zero(t, cfg.Cache.Redis.CompressMinBytes) } func TestLoad_CacheRedisRefusesUnknownKeys(t *testing.T) { @@ -171,7 +169,7 @@ func TestUnboundEnv_KnowsTheCacheRedisVariables(t *testing.T) { t.Parallel() assert.Empty(t, unboundEnv([]string{ "WH_CACHE_REDIS_ADDRS=r:6379", "WH_CACHE_REDIS_PASSWORD=x", "WH_CACHE_REDIS_TLS_CA_FILE=/ca.pem", - "WH_CACHE_REDIS_VERSION_TTL=1h", "WH_CACHE_REDIS_COMPRESS_MIN_BYTES=-1", + "WH_CACHE_REDIS_VERSION_TTL=1h", "WH_CACHE_REDIS_COMPRESS_MIN_BYTES=0", })) assert.Equal(t, []string{"WH_CACHE_REDIS_ADDR"}, unboundEnv([]string{"WH_CACHE_REDIS_ADDR=r:6379"})) } @@ -203,9 +201,8 @@ func TestValidate_CacheRedis(t *testing.T) { {"negative dial timeout", func(r *CacheRedisConfig) { r.DialTimeout = -time.Second }, "cache.redis.dial_timeout"}, {"short version ttl", func(r *CacheRedisConfig) { r.VersionTTL = time.Second }, "cache.redis.version_ttl (WH_CACHE_REDIS_VERSION_TTL) 1s is under 2s"}, {"zero max value", func(r *CacheRedisConfig) { r.MaxValueBytes = 0 }, "cache.redis.max_value_bytes"}, - {"compress 0", func(r *CacheRedisConfig) { r.CompressMinBytes = 0 }, "or -1 to never compress"}, - {"compress -2", func(r *CacheRedisConfig) { r.CompressMinBytes = -2 }, "or -1 to never compress"}, - {"compress never", func(r *CacheRedisConfig) { r.CompressMinBytes = -1 }, ""}, + {"compress never", func(r *CacheRedisConfig) { r.CompressMinBytes = 0 }, ""}, + {"compress negative", func(r *CacheRedisConfig) { r.CompressMinBytes = -1 }, "-1 is negative: want a size, or 0 to never compress"}, {"tls files while off", func(r *CacheRedisConfig) { r.TLS.CAFile = caFile }, "cache.redis.tls.enabled (WH_CACHE_REDIS_TLS_ENABLED) is off"}, {"tls system roots", func(r *CacheRedisConfig) { r.TLS.Enabled = true }, ""}, {"tls full", func(r *CacheRedisConfig) { diff --git a/internal/config/config.go b/internal/config/config.go index 95bfacb1..27f4188c 100644 --- a/internal/config/config.go +++ b/internal/config/config.go @@ -7,6 +7,7 @@ import ( "os" "slices" "strings" + "time" "github.com/ilyakaznacheev/cleanenv" ) @@ -255,9 +256,16 @@ func defaults() Config { Roles: AllRoles(), Server: Server{Port: 8080, ShutdownTimeout: 10}, MQ: MQ{Backend: MQEmbedded}, - Cache: Cache{Backend: CacheLocal, L1MaxCost: 64 << 20}, - Dedupe: Dedupe{Backend: DedupePebble}, - Coord: Coord{Backend: CoordLocal}, + Cache: Cache{ + Backend: CacheLocal, L1MaxCost: 64 << 20, + Redis: CacheRedisConfig{ + Mode: RedisStandalone, KeyPrefix: "wh", + Timeout: 100 * time.Millisecond, DialTimeout: time.Second, + MaxValueBytes: 1 << 20, CompressMinBytes: 1 << 10, VersionTTL: 168 * time.Hour, + }, + }, + Dedupe: Dedupe{Backend: DedupePebble}, + Coord: Coord{Backend: CoordLocal}, OTel: OTel{ Traces: OTelTraces{Enabled: true, SampleRate: 1.0}, Metrics: OTelMetrics{Enabled: true}, diff --git a/internal/config/defaults_test.go b/internal/config/defaults_test.go index 6877be4b..1a8fba33 100644 --- a/internal/config/defaults_test.go +++ b/internal/config/defaults_test.go @@ -9,6 +9,7 @@ import ( "strconv" "strings" "testing" + "time" "github.com/stretchr/testify/assert" "github.com/stretchr/testify/require" @@ -37,6 +38,15 @@ var zeroCases = []zeroCase{ {"cache.l1_max_cost", "WH_CACHE_L1_MAX_COST", int64(0), int64(64 << 20), "1024", int64(1024), func(c *Config) any { return c.Cache.L1MaxCost }}, {"prometheus.path", "WH_PROMETHEUS_PATH", "", "/metrics", "/prom", "/prom", func(c *Config) any { return c.Prometheus.Path }}, {"data_dir", "WH_DATA_DIR", "", "./data", "/var/lib/wh", "/var/lib/wh", func(c *Config) any { return c.DataDir }}, + // The cache.redis block is validated only under backend=redis, so under + // the default backend its zeros load as written. + {"cache.redis.mode", "WH_CACHE_REDIS_MODE", "", RedisStandalone, RedisCluster, RedisCluster, func(c *Config) any { return c.Cache.Redis.Mode }}, + {"cache.redis.key_prefix", "WH_CACHE_REDIS_KEY_PREFIX", "", "wh", "staging", "staging", func(c *Config) any { return c.Cache.Redis.KeyPrefix }}, + {"cache.redis.timeout", "WH_CACHE_REDIS_TIMEOUT", time.Duration(0), 100 * time.Millisecond, "250ms", 250 * time.Millisecond, func(c *Config) any { return c.Cache.Redis.Timeout }}, + {"cache.redis.dial_timeout", "WH_CACHE_REDIS_DIAL_TIMEOUT", time.Duration(0), time.Second, "500ms", 500 * time.Millisecond, func(c *Config) any { return c.Cache.Redis.DialTimeout }}, + {"cache.redis.max_value_bytes", "WH_CACHE_REDIS_MAX_VALUE_BYTES", 0, 1 << 20, "2048", 2048, func(c *Config) any { return c.Cache.Redis.MaxValueBytes }}, + {"cache.redis.compress_min_bytes", "WH_CACHE_REDIS_COMPRESS_MIN_BYTES", 0, 1 << 10, "2048", 2048, func(c *Config) any { return c.Cache.Redis.CompressMinBytes }}, + {"cache.redis.version_ttl", "WH_CACHE_REDIS_VERSION_TTL", time.Duration(0), 168 * time.Hour, "1h", time.Hour, func(c *Config) any { return c.Cache.Redis.VersionTTL }}, } // refusedZeros are the non-zero defaults whose zero Validate refuses: written @@ -296,11 +306,15 @@ func parseDocDefault(t *testing.T, key, cell string, like any) any { v, err = strconv.ParseInt(cell, 10, 64) case float64: v, err = strconv.ParseFloat(cell, 64) + case time.Duration: + v, err = time.ParseDuration(cell) default: rt := reflect.TypeOf(like) switch { case rt.Kind() == reflect.String: v = reflect.ValueOf(cell).Convert(rt).Interface() + case rt.Kind() == reflect.Slice && rt.Elem().Kind() == reflect.String && cell == "": + v = reflect.Zero(rt).Interface() // cache.redis.addrs: nil case rt.Kind() == reflect.Slice && rt.Elem().Kind() == reflect.String: // roles: a comma-separated cell parts := strings.Split(cell, ",") sv := reflect.MakeSlice(rt, len(parts), len(parts)) From 6fa1ae6eb949998c548c34bdc6e1cc623402029c Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:00:52 -0400 Subject: [PATCH 10/27] fix(cache): give the two-instance test its roles; redis is the shared cache app.New now wires the API, ingest and cache only for the roles a config names (#622), and refuses a Config with none, so the shared-cache integration test's instances run every role, as the suite's own app does. The refusal of api without ingest over a local cache now names cache.backend=redis, the shared cache it asks for, and the configuration and deployment pages say this build has a shared cache but still refuses every split while the queue and leases are in-process. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- CHANGELOG.md | 2 +- docs/src/content/docs/configuration.mdx | 2 +- docs/src/content/docs/deployment.md | 2 +- internal/config/config.go | 2 +- internal/config/roles_test.go | 4 +++- tests/integration/shared_cache_test.go | 1 + 6 files changed, 8 insertions(+), 5 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 7d0e6ea2..f8f2efab 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone`, `cluster`, `sentinel`), `sentinel_master`, `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. A malformed block — no address, an address without a port, an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. +- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone`, `cluster`, `sentinel`), `sentinel_master`, `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. - **A Redis-compatible shared cache backend** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row open a circuit breaker that skips the server until a probe succeeds, and deferred invalidations are retried until they land, collapsing to one tenant-wide bump per tenant past 100,000 keys. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. `cache.backend: redis` selects it (the entry above). Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression and server-stops-answering cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 0b534854..6425d6f7 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -70,7 +70,7 @@ Every process, whatever its roles, reads the settings directory and reloads it ( Boot refuses a role set the selected backends cannot serve: - **Any split with `mq.backend=embedded`.** The embedded queue lives inside its process and listens on no port, so a process without every role could not reach it. Until a shared `mq.backend` exists, every process runs every role. -- **`api` without `ingest`, or `ingest` without `api`, with `cache.backend=local`.** The ingest worker invalidates the cache the API reads, and a local cache in another process never sees that invalidation. Run `api` and `ingest` together, or choose a shared `cache.backend`. A `sweeper`-only process holds no cache, so this rule does not apply to it. +- **`api` without `ingest`, or `ingest` without `api`, with `cache.backend=local`.** The ingest worker invalidates the cache the API reads, and a local cache in another process never sees that invalidation. Run `api` and `ingest` together, or set [`cache.backend: redis`](#backends), one cache every process shares. A `sweeper`-only process holds no cache, so this rule does not apply to it. ### Server diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 32e508a6..2acf348c 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -347,7 +347,7 @@ By default one process runs all of WaveHouse. [`roles`](/configuration#process-r - **Ingest.** Every ingest pod consumes the same shared durable consumer and competes for its messages, so throughput scales with the pod count. The rows of one table are then split across pods: each pod writes smaller batches, and rows written by different pods do not reach ClickHouse in publish order. - **Sweeper.** The sweeper runs under a lease held in the shared `coord.backend`, so only one pod sweeps at a time. A second replica waits and takes over when the first stops. -A split needs backends that every process can reach: a shared `mq.backend`, so that every process reaches the same queue; a shared `cache.backend`, so that the ingest pods' invalidations reach the API pods' cache; and a shared `coord.backend`, so that the sweeper lease spans pods. **This build has only the in-process backends, so boot refuses any split** and names the backend to change. Until shared backends ship, run every role in one process, the default. +A split needs backends that every process can reach: a shared `mq.backend`, so that every process reaches the same queue; a shared `cache.backend` ([`redis`](#multiple-instances-and-the-shared-cache)), so that the ingest pods' invalidations reach the API pods' cache; and a shared `coord.backend`, so that the sweeper lease spans pods. **This build has a shared cache but only the in-process queue and leases, so boot refuses any split** and names the backend to change. Until a shared `mq.backend` and `coord.backend` ship, run every role in one process, the default. A pod without the `api` role serves an ops listener on `:8080`: `/livez`, `/readyz` and their aliases, `/version`, the metrics path when `prometheus.port` is `0`, and `POST /v1/ops/settings/reload`. Every other route answers 404 (under `/v1/ops`, 403 without the operator key, and 401 for a bearer token). Point the same probes at it as at an API pod. `/livez` does not wait for schema discovery there, because only the API runs it. `/readyz` checks ClickHouse in an ingest pod, and is ready once a sweeper pod has booted. Every pod reads the settings directory, so mount it in every Deployment. The reload route on the ops listener accepts only the operator key, so whatever reloads your API pods over HTTP must send the operator key to the worker pods too, or rely on `SIGHUP` (or, over a flat directory, the directory watcher) instead. diff --git a/internal/config/config.go b/internal/config/config.go index 27f4188c..0bc70f2a 100644 --- a/internal/config/config.go +++ b/internal/config/config.go @@ -218,7 +218,7 @@ func (c *Config) validateTopology() error { return fmt.Errorf("roles %s with mq.backend=embedded: the embedded MQ lives inside this process, and a process without it cannot reach its queue — run every role (%s), or set a shared mq.backend", joinRoles(c.Roles), joinRoles(allRoles)) } if c.splitsCache() && c.Cache.Backend == CacheLocal { - return fmt.Errorf("roles %s with cache.backend=local: api and ingest run in different processes, and the ingest worker's cache invalidation would never reach the API's cache — run api and ingest together, or set a shared cache.backend", joinRoles(c.Roles)) + return fmt.Errorf("roles %s with cache.backend=local: api and ingest run in different processes, and the ingest worker's cache invalidation would never reach the API's cache — run api and ingest together, or set cache.backend=redis, one cache every process shares", joinRoles(c.Roles)) } return nil } diff --git a/internal/config/roles_test.go b/internal/config/roles_test.go index 18b28267..914815bf 100644 --- a/internal/config/roles_test.go +++ b/internal/config/roles_test.go @@ -114,12 +114,14 @@ func TestValidate_RoleSplits(t *testing.T) { {"every role, shared queue", all, "shared", CacheLocal, ""}, {"api+ingest, shared queue", []Role{RoleAPI, RoleIngest}, "shared", CacheLocal, ""}, {"sweeper, shared queue", []Role{RoleSweeper}, "shared", CacheLocal, ""}, - {"api, local cache", []Role{RoleAPI}, "shared", CacheLocal, "roles api with cache.backend=local: api and ingest run in different processes"}, + {"api, local cache", []Role{RoleAPI}, "shared", CacheLocal, "roles api with cache.backend=local: api and ingest run in different processes, and the ingest worker's cache invalidation would never reach the API's cache — run api and ingest together, or set cache.backend=redis"}, {"ingest, local cache", []Role{RoleIngest}, "shared", CacheLocal, "roles ingest with cache.backend=local"}, {"api+sweeper, local cache", []Role{RoleAPI, RoleSweeper}, "shared", CacheLocal, "roles api,sweeper with cache.backend=local"}, {"ingest+sweeper, local cache", []Role{RoleIngest, RoleSweeper}, "shared", CacheLocal, "roles ingest,sweeper with cache.backend=local"}, {"api, shared cache", []Role{RoleAPI}, "shared", "shared", ""}, {"ingest, shared cache", []Role{RoleIngest}, "shared", "shared", ""}, + {"api, redis cache", []Role{RoleAPI}, "shared", CacheRedis, ""}, + {"ingest, redis cache", []Role{RoleIngest}, "shared", CacheRedis, ""}, } { t.Run(tc.name, func(t *testing.T) { t.Parallel() diff --git a/tests/integration/shared_cache_test.go b/tests/integration/shared_cache_test.go index b665ccb6..f4a884d0 100644 --- a/tests/integration/shared_cache_test.go +++ b/tests/integration/shared_cache_test.go @@ -85,6 +85,7 @@ func bootRedisApp(t *testing.T, redisAddr, prefix string, timeout time.Duration) }}, Dedupe: config.Dedupe{Backend: config.DedupePebble}, Coord: config.Coord{Backend: config.CoordLocal}, + Roles: config.AllRoles(), Settings: config.Settings{Dir: settingsDir}, } a, err := app.New(ctx, app.Options{Config: cfg, Listener: ln}) From 712b1db250585dd0dfed112e76754763cdc50c8e Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:02:57 -0400 Subject: [PATCH 11/27] fix(config): refuse cache.redis.mode=sentinel until #656 Sentinel mode passed only the master set name: the backend does not authenticate to the sentinels, sets no topology refresh, and has no test (#656). Rather than make it selectable, validation refuses it and points at the issue; standalone and cluster are unchanged. cache.redis.sentinel_master is removed rather than kept and refused: a key that no config can use is noise in the reference, and the strict loader now reports it (and WH_CACHE_REDIS_SENTINEL_MASTER) as unknown, so nobody sets it believing it works. #656 can add it back with the sentinel credentials it also needs. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/configuration.mdx | 11 ++++----- docs/src/content/docs/deployment.md | 2 +- internal/app/wire.go | 1 - internal/config/cache_redis.go | 30 ++++++++++++------------- internal/config/cache_redis_test.go | 17 +++++++------- 7 files changed, 30 insertions(+), 35 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index f8f2efab..ac09db1a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone`, `cluster`, `sentinel`), `sentinel_master`, `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. +- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. - **A Redis-compatible shared cache backend** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row open a circuit breaker that skips the server until a probe succeeds, and deferred invalidations are retried until they land, collapsing to one tenant-wide bump per tenant past 100,000 keys. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. `cache.backend: redis` selects it (the entry above). Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression and server-stops-answering cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 2aba0526..4c627ee7 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -126,7 +126,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — the implementation each layer runs on, the process's `roles`, resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `cache.redis.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). - **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load` when `Config.NeedsDataDir` says a selected backend keeps state there, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. - **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` in `Validate`, between `validateRoles` and `validateTopology`), which refuses a value not on the list and names the ones that are. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns what a valid configuration is still likely to get wrong — the combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), a `cache.redis` block that is not read, certificate verification turned off — which `app.New` logs at `WARN`. -- **cache_redis.go** — `CacheRedisConfig`, the `cache.redis` sub-block, and its checks: an address, each `host:port`, a known mode, a sentinel's master name, `db` 0 in cluster mode, positive timeouts and sizes, a `version_ttl` of at least 2 s, and a `compress_min_bytes` that is not negative (`0` never compresses); its defaults are in `defaults()` with the rest. `CacheRedisTLS.Config` builds the `tls.Config`, reading the files; `Validate` calls it so an unreadable file refuses boot, and `internal/app` calls it again to build the connection. A TLS key set while `tls.enabled` is off is an error rather than a plaintext connection. +- **cache_redis.go** — `CacheRedisConfig`, the `cache.redis` sub-block, and its checks: an address, each `host:port`, a known mode (`sentinel` is refused until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `db` 0 in cluster mode, positive timeouts and sizes, a `version_ttl` of at least 2 s, and a `compress_min_bytes` that is not negative (`0` never compresses); its defaults are in `defaults()` with the rest. `CacheRedisTLS.Config` builds the `tls.Config`, reading the files; `Validate` calls it so an unreadable file refuses boot, and `internal/app` calls it again to build the connection. A TLS key set while `tls.enabled` is off is an error rather than a plaintext connection. - **config.go**, roles — `roles` (`[]Role`: `api`, `ingest`, `sweeper`; `AllRoles` by default; `Has(Role)`) picks which components `internal/app` wires, and `instance_id` names the process (`-<8 hex>` when empty, resolved in `Load`; today only logged at boot, and a distributed coordinator will record it as a lease's holder). `validateRoles` refuses an empty list, an empty entry, an unknown or a repeated role; `validateTopology` refuses a role set the backends cannot serve: any split over the embedded MQ, and a process with exactly one of `api` and `ingest` over a local cache. `NeedsDataDir` counts Pebble only for a process running `api`, and `Warnings` is empty without `api`, since only that role opens a cache it reads or a dedupe store. - **strict.go** — `rejectUnknownKeys`, the YAML half: re-reads the file as a generic tree and walks it against the struct's `yaml` tags, listing every key the struct doesn't declare. cleanenv itself is lenient by design, which is exactly wrong for boot config once keys have moved to the settings directory. - **persistence.go** — `WarnIfFreshDataDir` logs the startup `WARN` when `data_dir` is missing or empty (on a redeploy, the sign that the volume didn't persist); `LogStorageInitError` attaches the UID-65532 `permissionHint` to a NATS or Pebble open failure that looks like a permission denial — the same hint string `CheckDataDir` uses. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 6425d6f7..700045fc 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -150,12 +150,11 @@ The `redis` backend's settings, read only when `cache.backend` is `redis`. It is | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | -| `cache.redis.addrs` | `WH_CACHE_REDIS_ADDRS` | *(empty)* | **Required** with `backend: redis`. `host:port` of the server; several are a cluster's seed nodes or the sentinels. Comma-separated in the env var. | -| `cache.redis.mode` | `WH_CACHE_REDIS_MODE` | `standalone` | `standalone`, `cluster` or `sentinel`. | -| `cache.redis.sentinel_master` | `WH_CACHE_REDIS_SENTINEL_MASTER` | *(empty)* | The master set name. Required with `mode: sentinel`. | +| `cache.redis.addrs` | `WH_CACHE_REDIS_ADDRS` | *(empty)* | **Required** with `backend: redis`. `host:port` of the server; several are a cluster's seed nodes. Comma-separated in the env var. | +| `cache.redis.mode` | `WH_CACHE_REDIS_MODE` | `standalone` | `standalone` or `cluster`. `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656): the cache does not yet authenticate to the sentinels or refresh their topology, so it could not be trusted to follow a failover. | | `cache.redis.username` | `WH_CACHE_REDIS_USERNAME` | *(empty)* | ACL user. Empty uses the server's `default` user. | | `cache.redis.password` | `WH_CACHE_REDIS_PASSWORD` | *(empty)* | A secret: set it through the environment (or your secret store's env injection), not in a tracked `config.yaml`. | -| `cache.redis.db` | `WH_CACHE_REDIS_DB` | `0` | Database number (`SELECT`). Standalone and sentinel only: a cluster has only database `0`, and any other value refuses boot. | +| `cache.redis.db` | `WH_CACHE_REDIS_DB` | `0` | Database number (`SELECT`). Standalone only: a cluster has only database `0`, and any other value refuses boot. | | `cache.redis.tls.enabled` | `WH_CACHE_REDIS_TLS_ENABLED` | `false` | Connect over TLS, verifying the server against the system roots or `ca_file`. Any other `tls` key set while this is off refuses boot, rather than connecting in plaintext. | | `cache.redis.tls.ca_file` | `WH_CACHE_REDIS_TLS_CA_FILE` | *(empty)* | PEM file of the authorities to trust instead of the system roots. | | `cache.redis.tls.cert_file` | `WH_CACHE_REDIS_TLS_CERT_FILE` | *(empty)* | Client certificate (PEM) for mutual TLS. Set together with `key_file`. | @@ -268,8 +267,7 @@ cache: l1_max_cost: 67108864 # the local backend's size redis: # read only with backend: redis addrs: [] # required with backend: redis, e.g. ["redis:6379"] - mode: standalone # standalone | cluster | sentinel - sentinel_master: "" + mode: standalone # standalone | cluster username: "" password: "" # a secret: set WH_CACHE_REDIS_PASSWORD instead db: 0 @@ -348,7 +346,6 @@ WH_CACHE_L1_MAX_COST=67108864 # Read only with WH_CACHE_BACKEND=redis; WH_CACHE_REDIS_ADDRS is then required. WH_CACHE_REDIS_ADDRS= WH_CACHE_REDIS_MODE=standalone -WH_CACHE_REDIS_SENTINEL_MASTER= WH_CACHE_REDIS_USERNAME= WH_CACHE_REDIS_PASSWORD= WH_CACHE_REDIS_DB=0 diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 2acf348c..92c0cb95 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -427,7 +427,7 @@ The folder name is the tenant id, and each folder is a complete settings directo Several WaveHouse instances can serve one ClickHouse behind a load balancer, but most of what each one holds is its own. The message queue is embedded, so an event is inserted by the instance that took its `POST /v1/ingest`, and reaches only that instance's SSE subscribers. The dedupe store is per instance too, so an id one instance has seen is new to another. -The query-result cache is the layer that can be shared today. With the default `cache.backend: local`, each instance caches in its own memory, and an insert invalidates only the cache of the instance that made it. Every other instance keeps serving its cached results for the rows before the insert until each entry's TTL runs out, between 10 s and 1 h depending on how long the query took. With [`cache.backend: redis`](/configuration#cache), every instance reads and fills one Redis-compatible server, and an insert on any instance invalidates the cached results of every instance. +The query-result cache is the layer that can be shared today. With the default `cache.backend: local`, each instance caches in its own memory, and an insert invalidates only the cache of the instance that made it. Every other instance keeps serving its cached results for the rows before the insert until each entry's TTL runs out, between 10 s and 1 h depending on how long the query took. With [`cache.backend: redis`](/configuration#cache), every instance reads and fills one Redis-compatible server, and an insert on any instance invalidates the cached results of every instance. The server is a standalone one or a Redis Cluster; Sentinel (`mode: sentinel`) refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656), since the cache does not yet authenticate to the sentinels or refresh their topology. **What another instance can see.** Ingest is already asynchronous: `/v1/ingest` answers before the batch is inserted. Once the inserting instance's worker has written the batch to ClickHouse, it replaces the table's version token in Redis, and from then on a lookup on any instance misses and reads the new rows. The cache adds no delay of its own beyond that single write. The exceptions: diff --git a/internal/app/wire.go b/internal/app/wire.go index ea23606d..2ba112a4 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -667,7 +667,6 @@ func redisConfig(r config.CacheRedisConfig) (cache.RedisConfig, error) { return cache.RedisConfig{ Addrs: r.Addrs, Mode: r.Mode, - SentinelMaster: r.SentinelMaster, Username: r.Username, Password: r.Password, DB: r.DB, diff --git a/internal/config/cache_redis.go b/internal/config/cache_redis.go index 997d19a3..83ff056c 100644 --- a/internal/config/cache_redis.go +++ b/internal/config/cache_redis.go @@ -11,7 +11,8 @@ import ( "time" ) -// Redis deployment modes for cache.redis.mode. +// Redis deployment modes for cache.redis.mode. RedisSentinel is refused +// until the backend supports it (#656). const ( RedisStandalone = "standalone" RedisCluster = "cluster" @@ -22,15 +23,13 @@ const ( // (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB) shared by every process. // Read only when that backend is selected. type CacheRedisConfig struct { - // Addrs are host:port pairs: the server, or seeds for a cluster, or the - // sentinels. - Addrs []string `yaml:"addrs" env:"WH_CACHE_REDIS_ADDRS"` - Mode string `yaml:"mode" env:"WH_CACHE_REDIS_MODE"` - SentinelMaster string `yaml:"sentinel_master" env:"WH_CACHE_REDIS_SENTINEL_MASTER"` - Username string `yaml:"username" env:"WH_CACHE_REDIS_USERNAME"` - Password string `yaml:"password" env:"WH_CACHE_REDIS_PASSWORD"` - DB int `yaml:"db" env:"WH_CACHE_REDIS_DB"` - TLS CacheRedisTLS `yaml:"tls"` + // Addrs are host:port pairs: the server, or seeds for a cluster. + Addrs []string `yaml:"addrs" env:"WH_CACHE_REDIS_ADDRS"` + Mode string `yaml:"mode" env:"WH_CACHE_REDIS_MODE"` + Username string `yaml:"username" env:"WH_CACHE_REDIS_USERNAME"` + Password string `yaml:"password" env:"WH_CACHE_REDIS_PASSWORD"` + DB int `yaml:"db" env:"WH_CACHE_REDIS_DB"` + TLS CacheRedisTLS `yaml:"tls"` // KeyPrefix leads every key, so deployments can share one server. KeyPrefix string `yaml:"key_prefix" env:"WH_CACHE_REDIS_KEY_PREFIX"` Timeout time.Duration `yaml:"timeout" env:"WH_CACHE_REDIS_TIMEOUT"` @@ -61,7 +60,7 @@ func (r CacheRedisConfig) hasAddrs() bool { func (r CacheRedisConfig) validate() error { if !r.hasAddrs() { - return errors.New("cache.backend=redis needs cache.redis.addrs (WH_CACHE_REDIS_ADDRS): the server's host:port, or a cluster's seeds, or the sentinels") + return errors.New("cache.backend=redis needs cache.redis.addrs (WH_CACHE_REDIS_ADDRS): the server's host:port, or a cluster's seeds") } for _, a := range r.Addrs { if strings.TrimSpace(a) != a { @@ -72,12 +71,11 @@ func (r CacheRedisConfig) validate() error { } } switch r.Mode { - case RedisStandalone, RedisCluster, RedisSentinel: + case RedisStandalone, RedisCluster: + case RedisSentinel: + return fmt.Errorf("cache.redis.mode (WH_CACHE_REDIS_MODE) %q is not supported yet: the cache neither authenticates to the sentinels nor refreshes their topology (https://github.com/Wave-RF/WaveHouse/issues/656); valid: %s, %s", r.Mode, RedisStandalone, RedisCluster) default: - return fmt.Errorf("cache.redis.mode (WH_CACHE_REDIS_MODE) %q: valid: %s, %s, %s", r.Mode, RedisStandalone, RedisCluster, RedisSentinel) - } - if r.Mode == RedisSentinel && r.SentinelMaster == "" { - return errors.New("cache.redis.mode=sentinel needs cache.redis.sentinel_master (WH_CACHE_REDIS_SENTINEL_MASTER), the master set name") + return fmt.Errorf("cache.redis.mode (WH_CACHE_REDIS_MODE) %q: valid: %s, %s", r.Mode, RedisStandalone, RedisCluster) } if r.DB < 0 { return fmt.Errorf("cache.redis.db (WH_CACHE_REDIS_DB) %d is negative", r.DB) diff --git a/internal/config/cache_redis_test.go b/internal/config/cache_redis_test.go index 5389715a..d9d3adab 100644 --- a/internal/config/cache_redis_test.go +++ b/internal/config/cache_redis_test.go @@ -45,9 +45,8 @@ func TestLoad_CacheRedisFromEnv(t *testing.T) { caFile, certFile, keyFile := writeTestPKI(t, dir) for k, v := range map[string]string{ "WH_CACHE_BACKEND": "redis", - "WH_CACHE_REDIS_ADDRS": "s1:26379,s2:26379", - "WH_CACHE_REDIS_MODE": "sentinel", - "WH_CACHE_REDIS_SENTINEL_MASTER": "mymaster", + "WH_CACHE_REDIS_ADDRS": "r1:6379,r2:6379", + "WH_CACHE_REDIS_MODE": "standalone", "WH_CACHE_REDIS_USERNAME": "wavehouse", "WH_CACHE_REDIS_PASSWORD": "s3cret", "WH_CACHE_REDIS_DB": "2", @@ -69,7 +68,7 @@ func TestLoad_CacheRedisFromEnv(t *testing.T) { cfg, err := Load("nonexistent.yaml") require.NoError(t, err) assert.Equal(t, CacheRedisConfig{ - Addrs: []string{"s1:26379", "s2:26379"}, Mode: RedisSentinel, SentinelMaster: "mymaster", + Addrs: []string{"r1:6379", "r2:6379"}, Mode: RedisStandalone, Username: "wavehouse", Password: "s3cret", DB: 2, TLS: CacheRedisTLS{ Enabled: true, CAFile: caFile, CertFile: certFile, KeyFile: keyFile, ServerName: "redis.internal", @@ -142,6 +141,7 @@ cache: addr: r:6379 near_cache: max_cost: 1 + sentinel_master: mymaster tls: ca: /x memcached: @@ -149,7 +149,7 @@ cache: `), 0o600)) _, err := Load(path) require.Error(t, err) - assert.Contains(t, err.Error(), "cache.memcached, cache.redis.addr, cache.redis.near_cache, cache.redis.tls.ca") + assert.Contains(t, err.Error(), "cache.memcached, cache.redis.addr, cache.redis.near_cache, cache.redis.sentinel_master, cache.redis.tls.ca") } // The documented env file lists WH_CACHE_REDIS_ADDRS blank: that is no @@ -172,6 +172,7 @@ func TestUnboundEnv_KnowsTheCacheRedisVariables(t *testing.T) { "WH_CACHE_REDIS_VERSION_TTL=1h", "WH_CACHE_REDIS_COMPRESS_MIN_BYTES=0", })) assert.Equal(t, []string{"WH_CACHE_REDIS_ADDR"}, unboundEnv([]string{"WH_CACHE_REDIS_ADDR=r:6379"})) + assert.Equal(t, []string{"WH_CACHE_REDIS_SENTINEL_MASTER"}, unboundEnv([]string{"WH_CACHE_REDIS_SENTINEL_MASTER=m"}), "no sentinel mode until #656") } func TestValidate_CacheRedis(t *testing.T) { @@ -189,9 +190,9 @@ func TestValidate_CacheRedis(t *testing.T) { {"no addrs", func(r *CacheRedisConfig) { r.Addrs = nil }, "cache.backend=redis needs cache.redis.addrs (WH_CACHE_REDIS_ADDRS)"}, {"addr with space", func(r *CacheRedisConfig) { r.Addrs = []string{"a:6379", " b:6379"} }, "no spaces around an address"}, {"addr without port", func(r *CacheRedisConfig) { r.Addrs = []string{"redis"} }, `cache.redis.addrs (WH_CACHE_REDIS_ADDRS) "redis": want host:port`}, - {"mode", func(r *CacheRedisConfig) { r.Mode = "replica" }, `cache.redis.mode (WH_CACHE_REDIS_MODE) "replica": valid: standalone, cluster, sentinel`}, - {"sentinel without master", func(r *CacheRedisConfig) { r.Mode = RedisSentinel }, "needs cache.redis.sentinel_master"}, - {"sentinel", func(r *CacheRedisConfig) { r.Mode, r.SentinelMaster = RedisSentinel, "m" }, ""}, + {"mode", func(r *CacheRedisConfig) { r.Mode = "replica" }, `cache.redis.mode (WH_CACHE_REDIS_MODE) "replica": valid: standalone, cluster`}, + {"sentinel refused", func(r *CacheRedisConfig) { r.Mode = RedisSentinel }, `cache.redis.mode (WH_CACHE_REDIS_MODE) "sentinel" is not supported yet: the cache neither authenticates to the sentinels nor refreshes their topology (https://github.com/Wave-RF/WaveHouse/issues/656)`}, + {"cluster", func(r *CacheRedisConfig) { r.Mode = RedisCluster }, ""}, {"cluster db", func(r *CacheRedisConfig) { r.Mode, r.DB = RedisCluster, 1 }, "a Redis cluster has only database 0"}, {"standalone db", func(r *CacheRedisConfig) { r.DB = 3 }, ""}, {"negative db", func(r *CacheRedisConfig) { r.DB = -1 }, "is negative"}, From a7ef71fccd3d13f647d21e2a9c12ff7a42363692 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:03:55 -0400 Subject: [PATCH 12/27] fix(config): refuse a URL-style cache.redis address without echoing it WH_CACHE_REDIS_ADDRS=redis://default:s3cret@redis:6379 failed boot with an error that quoted the address, password included, twice: once in the key's message and once in net.SplitHostPort's. An address holding "://" or "@" is now refused by its position alone, pointing at the username, password and TLS settings that take those parts. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/configuration.mdx | 2 +- internal/config/cache_redis.go | 7 ++++++- internal/config/cache_redis_test.go | 21 +++++++++++++++++++++ 5 files changed, 30 insertions(+), 4 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index ac09db1a..471ccc17 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. +- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. - **A Redis-compatible shared cache backend** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row open a circuit breaker that skips the server until a probe succeeds, and deferred invalidations are retried until they land, collapsing to one tenant-wide bump per tenant past 100,000 keys. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. `cache.backend: redis` selects it (the entry above). Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression and server-stops-answering cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 4c627ee7..c0f83326 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -126,7 +126,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — the implementation each layer runs on, the process's `roles`, resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `cache.redis.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). - **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load` when `Config.NeedsDataDir` says a selected backend keeps state there, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. - **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` in `Validate`, between `validateRoles` and `validateTopology`), which refuses a value not on the list and names the ones that are. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns what a valid configuration is still likely to get wrong — the combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), a `cache.redis` block that is not read, certificate verification turned off — which `app.New` logs at `WARN`. -- **cache_redis.go** — `CacheRedisConfig`, the `cache.redis` sub-block, and its checks: an address, each `host:port`, a known mode (`sentinel` is refused until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `db` 0 in cluster mode, positive timeouts and sizes, a `version_ttl` of at least 2 s, and a `compress_min_bytes` that is not negative (`0` never compresses); its defaults are in `defaults()` with the rest. `CacheRedisTLS.Config` builds the `tls.Config`, reading the files; `Validate` calls it so an unreadable file refuses boot, and `internal/app` calls it again to build the connection. A TLS key set while `tls.enabled` is off is an error rather than a plaintext connection. +- **cache_redis.go** — `CacheRedisConfig`, the `cache.redis` sub-block, and its checks: an address, each `host:port` (a URL or `user:password@` form refused without repeating it, since it may hold a password), a known mode (`sentinel` is refused until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `db` 0 in cluster mode, positive timeouts and sizes, a `version_ttl` of at least 2 s, and a `compress_min_bytes` that is not negative (`0` never compresses); its defaults are in `defaults()` with the rest. `CacheRedisTLS.Config` builds the `tls.Config`, reading the files; `Validate` calls it so an unreadable file refuses boot, and `internal/app` calls it again to build the connection. A TLS key set while `tls.enabled` is off is an error rather than a plaintext connection. - **config.go**, roles — `roles` (`[]Role`: `api`, `ingest`, `sweeper`; `AllRoles` by default; `Has(Role)`) picks which components `internal/app` wires, and `instance_id` names the process (`-<8 hex>` when empty, resolved in `Load`; today only logged at boot, and a distributed coordinator will record it as a lease's holder). `validateRoles` refuses an empty list, an empty entry, an unknown or a repeated role; `validateTopology` refuses a role set the backends cannot serve: any split over the embedded MQ, and a process with exactly one of `api` and `ingest` over a local cache. `NeedsDataDir` counts Pebble only for a process running `api`, and `Warnings` is empty without `api`, since only that role opens a cache it reads or a dedupe store. - **strict.go** — `rejectUnknownKeys`, the YAML half: re-reads the file as a generic tree and walks it against the struct's `yaml` tags, listing every key the struct doesn't declare. cleanenv itself is lenient by design, which is exactly wrong for boot config once keys have moved to the settings directory. - **persistence.go** — `WarnIfFreshDataDir` logs the startup `WARN` when `data_dir` is missing or empty (on a redeploy, the sign that the volume didn't persist); `LogStorageInitError` attaches the UID-65532 `permissionHint` to a NATS or Pebble open failure that looks like a permission denial — the same hint string `CheckDataDir` uses. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 700045fc..f15788f3 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -150,7 +150,7 @@ The `redis` backend's settings, read only when `cache.backend` is `redis`. It is | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | -| `cache.redis.addrs` | `WH_CACHE_REDIS_ADDRS` | *(empty)* | **Required** with `backend: redis`. `host:port` of the server; several are a cluster's seed nodes. Comma-separated in the env var. | +| `cache.redis.addrs` | `WH_CACHE_REDIS_ADDRS` | *(empty)* | **Required** with `backend: redis`. `host:port` of the server; several are a cluster's seed nodes. Comma-separated in the env var. Not a URL: `redis://user:password@host:port` refuses boot, without repeating the value, so set the credentials through `username` and `password`, and `rediss://` through `tls.enabled`. | | `cache.redis.mode` | `WH_CACHE_REDIS_MODE` | `standalone` | `standalone` or `cluster`. `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656): the cache does not yet authenticate to the sentinels or refresh their topology, so it could not be trusted to follow a failover. | | `cache.redis.username` | `WH_CACHE_REDIS_USERNAME` | *(empty)* | ACL user. Empty uses the server's `default` user. | | `cache.redis.password` | `WH_CACHE_REDIS_PASSWORD` | *(empty)* | A secret: set it through the environment (or your secret store's env injection), not in a tracked `config.yaml`. | diff --git a/internal/config/cache_redis.go b/internal/config/cache_redis.go index 83ff056c..0e41b0ff 100644 --- a/internal/config/cache_redis.go +++ b/internal/config/cache_redis.go @@ -62,7 +62,12 @@ func (r CacheRedisConfig) validate() error { if !r.hasAddrs() { return errors.New("cache.backend=redis needs cache.redis.addrs (WH_CACHE_REDIS_ADDRS): the server's host:port, or a cluster's seeds") } - for _, a := range r.Addrs { + for i, a := range r.Addrs { + // A redis:// URL, or user:pass@host, may carry a password: refuse it + // without echoing it into the boot error and the logs. + if strings.Contains(a, "://") || strings.Contains(a, "@") { + return fmt.Errorf("cache.redis.addrs (WH_CACHE_REDIS_ADDRS) entry %d is a URL or holds credentials (not echoed): give host:port, and set the user and password with WH_CACHE_REDIS_USERNAME and WH_CACHE_REDIS_PASSWORD, and TLS (rediss://) with WH_CACHE_REDIS_TLS_ENABLED", i+1) + } if strings.TrimSpace(a) != a { return fmt.Errorf("cache.redis.addrs (WH_CACHE_REDIS_ADDRS) %q: no spaces around an address", a) } diff --git a/internal/config/cache_redis_test.go b/internal/config/cache_redis_test.go index d9d3adab..21ec15fc 100644 --- a/internal/config/cache_redis_test.go +++ b/internal/config/cache_redis_test.go @@ -165,6 +165,27 @@ func TestLoad_CacheRedisBlankAddrs(t *testing.T) { require.ErrorContains(t, err, "cache.backend=redis needs cache.redis.addrs") } +// A URL-style address carries its password, and a boot error reaches the +// logs: the refusal names the entry, never its value. +func TestLoad_CacheRedisRefusesURLAddrsWithoutEchoingThem(t *testing.T) { + for _, addr := range []string{ + "redis://default:s3cret@redis:6379", + "rediss://default:s3cret@redis:6380", + "default:s3cret@redis:6379", + "redis://redis:6379", + } { + t.Run(addr, func(t *testing.T) { + t.Setenv("WH_CACHE_BACKEND", "redis") + t.Setenv("WH_CACHE_REDIS_ADDRS", "ok:6379,"+addr) + _, err := Load("nonexistent.yaml") + require.ErrorContains(t, err, "cache.redis.addrs (WH_CACHE_REDIS_ADDRS) entry 2 is a URL or holds credentials") + assert.Contains(t, err.Error(), "WH_CACHE_REDIS_USERNAME and WH_CACHE_REDIS_PASSWORD") + assert.NotContains(t, err.Error(), "s3cret") + assert.NotContains(t, err.Error(), addr) + }) + } +} + func TestUnboundEnv_KnowsTheCacheRedisVariables(t *testing.T) { t.Parallel() assert.Empty(t, unboundEnv([]string{ From de27b600f670ec6ae53777836fd3a1f85951eb1c Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:04:55 -0400 Subject: [PATCH 13/27] fix(config): cap cache.redis.dial_timeout at 2s A dial is a TCP connect and a handshake, each bounded by dial_timeout. NewRedis dials once synchronously, so boot waits out one, and the cache's Close waits out a reconnect dial in flight: rueidis.NewClient takes no context. An unbounded dial_timeout could therefore push the cache's close past the 5s release budget and leave the queue, dedupe and ClickHouse pools unreleased. At 2s the wait is at most 4s; the default (1s) is unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/configuration.mdx | 2 +- internal/config/cache_redis.go | 9 +++++++++ internal/config/cache_redis_test.go | 6 ++++-- 5 files changed, 16 insertions(+), 5 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 471ccc17..1454d47a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. +- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`, at most `2s`: boot and shutdown each wait out a dial), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. - **A Redis-compatible shared cache backend** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row open a circuit breaker that skips the server until a probe succeeds, and deferred invalidations are retried until they land, collapsing to one tenant-wide bump per tenant past 100,000 keys. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. `cache.backend: redis` selects it (the entry above). Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression and server-stops-answering cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index c0f83326..cd44c8d1 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -126,7 +126,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — the implementation each layer runs on, the process's `roles`, resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `cache.redis.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). - **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load` when `Config.NeedsDataDir` says a selected backend keeps state there, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. - **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` in `Validate`, between `validateRoles` and `validateTopology`), which refuses a value not on the list and names the ones that are. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns what a valid configuration is still likely to get wrong — the combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), a `cache.redis` block that is not read, certificate verification turned off — which `app.New` logs at `WARN`. -- **cache_redis.go** — `CacheRedisConfig`, the `cache.redis` sub-block, and its checks: an address, each `host:port` (a URL or `user:password@` form refused without repeating it, since it may hold a password), a known mode (`sentinel` is refused until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `db` 0 in cluster mode, positive timeouts and sizes, a `version_ttl` of at least 2 s, and a `compress_min_bytes` that is not negative (`0` never compresses); its defaults are in `defaults()` with the rest. `CacheRedisTLS.Config` builds the `tls.Config`, reading the files; `Validate` calls it so an unreadable file refuses boot, and `internal/app` calls it again to build the connection. A TLS key set while `tls.enabled` is off is an error rather than a plaintext connection. +- **cache_redis.go** — `CacheRedisConfig`, the `cache.redis` sub-block, and its checks: an address, each `host:port` (a URL or `user:password@` form refused without repeating it, since it may hold a password), a known mode (`sentinel` is refused until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `db` 0 in cluster mode, positive timeouts and sizes, a `dial_timeout` of at most 2 s (boot and `Close` each wait out a dial, a connect and a handshake bounded by it), a `version_ttl` of at least 2 s, and a `compress_min_bytes` that is not negative (`0` never compresses); its defaults are in `defaults()` with the rest. `CacheRedisTLS.Config` builds the `tls.Config`, reading the files; `Validate` calls it so an unreadable file refuses boot, and `internal/app` calls it again to build the connection. A TLS key set while `tls.enabled` is off is an error rather than a plaintext connection. - **config.go**, roles — `roles` (`[]Role`: `api`, `ingest`, `sweeper`; `AllRoles` by default; `Has(Role)`) picks which components `internal/app` wires, and `instance_id` names the process (`-<8 hex>` when empty, resolved in `Load`; today only logged at boot, and a distributed coordinator will record it as a lease's holder). `validateRoles` refuses an empty list, an empty entry, an unknown or a repeated role; `validateTopology` refuses a role set the backends cannot serve: any split over the embedded MQ, and a process with exactly one of `api` and `ingest` over a local cache. `NeedsDataDir` counts Pebble only for a process running `api`, and `Warnings` is empty without `api`, since only that role opens a cache it reads or a dedupe store. - **strict.go** — `rejectUnknownKeys`, the YAML half: re-reads the file as a generic tree and walks it against the struct's `yaml` tags, listing every key the struct doesn't declare. cleanenv itself is lenient by design, which is exactly wrong for boot config once keys have moved to the settings directory. - **persistence.go** — `WarnIfFreshDataDir` logs the startup `WARN` when `data_dir` is missing or empty (on a redeploy, the sign that the volume didn't persist); `LogStorageInitError` attaches the UID-65532 `permissionHint` to a NATS or Pebble open failure that looks like a permission denial — the same hint string `CheckDataDir` uses. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index f15788f3..e00fdaec 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -163,7 +163,7 @@ The `redis` backend's settings, read only when `cache.backend` is `redis`. It is | `cache.redis.tls.insecure_skip_verify` | `WH_CACHE_REDIS_TLS_INSECURE_SKIP_VERIFY` | `false` | Accept any server certificate. Logged at `WARN` at boot: whoever can intercept the connection can read and replace cached results. | | `cache.redis.key_prefix` | `WH_CACHE_REDIS_KEY_PREFIX` | `wh` | Leads every key, so several deployments can share one server, provided you trust each as much as the others: any of them can overwrite what the rest serve. No `{` or `}`. | | `cache.redis.timeout` | `WH_CACHE_REDIS_TIMEOUT` | `100ms` | Per operation. A lookup or fill that takes longer is a miss or a skipped fill, never a failed query. | -| `cache.redis.dial_timeout` | `WH_CACHE_REDIS_DIAL_TIMEOUT` | `1s` | Per connection attempt. | +| `cache.redis.dial_timeout` | `WH_CACHE_REDIS_DIAL_TIMEOUT` | `1s` | Per connection attempt, at most `2s`. An attempt is a connect and a handshake, each bounded by this, and both boot and shutdown wait out one in flight, so the cap keeps a stop inside the release's fixed 5s budget (see [Stopping](/deployment#stopping)). | | `cache.redis.max_value_bytes` | `WH_CACHE_REDIS_MAX_VALUE_BYTES` | `1048576` | Largest result stored, after compression (1 MiB). A larger one is returned to the caller but not cached. | | `cache.redis.compress_min_bytes` | `WH_CACHE_REDIS_COMPRESS_MIN_BYTES` | `1024` | Results at least this large are zstd-compressed when that makes them smaller. `0` never compresses. | | `cache.redis.version_ttl` | `WH_CACHE_REDIS_VERSION_TTL` | `168h` | How long a table's or tenant's version token outlives its last write, so the tokens of dropped tables and removed tenants eventually expire. At least `2s`. An expired token only causes misses. | diff --git a/internal/config/cache_redis.go b/internal/config/cache_redis.go index 0e41b0ff..92063015 100644 --- a/internal/config/cache_redis.go +++ b/internal/config/cache_redis.go @@ -19,6 +19,12 @@ const ( RedisSentinel = "sentinel" ) +// maxRedisDialTimeout caps cache.redis.dial_timeout. A dial is a connect +// and a handshake, each bounded by it, and both boot and the cache's Close +// wait out one in flight: at 2s that is at most 4s, inside the 5s budget +// Close shares with the stores released after the cache. +const maxRedisDialTimeout = 2 * time.Second + // CacheRedisConfig configures cache.backend=redis: one Redis-compatible server // (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB) shared by every process. // Read only when that backend is selected. @@ -102,6 +108,9 @@ func (r CacheRedisConfig) validate() error { return fmt.Errorf("%s %s must be positive", d.key, d.v) } } + if r.DialTimeout > maxRedisDialTimeout { + return fmt.Errorf("cache.redis.dial_timeout (WH_CACHE_REDIS_DIAL_TIMEOUT) %s is over %s: boot and shutdown each wait out a dial, up to twice this", r.DialTimeout, maxRedisDialTimeout) + } if r.VersionTTL < 2*time.Second { return fmt.Errorf("cache.redis.version_ttl (WH_CACHE_REDIS_VERSION_TTL) %s is under 2s", r.VersionTTL) } diff --git a/internal/config/cache_redis_test.go b/internal/config/cache_redis_test.go index 21ec15fc..7f91a049 100644 --- a/internal/config/cache_redis_test.go +++ b/internal/config/cache_redis_test.go @@ -57,7 +57,7 @@ func TestLoad_CacheRedisFromEnv(t *testing.T) { "WH_CACHE_REDIS_TLS_SERVER_NAME": "redis.internal", "WH_CACHE_REDIS_KEY_PREFIX": "staging", "WH_CACHE_REDIS_TIMEOUT": "250ms", - "WH_CACHE_REDIS_DIAL_TIMEOUT": "3s", + "WH_CACHE_REDIS_DIAL_TIMEOUT": "2s", "WH_CACHE_REDIS_MAX_VALUE_BYTES": "2048", "WH_CACHE_REDIS_COMPRESS_MIN_BYTES": "0", "WH_CACHE_REDIS_VERSION_TTL": "24h", @@ -73,7 +73,7 @@ func TestLoad_CacheRedisFromEnv(t *testing.T) { TLS: CacheRedisTLS{ Enabled: true, CAFile: caFile, CertFile: certFile, KeyFile: keyFile, ServerName: "redis.internal", }, - KeyPrefix: "staging", Timeout: 250 * time.Millisecond, DialTimeout: 3 * time.Second, + KeyPrefix: "staging", Timeout: 250 * time.Millisecond, DialTimeout: 2 * time.Second, MaxValueBytes: 2048, CompressMinBytes: 0, VersionTTL: 24 * time.Hour, }, cfg.Cache.Redis) tc, err := cfg.Cache.Redis.TLS.Config() @@ -221,6 +221,8 @@ func TestValidate_CacheRedis(t *testing.T) { {"brace prefix", func(r *CacheRedisConfig) { r.KeyPrefix = "a{b}" }, "hash-tag brace"}, {"zero timeout", func(r *CacheRedisConfig) { r.Timeout = 0 }, "cache.redis.timeout (WH_CACHE_REDIS_TIMEOUT) 0s must be positive"}, {"negative dial timeout", func(r *CacheRedisConfig) { r.DialTimeout = -time.Second }, "cache.redis.dial_timeout"}, + {"dial timeout at the cap", func(r *CacheRedisConfig) { r.DialTimeout = 2 * time.Second }, ""}, + {"dial timeout over the cap", func(r *CacheRedisConfig) { r.DialTimeout = 2*time.Second + time.Millisecond }, "cache.redis.dial_timeout (WH_CACHE_REDIS_DIAL_TIMEOUT) 2.001s is over 2s"}, {"short version ttl", func(r *CacheRedisConfig) { r.VersionTTL = time.Second }, "cache.redis.version_ttl (WH_CACHE_REDIS_VERSION_TTL) 1s is under 2s"}, {"zero max value", func(r *CacheRedisConfig) { r.MaxValueBytes = 0 }, "cache.redis.max_value_bytes"}, {"compress never", func(r *CacheRedisConfig) { r.CompressMinBytes = 0 }, ""}, From 8189783eb2ab4972753c27357c3bedfd8351690a Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:08:18 -0400 Subject: [PATCH 14/27] test(cache): the local backend's insert-invalidates lifecycle, end to end With e2e on Redis, no test above unit level showed an ingest invalidating LocalCache, the default backend. The integration suite's own app runs cache.backend=local, so it gets the shared-cache test's twin: a fill is a hit, an ingested row is first served as a miss well inside the stale entry's TTL, and the refill is a hit again. With LocalCache.Invalidate made a no-op it fails: the row landed only ~10.0s after the fill, every round. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- CHANGELOG.md | 2 +- tests/integration/shared_cache_test.go | 42 ++++++++++++++++++++++++++ 2 files changed, 43 insertions(+), 1 deletion(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 1454d47a..48eb1014 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`, at most `2s`: boot and shutdown each wait out a dial), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. +- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`, at most `2s`: boot and shutdown each wait out a dial), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits; a third runs the first one's hit, insert and fresh-miss lifecycle on the suite's own `cache.backend: local` app, since e2e no longer exercises that backend. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. - **A Redis-compatible shared cache backend** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row open a circuit breaker that skips the server until a probe succeeds, and deferred invalidations are retried until they land, collapsing to one tenant-wide bump per tenant past 100,000 keys. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. `cache.backend: redis` selects it (the entry above). Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression and server-stops-answering cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/tests/integration/shared_cache_test.go b/tests/integration/shared_cache_test.go index f4a884d0..e9857b72 100644 --- a/tests/integration/shared_cache_test.go +++ b/tests/integration/shared_cache_test.go @@ -216,6 +216,48 @@ func TestSharedCache_IngestOnOneInstanceInvalidatesAnother(t *testing.T) { } } +// The same lifecycle on the default backend, against the suite's own app +// (cache.backend=local), which e2e no longer runs: a fill is a hit, and an +// ingest invalidates it, so the first answer carrying the new row is a miss +// well inside the stale entry's TTL, and the refill is a hit again. +func TestLocalCache_IngestInvalidates(t *testing.T) { + base := env(t).baseURL + // As above: a round whose row lands past the TTL floor proves nothing, + // and runs again on a fresh table, so a fresh fill. + for round := 1; ; round++ { + table := createTable(t, "user_id String, value Float64", "ORDER BY user_id") + status, xc, body := structuredQuery(t, base, table) + require.Equal(t, http.StatusOK, status, body) + require.Equal(t, "MISS", xc) + filled := time.Now() + status, xc, _ = structuredQuery(t, base, table) + require.Equal(t, http.StatusOK, status) + require.Equal(t, "HIT", xc) + + ingestRow(t, base, table, "u1") + var seenAt time.Time + require.Eventually(t, func() bool { + var err error + status, xc, body, err = tryStructuredQuery(base, table) + if err == nil && status == http.StatusOK && strings.Contains(body, "u1") { + seenAt = time.Now() + return true + } + return false + }, 30*time.Second, 100*time.Millisecond, "the ingested row is never served") + if seenAt.Sub(filled) < minCacheTTL-time.Second { + assert.Equal(t, "MISS", xc, "the first answer carrying the new row is a refill") + status, xc, body = structuredQuery(t, base, table) + require.Equal(t, http.StatusOK, status) + assert.Equal(t, "HIT", xc, "the refill is cached again") + assert.Contains(t, body, "u1") + return + } + require.Less(t, round, 3, "the new row was served only once the stale entry could have expired, in every round: the ingest never invalidated it, or ingest is too slow here to tell") + t.Logf("round %d: row landed %s after the fill, past the TTL floor; retrying", round, seenAt.Sub(filled)) + } +} + // A Redis that stops answering costs queries nothing but the cache: they // keep succeeding, straight from ClickHouse, each a miss; an ingest made // meanwhile is visible at once. Once it answers again, the cache serves hits. From 5e8476271d03db7fc5ca51aba68e3a48db5353c4 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:10:05 -0400 Subject: [PATCH 15/27] fix(ingest): log an invalidation that did not land at WARN With cache.backend=redis a bump the server does not take is deferred and retried, so a Redis outage logged an ERROR for every batch the worker inserted. It is a WARN now; wavehouse_cache_invalidations_pending is the signal to alert on. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- CHANGELOG.md | 2 +- internal/ingest/worker.go | 4 +++- 2 files changed, 4 insertions(+), 2 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 48eb1014..27ee2eaf 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`, at most `2s`: boot and shutdown each wait out a dial), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits; a third runs the first one's hit, insert and fresh-miss lifecycle on the suite's own `cache.backend: local` app, since e2e no longer exercises that backend. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. +- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `internal/ingest/worker.go`, `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`, at most `2s`: boot and shutdown each wait out a dial), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The ingest worker's log of an invalidation that did not land drops from `ERROR` to `WARN`, since the shared backend defers and retries it: an outage would otherwise log an `ERROR` for every batch. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits; a third runs the first one's hit, insert and fresh-miss lifecycle on the suite's own `cache.backend: local` app, since e2e no longer exercises that backend. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. - **A Redis-compatible shared cache backend** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row open a circuit breaker that skips the server until a probe succeeds, and deferred invalidations are retried until they land, collapsing to one tenant-wide bump per tenant past 100,000 keys. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. `cache.backend: redis` selects it (the entry above). Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression and server-stops-answering cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/internal/ingest/worker.go b/internal/ingest/worker.go index 70517cd2..529b8485 100644 --- a/internal/ingest/worker.go +++ b/internal/ingest/worker.go @@ -910,7 +910,9 @@ func (w *IngestWorker) invalidate(ctx context.Context, id tenant.ID, tableName s } invCtx := trace.ContextWithSpanContext(context.WithoutCancel(ctx), trace.SpanContextFromContext(ctx)) if _, err := w.cache.Invalidate(invCtx, namespaces); err != nil { - slog.ErrorContext(invCtx, "failed to invalidate cache after insert - your cache is holding stale data now!", "tenant", id, "table", tableName, "error", err) + // WARN, not ERROR: a shared backend defers and retries the bump, and + // an outage would otherwise log an ERROR for every batch. + slog.WarnContext(invCtx, "cache invalidation after insert did not land; the table's cached results may be stale until it does", "tenant", id, "table", tableName, "error", err) } } From 4df73f995fd14139088e310b2175faf27eebc69d Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:46:54 -0400 Subject: [PATCH 16/27] docs(cache): clarify which bumps a no-index tenant absorbs MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The version_manager.go bullet read "a bump of a tenant with no index is a no-op" right after introducing BumpTenant, so it looked scoped to that one call. It actually covers any bump — table, scope or tenant — against a tenant with no index, which is what lets Prune and a departed tenant's index stay gone (neither the sharedTables fan-out nor an insert still in flight for a just-pruned tenant can revive it). Found by pre-push review of #621 (origin/feat/cache-snapshot...HEAD). Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- docs/src/content/docs/architecture.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 4690761d..7aede5ba 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -114,7 +114,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on — one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)) — each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope by their raw names, which the cache escapes where it builds a key. `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The snapshot is taken before any input a bump invalidates is chosen, the tenant's connection included: a reload that moves the tenant to another address or database runs `Pools.Reconcile` and then `InvalidateTenant` (a repoint that keeps both, such as a username or `tls` change, reads the same tables and bumps nothing), so a request that took the old pool files the old database's rows under a version that bump orphans, whether its `Set` lands before the bump or after. The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. -- **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: one version per tenant, per (tenant, table) and per (tenant, table, scope), each keyed by its name alone and bumped in place, so the index holds one entry per live tenant, table and scope however often each is bumped ([#262](https://github.com/Wave-RF/WaveHouse/issues/262)). A query key folds the tenant's version and, for each dependency, its tenant's, table's and scope's, so bumping a table (a scopeless write) orphans every scope of it, and bumping one scope orphans that scope and the whole-table view — scope is reserved and empty today, so every write is the whole-table bump — all without touching the pool. Every field — the caller's query key, the tenant id, and each dependency's table and scope — is escaped and joined by `internal/keyenc` where the key is built, so a dot, a space or a `%` in a name is never read as a separator: each dependency renders as `..
.
..`, and the whole entry key is `|.||…`. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results. A tenant's version is a *generation*, unique within the process and handed out by the first key built for the tenant; `BumpTenant` (behind `InvalidateTenant`) drops the tenant's whole index, so the next key gets a fresh generation no cached entry folds, orphaning every cached result of the tenant in one step — a pipe result with no dependencies, and a table no bump ever keyed, included — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). `LocalCache.Prune` does the same for every tenant no longer served, which `internal/app` runs after each settings reload, so a tenant removed or rejected stops holding its index. A table bump drops the table's scope versions with it, since every key they were folded into also folds the old table version; and a bump of a tenant with no index is a no-op, since no key folds its next generation yet. The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. +- **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: one version per tenant, per (tenant, table) and per (tenant, table, scope), each keyed by its name alone and bumped in place, so the index holds one entry per live tenant, table and scope however often each is bumped ([#262](https://github.com/Wave-RF/WaveHouse/issues/262)). A query key folds the tenant's version and, for each dependency, its tenant's, table's and scope's, so bumping a table (a scopeless write) orphans every scope of it, and bumping one scope orphans that scope and the whole-table view — scope is reserved and empty today, so every write is the whole-table bump — all without touching the pool. Every field — the caller's query key, the tenant id, and each dependency's table and scope — is escaped and joined by `internal/keyenc` where the key is built, so a dot, a space or a `%` in a name is never read as a separator: each dependency renders as `..
.
..`, and the whole entry key is `|.||…`. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results. A tenant's version is a *generation*, unique within the process and handed out by the first key built for the tenant; `BumpTenant` (behind `InvalidateTenant`) drops the tenant's whole index, so the next key gets a fresh generation no cached entry folds, orphaning every cached result of the tenant in one step — a pipe result with no dependencies, and a table no bump ever keyed, included — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). `LocalCache.Prune` does the same for every tenant no longer served, which `internal/app` runs after each settings reload, so a tenant removed or rejected stops holding its index. A table bump drops the table's scope versions with it, since every key they were folded into also folds the old table version; and any bump (of a table, a scope or the tenant) under a tenant with no index is a no-op that records nothing, since the next key built for it gets a fresh generation no cached entry folds — so neither the `sharedTables` fan-out nor an insert still in flight for a tenant just pruned brings its index back. The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. ### `config/` — Configuration From 95575f27a49982cddd1c1febd72028ff669a901b Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:54:37 -0400 Subject: [PATCH 17/27] docs(cache): CHANGELOG says #262 is part-fixed, not fixed The #262 entry said "fixes #262 for the in-process cache", which would close the issue on merge (delete_branch_on_merge + squash_merge_commit_message: PR_BODY carry the PR body's own "Fixes #262" into main's history the same way). #262 has a residual left open on purpose: per-table scope cardinality isn't capped, since scope stays empty until #235 populates it. Reworded to "part of #262" and named the residual, matching the PR body and the issue comment that records the trigger. Found by pre-push review of #621 (origin/feat/cache-snapshot...HEAD). Co-Authored-By: Claude Sonnet 5 --- CHANGELOG.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index a89a0d2b..6631b3c4 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -89,7 +89,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **An unavailable ClickHouse is retried with backoff instead of dead-lettering every row** (`internal/chconn/errclass.go` (new, + tests), `internal/ingest/{worker,backoff}.go` (`backoff.go` new, + tests), `internal/mq/{mq,embedded}.go`, `internal/testutil/mocks.go`, `tests/integration/ingest_outage_test.go` (new), `AGENTS.md`, `README.md`, `docs/src/content/docs/{ingest-pipeline,architecture,api,deployment,why-wavehouse}.md`, `docs/src/content/docs/{settings-directory,index,access-control}.mdx`): workstream A of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). A failed batch insert used to go through row-by-row isolation whatever the failure, so a ClickHouse that was down, overloaded or read-only failed every row twice and parked the whole batch on the DLQ. `chconn.Classify` now classes the failure first — `Rejected` (any ClickHouse exception code outside the availability and credential lists: the server read the row and refused it), `Unavailable` (connection refused/reset, timeouts, `TOO_MANY_SIMULTANEOUS_QUERIES`, `SERVER_OVERLOADED`, `MEMORY_LIMIT_EXCEEDED`, `TOO_MANY_PARTS`, `READONLY`, `TABLE_IS_READ_ONLY`, `KEEPER_EXCEPTION`, …), `Denied` (`AUTHENTICATION_FAILED`, `ACCESS_DENIED`, …) or `Unknown` (no code, no recognizable transport failure). Only `Rejected` is isolated and dead-lettered as before, and a multi-row batch refused with `TOO_MANY_PARTS` or `MEMORY_LIMIT_EXCEEDED` is split row by row first (`chconn.Splittable`), because a batch spanning too many partitions or too much memory can fail where each of its rows inserts; every other class hands the batch back to the queue with a delayed nak (`mq.Message.NakWithDelay`, new) under a jittered 1 s → 30 s backoff shared by every table on the same ClickHouse pool (a failure of one table — read-only, too many parts or mutations, a grant missing on it, `chconn.TableScoped` — backs off that table alone), which turns rows away without a request while it runs and probes once per window, and ClickHouse going away mid-isolation stops isolation and retries the rows it had not settled. Counted by the new `wavehouse_ingest_retries_total{table, reason}`; logged at `WARN` when an outage starts and at most every 30 s during it. A long outage now shows as a growing ingest stream and, at `mq.max_bytes_gb`, ingest `503`s — not as a full DLQ; a lasting failure of one table holds back its tenant's other tables once its waiting rows reach `maxAckPending`. Retried rows come back out of arrival order, which matters only to a `ReplacingMergeTree` without a version column or a `CollapsingMergeTree`. - **Schema discovery's retry loop jitters its backoff** (`internal/discovery/discovery.go` (+ tests), `internal/app/wire.go`, `internal/api/errors.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`): `RetryRefresh` slept exactly `2s * 2^n` capped at 60s, so instances retrying against one recovering ClickHouse fired in lockstep, every 60s on the same second. Each sleep is now drawn uniformly from below the backoff (full jitter), spreading the retries over the whole window and halving the mean wait — so a failing tenant's retries, their log lines and `wavehouse_schema_refresh_failures_total` come about twice as often ([#141](https://github.com/Wave-RF/WaveHouse/issues/141)). - **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/query/ident.go` (removed, + tests), `internal/app/wire.go`, `docs/src/content/docs/{api,architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL (a pipe result's key folded no version, so pipes were unaffected; with the tenant's version in every key they now take the same snapshot). The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. The snapshot is taken before any input a bump invalidates is chosen, the tenant's ClickHouse connection included: both handlers now look up before they resolve the tenant's pool, so a reload that moves the tenant to another address or database after a request took the old pool orphans that request's fill instead of filing the old database's rows as fresh under the new tenant version. A tenant on no pool is still a `503` before a cached result is served or a query runs. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. The cache now escapes table and scope names itself: a `Namespace` carries them raw and the version index builds its keys with `internal/keyenc`, so neither the structured-query read nor the ingest worker's invalidation escapes them (`query.SafeEncodeToken` is gone) and no name reaches a key unescaped. The suite pins that a name holding a dot, a space or a `%` is read and bumped under one key, and that names which would run together unescaped (`a.0.b` against `a` with scope `b.0.`) stay two entries. The version index's own entries are unaffected by the escaping; only the rendered key changes: it now carries the tenant's version and escapes the caller's query key whole. All of them live in the process, so nothing stored is orphaned. -- **The cache's version index no longer grows with every bump, and forgets a tenant no longer served** (`internal/cache/{local,version_manager}.go` (+ tests), `internal/app/wire.go` (+ tests), `docs/src/content/docs/architecture.md`, `AGENTS.md`): fixes [#262](https://github.com/Wave-RF/WaveHouse/issues/262) for the in-process cache, part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The index nested each table under its tenant's version and each scope under its table's, and never pruned, so every tenant invalidation left the tenant's whole index behind, and it grew with every tenant ever served. It now holds one version per tenant, per (tenant, table) and per (tenant, table, scope), bumped in place. A tenant invalidation drops the tenant's index and hands its next key a generation unique within the process, so nothing cached before it can match again, and a table bump drops the table's scope versions. After each settings reload the index of every tenant no longer served, removed or rejected, is dropped the same way; its cached results are orphaned with it, as they already were when such a tenant came back on a pool. No change to what is cached or served. The Redis-compatible backend (#613) will bound its versions with a TTL instead. +- **The cache's version index no longer grows with every bump, and forgets a tenant no longer served** (`internal/cache/{local,version_manager}.go` (+ tests), `internal/app/wire.go` (+ tests), `docs/src/content/docs/architecture.md`, `AGENTS.md`): part of [#262](https://github.com/Wave-RF/WaveHouse/issues/262) (growth across bumps and a departed tenant's memory; per-table scope cardinality is left open, see the issue) and of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The index nested each table under its tenant's version and each scope under its table's, and never pruned, so every tenant invalidation left the tenant's whole index behind, and it grew with every tenant ever served. It now holds one version per tenant, per (tenant, table) and per (tenant, table, scope), bumped in place. A tenant invalidation drops the tenant's index and hands its next key a generation unique within the process, so nothing cached before it can match again, and a table bump drops the table's scope versions. After each settings reload the index of every tenant no longer served, removed or rejected, is dropped the same way; its cached results are orphaned with it, as they already were when such a tenant came back on a pool. No change to what is cached or served. The Redis-compatible backend (#613) will bound its versions with a TTL instead. - **An explicit `false`, `0` or `""` in `config.yaml` is no longer replaced by the key's default** (`internal/config/config.go`, `internal/config/defaults_test.go` (new), `docs/src/content/docs/configuration.mdx`, `AGENTS.md`): [#631](https://github.com/Wave-RF/WaveHouse/issues/631). Defaults lived in cleanenv `env-default` tags, which cleanenv applies after the YAML decode to any field still at its zero value, so it could not tell a key the file set to its zero value from one the file left out. `otel.traces.enabled: false`, `otel.metrics.enabled: false` and `otel.logs.enabled: false` came back `true`; `otel.traces.sample_rate: 0` and `otel.logs.sample_rate: 0` came back `1.0`; `server.shutdown_timeout: 0` came back `10`; `cache.l1_max_cost: 0`, `prometheus.path: ""` and `data_dir: ""` came back as their defaults; `server.port: 0` came back `8080`. All of it was silent. Defaults now live in one Go function, `defaults()`, which `Load` starts from before decoding the file and then applying `WH_*` variables, so the order is env > YAML > default and a key the file sets always wins. **Behaviour change if your file relied on the bug:** a zero you wrote now takes effect. A file that says `sample_rate: 0` now exports no traces (or no DEBUG/INFO logs), where it silently exported everything; a signal set `enabled: false` is now off; `shutdown_timeout: 0` now skips the drain. `cache.l1_max_cost: 0`, `server.port: 0`, and `data_dir: ""` now refuse boot (`cache init: MaxCost can't be zero`, `server.port 0 out of range`, `data_dir (WH_DATA_DIR) is required`) instead of running on the default; an empty `prometheus.path` refuses boot when `prometheus.enabled` is true. Delete the key to get the default back. Env vars are unchanged: they already honoured an explicit zero. New tests load through `config.Load` for every affected key (a YAML zero is kept, an absent key gets the default, env wins in both directions), refuse an `env-default` tag on any field, and pin each documented default in `configuration.mdx` to `defaults()`. - **An embedded queue store that cannot be created fails boot at once, naming the cause** (`internal/mq/embedded.go` (+ tests)): part of [#617](https://github.com/Wave-RF/WaveHouse/issues/617). A regular file at `/nats`, or a `nats` directory that could not be created there, failed JetStream in the background, so boot waited out the server's 5s readiness check and reported only `nats server not ready`. `NewEmbedded` now creates the directory first (at `0700`, as the server does) and refuses boot with the mkdir error. An existing but unwritable `nats` directory still takes the old path. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its cached results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. From 5b31ce6d97bc0b176e172219e2ee9cf63a7c4db0 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 06:23:37 -0400 Subject: [PATCH 18/27] fix(config): cap cache.redis timeout and dial_timeout at 1s Since #626 bounds a cluster client's topology read by the larger of the two, a dial held by boot or Close is a connect and a handshake, each up to dial_timeout, plus that read. At a 2s dial cap it could reach 6s, past the 5s release budget. Both caps are now 1s, so at most 3s. The shared-cache docs now describe #626's behavior: a refused write opens the breaker, the owing instance bypasses the lookups it would orphan, and connections are replaced every minute after a failover behind a stable address. Co-Authored-By: Claude Opus 5.5 (1M context) --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/configuration.mdx | 6 +++--- docs/src/content/docs/deployment.md | 6 +++--- internal/config/cache_redis.go | 23 ++++++++++++++++------- internal/config/cache_redis_test.go | 9 +++++---- tests/integration/shared_cache_test.go | 6 +++--- 7 files changed, 32 insertions(+), 22 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index ae2b6ac5..51ed8865 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `internal/ingest/worker.go`, `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`, at most `2s`: boot and shutdown each wait out a dial), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The ingest worker's log of an invalidation that did not land drops from `ERROR` to `WARN`, since the shared backend defers and retries it: an outage would otherwise log an `ERROR` for every batch. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits; a third runs the first one's hit, insert and fresh-miss lifecycle on the suite's own `cache.backend: local` app, since e2e no longer exercises that backend. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. +- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `internal/ingest/worker.go`, `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`) and `dial_timeout` (`1s`), each at most `1s` since boot and shutdown each wait out a connection attempt they bound, `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The ingest worker's log of an invalidation that did not land drops from `ERROR` to `WARN`, since the shared backend defers and retries it: an outage would otherwise log an `ERROR` for every batch. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits; a third runs the first one's hit, insert and fresh-miss lifecycle on the suite's own `cache.backend: local` app, since e2e no longer exercises that backend. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. - **A Redis-compatible shared cache backend** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/cache.go`, `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`, `CONTRIBUTING.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), open a circuit breaker that skips the server until a probe write succeeds (connections are replaced every minute, so after a failover behind a stable address the process reaches the new primary, and delivers the bumps it owes, within about a minute, provided a fresh connection's dial and handshake fit in the per-operation timeout), and deferred invalidations are retried until they land — the first at once, and at the probe's cadence while the breaker is open — collapsing to one tenant-wide bump per tenant past 100,000 keys; until one lands, the process that owes it bypasses the lookups it would orphan. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. `cache.backend: redis` selects it (the entry above). Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression, server-stops-answering, refused-writes (a demoted primary, a full `noeviction` server), failover-behind-a-stable-address and owed-invalidation cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index fdab8bbe..d3e47962 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -126,7 +126,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — the implementation each layer runs on, the process's `roles`, resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `cache.redis.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). - **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load` when `Config.NeedsDataDir` says a selected backend keeps state there, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. - **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` in `Validate`, between `validateRoles` and `validateTopology`), which refuses a value not on the list and names the ones that are. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns what a valid configuration is still likely to get wrong — the combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), a `cache.redis` block that is not read, certificate verification turned off — which `app.New` logs at `WARN`. -- **cache_redis.go** — `CacheRedisConfig`, the `cache.redis` sub-block, and its checks: an address, each `host:port` (a URL or `user:password@` form refused without repeating it, since it may hold a password), a known mode (`sentinel` is refused until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `db` 0 in cluster mode, positive timeouts and sizes, a `dial_timeout` of at most 2 s (boot and `Close` each wait out a dial, a connect and a handshake bounded by it), a `version_ttl` of at least 2 s, and a `compress_min_bytes` that is not negative (`0` never compresses); its defaults are in `defaults()` with the rest. `CacheRedisTLS.Config` builds the `tls.Config`, reading the files; `Validate` calls it so an unreadable file refuses boot, and `internal/app` calls it again to build the connection. A TLS key set while `tls.enabled` is off is an error rather than a plaintext connection. +- **cache_redis.go** — `CacheRedisConfig`, the `cache.redis` sub-block, and its checks: an address, each `host:port` (a URL or `user:password@` form refused without repeating it, since it may hold a password), a known mode (`sentinel` is refused until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `db` 0 in cluster mode, positive timeouts and sizes, a `timeout` and `dial_timeout` of at most 1 s each (boot and `Close` each wait out a dial: a connect and a handshake bounded by `dial_timeout`, and a cluster's topology read bounded by the larger of the two), a `version_ttl` of at least 2 s, and a `compress_min_bytes` that is not negative (`0` never compresses); its defaults are in `defaults()` with the rest. `CacheRedisTLS.Config` builds the `tls.Config`, reading the files; `Validate` calls it so an unreadable file refuses boot, and `internal/app` calls it again to build the connection. A TLS key set while `tls.enabled` is off is an error rather than a plaintext connection. - **config.go**, roles — `roles` (`[]Role`: `api`, `ingest`, `sweeper`; `AllRoles` by default; `Has(Role)`) picks which components `internal/app` wires, and `instance_id` names the process (`-<8 hex>` when empty, resolved in `Load`; today only logged at boot, and a distributed coordinator will record it as a lease's holder). `validateRoles` refuses an empty list, an empty entry, an unknown or a repeated role; `validateTopology` refuses a role set the backends cannot serve: any split over the embedded MQ, and a process with exactly one of `api` and `ingest` over a local cache. `NeedsDataDir` counts Pebble only for a process running `api`, and `Warnings` is empty without `api`, since only that role opens a cache it reads or a dedupe store. - **strict.go** — `rejectUnknownKeys`, the YAML half: re-reads the file as a generic tree and walks it against the struct's `yaml` tags, listing every key the struct doesn't declare. cleanenv itself is lenient by design, which is exactly wrong for boot config once keys have moved to the settings directory. - **persistence.go** — `WarnIfFreshDataDir` logs the startup `WARN` when `data_dir` is missing or empty (on a redeploy, the sign that the volume didn't persist); `LogStorageInitError` attaches the UID-65532 `permissionHint` to a NATS or Pebble open failure that looks like a permission denial — the same hint string `CheckDataDir` uses. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index d55936e5..509088ae 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -162,13 +162,13 @@ The `redis` backend's settings, read only when `cache.backend` is `redis`. It is | `cache.redis.tls.server_name` | `WH_CACHE_REDIS_TLS_SERVER_NAME` | *(empty)* | Name to verify the server's certificate against, when it differs from the address. | | `cache.redis.tls.insecure_skip_verify` | `WH_CACHE_REDIS_TLS_INSECURE_SKIP_VERIFY` | `false` | Accept any server certificate. Logged at `WARN` at boot: whoever can intercept the connection can read and replace cached results. | | `cache.redis.key_prefix` | `WH_CACHE_REDIS_KEY_PREFIX` | `wh` | Leads every key, so several deployments can share one server, provided you trust each as much as the others: any of them can overwrite what the rest serve. No `{` or `}`. | -| `cache.redis.timeout` | `WH_CACHE_REDIS_TIMEOUT` | `100ms` | Per operation. A lookup or fill that takes longer is a miss or a skipped fill, never a failed query. | -| `cache.redis.dial_timeout` | `WH_CACHE_REDIS_DIAL_TIMEOUT` | `1s` | Per connection attempt, at most `2s`. An attempt is a connect and a handshake, each bounded by this, and both boot and shutdown wait out one in flight, so the cap keeps a stop inside the release's fixed 5s budget (see [Stopping](/deployment#stopping)). | +| `cache.redis.timeout` | `WH_CACHE_REDIS_TIMEOUT` | `100ms` | Per operation, at most `1s`. A lookup or fill that takes longer is a miss or a skipped fill, never a failed query. | +| `cache.redis.dial_timeout` | `WH_CACHE_REDIS_DIAL_TIMEOUT` | `1s` | Per connection attempt, at most `1s`. An attempt is a connect and a handshake, each bounded by this, and in `cluster` mode a topology read bounded by the larger of this and `timeout`. Boot and shutdown each wait out one in flight, so the two caps keep that under 3s, inside a stop's fixed 5s release budget (see [Stopping](/deployment#stopping)). | | `cache.redis.max_value_bytes` | `WH_CACHE_REDIS_MAX_VALUE_BYTES` | `1048576` | Largest result stored, after compression (1 MiB). A larger one is returned to the caller but not cached. | | `cache.redis.compress_min_bytes` | `WH_CACHE_REDIS_COMPRESS_MIN_BYTES` | `1024` | Results at least this large are zstd-compressed when that makes them smaller. `0` never compresses. | | `cache.redis.version_ttl` | `WH_CACHE_REDIS_VERSION_TTL` | `168h` | How long a table's or tenant's version token outlives its last write, so the tokens of dropped tables and removed tenants eventually expire. At least `2s`. An expired token only causes misses. | -**When the server is unreachable or misbehaves, the cache is bypassed; queries are not.** A failure or a timeout makes the lookup a miss and the fill a no-op, and five in a row open a circuit breaker that skips the server entirely until a probe, every 5 s, gets an answer. Queries then go straight to ClickHouse, still coalesced per instance by `singleflight`. An invalidation the server did not take is kept and retried until it lands. `/readyz` does not depend on the cache. +**When the server is unreachable or misbehaves, the cache is bypassed; queries are not.** A failure or a timeout makes the lookup a miss and the fill a no-op. Five in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), open a circuit breaker that skips the server entirely until a probe write, every 5 s, succeeds. Queries then go straight to ClickHouse, still coalesced per instance by `singleflight`. An invalidation the server did not take is kept and retried until it lands, and until then the instance that owes it bypasses the lookups it would orphan. `/readyz` does not depend on the cache. **Boot does not wait for the server.** A malformed block (an address without a port, `mode: cluster` with `db` other than `0`, an unreadable or unparsable TLS file) refuses boot. A server that cannot be reached, or that refuses the credentials, does not: the process boots with the cache bypassed and keeps reconnecting, with backoff up to 30 s. A rejected password (`WRONGPASS`, `NOAUTH`) is logged at `ERROR` on every attempt; any other failure at `WARN`. This is deliberate: a rotated Redis password must not crash-loop every instance at once. Watch `wavehouse_cache_breaker_open`, which reads `1` while the cache is bypassed, including before the first connection. diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index fc366405..1e14399b 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -432,12 +432,12 @@ The query-result cache is the layer that can be shared today. With the default ` **What another instance can see.** Ingest is already asynchronous: `/v1/ingest` answers before the batch is inserted. Once the inserting instance's worker has written the batch to ClickHouse, it replaces the table's version token in Redis, and from then on a lookup on any instance misses and reads the new rows. The cache adds no delay of its own beyond that single write. The exceptions: - **The server is unreachable from the inserting instance.** The invalidation is kept and retried until it lands (`wavehouse_cache_invalidations_pending` counts what is owed). Meanwhile other instances that can still reach the server keep serving the older results, for as long as the outage lasts and at most until each entry's TTL. An instance that stops while invalidations are still owed loses them, with the same bound. The same thing happens today when a process stops between an insert and its invalidation. -- **A failover to a replica that had not yet received the latest token writes** can bring back entries filed under the older tokens, bounded by the replication lag at the moment of failover and those entries' TTL. WaveHouse never reads from replicas. -- **The server is full and `maxmemory-policy` is `noeviction`.** It refuses the token writes, so invalidations are kept and retried, and until one lands every instance serves the results from before the insert, up to their TTL. +- **A failover to a replica that had not yet received the latest token writes** can bring back entries filed under the older tokens, bounded by the replication lag at the moment of failover and those entries' TTL. WaveHouse never reads from replicas. Behind a stable address (a managed primary endpoint), an instance still connected to the demoted node has its writes refused, which bypasses its cache; connections are replaced every minute, so it reaches the new primary and delivers the invalidations it owes within about that long. +- **The server is full and `maxmemory-policy` is `noeviction`.** It refuses the token writes. The inserting instance keeps its invalidations and retries them, bypassing its cache meanwhile, but every other instance serves the results from before the insert until one lands, up to their TTL. - **A pipe that writes** (an `INSERT` in `pipes.json`) has its result cached like a read, so a repeated identical call is answered from the cache and the write does not run again ([#386](https://github.com/Wave-RF/WaveHouse/issues/386)). With a shared cache that holds on every instance, until the entry's TTL. - **Admin writes through `POST /v1/ops/query`** do not invalidate the cache ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). With a shared cache, the stale results they leave are served by every instance, not only one. -**Sizing the server.** Every key WaveHouse writes has a TTL, and a version token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses, never bring back an entry it had invalidated. So set `maxmemory` and let the server evict: `maxmemory-policy allkeys-lru` (or `allkeys-lfu`, `volatile-lru`, `volatile-lfu`). Under `noeviction`, a full server refuses the writes. Fills then fail (counted by `wavehouse_cache_set_failures_total{reason="oom"}`), only lookups whose version tokens already exist keep working, and invalidations are kept and retried, so the pre-insert results above stay served: avoid `noeviction`. A stored result is capped at `cache.redis.max_value_bytes` (1 MiB compressed). A tenant's version tokens share one hash tag, so each lookup reads them in one `MGET` in cluster mode as well. The results themselves carry no hash tag and spread across shards. Persistence is not needed: an empty server after a restart is a cold cache, not a wrong one. +**Sizing the server.** Every key WaveHouse writes has a TTL, and a version token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses, never bring back an entry it had invalidated. So set `maxmemory` and let the server evict: `maxmemory-policy allkeys-lru` (or `allkeys-lfu`, `volatile-lru`, `volatile-lfu`). Under `noeviction`, a full server refuses the writes: each refusal (a fill's is counted by `wavehouse_cache_set_failures_total{reason="oom"}`) bypasses the cache of the instance that got it, and invalidations are kept and retried, so the pre-insert results above stay served by the others: avoid `noeviction`. A stored result is capped at `cache.redis.max_value_bytes` (1 MiB compressed). A tenant's version tokens share one hash tag, so each lookup reads them in one `MGET` in cluster mode as well. The results themselves carry no hash tag and spread across shards. Persistence is not needed: an empty server after a restart is a cold cache, not a wrong one. **The server is inside the trust boundary.** A cached result is served after the access policy has filtered it, so whoever can write to the server can change what any caller reads. Keep it on a private network, require a password or ACL user (`WH_CACHE_REDIS_PASSWORD`), use TLS across links you do not trust, and share it only with deployments you trust as much as this one. diff --git a/internal/config/cache_redis.go b/internal/config/cache_redis.go index 92063015..df817435 100644 --- a/internal/config/cache_redis.go +++ b/internal/config/cache_redis.go @@ -19,11 +19,12 @@ const ( RedisSentinel = "sentinel" ) -// maxRedisDialTimeout caps cache.redis.dial_timeout. A dial is a connect -// and a handshake, each bounded by it, and both boot and the cache's Close -// wait out one in flight: at 2s that is at most 4s, inside the 5s budget -// Close shares with the stores released after the cache. -const maxRedisDialTimeout = 2 * time.Second +// maxRedisTimeout caps cache.redis.timeout and cache.redis.dial_timeout. +// Boot and the cache's Close each wait out a dial in flight: a connect and +// a handshake, each bounded by dial_timeout, then for a cluster a topology +// read bounded by the larger of the two. At the caps that is at most 3s, +// inside the 5s budget Close shares with the stores released after it. +const maxRedisTimeout = time.Second // CacheRedisConfig configures cache.backend=redis: one Redis-compatible server // (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB) shared by every process. @@ -108,8 +109,16 @@ func (r CacheRedisConfig) validate() error { return fmt.Errorf("%s %s must be positive", d.key, d.v) } } - if r.DialTimeout > maxRedisDialTimeout { - return fmt.Errorf("cache.redis.dial_timeout (WH_CACHE_REDIS_DIAL_TIMEOUT) %s is over %s: boot and shutdown each wait out a dial, up to twice this", r.DialTimeout, maxRedisDialTimeout) + for _, d := range []struct { + key string + v time.Duration + }{ + {"cache.redis.timeout (WH_CACHE_REDIS_TIMEOUT)", r.Timeout}, + {"cache.redis.dial_timeout (WH_CACHE_REDIS_DIAL_TIMEOUT)", r.DialTimeout}, + } { + if d.v > maxRedisTimeout { + return fmt.Errorf("%s %s is over %s: boot and shutdown each wait out a connection attempt, which both bound", d.key, d.v, maxRedisTimeout) + } } if r.VersionTTL < 2*time.Second { return fmt.Errorf("cache.redis.version_ttl (WH_CACHE_REDIS_VERSION_TTL) %s is under 2s", r.VersionTTL) diff --git a/internal/config/cache_redis_test.go b/internal/config/cache_redis_test.go index 7f91a049..264ed31d 100644 --- a/internal/config/cache_redis_test.go +++ b/internal/config/cache_redis_test.go @@ -57,7 +57,7 @@ func TestLoad_CacheRedisFromEnv(t *testing.T) { "WH_CACHE_REDIS_TLS_SERVER_NAME": "redis.internal", "WH_CACHE_REDIS_KEY_PREFIX": "staging", "WH_CACHE_REDIS_TIMEOUT": "250ms", - "WH_CACHE_REDIS_DIAL_TIMEOUT": "2s", + "WH_CACHE_REDIS_DIAL_TIMEOUT": "500ms", "WH_CACHE_REDIS_MAX_VALUE_BYTES": "2048", "WH_CACHE_REDIS_COMPRESS_MIN_BYTES": "0", "WH_CACHE_REDIS_VERSION_TTL": "24h", @@ -73,7 +73,7 @@ func TestLoad_CacheRedisFromEnv(t *testing.T) { TLS: CacheRedisTLS{ Enabled: true, CAFile: caFile, CertFile: certFile, KeyFile: keyFile, ServerName: "redis.internal", }, - KeyPrefix: "staging", Timeout: 250 * time.Millisecond, DialTimeout: 2 * time.Second, + KeyPrefix: "staging", Timeout: 250 * time.Millisecond, DialTimeout: 500 * time.Millisecond, MaxValueBytes: 2048, CompressMinBytes: 0, VersionTTL: 24 * time.Hour, }, cfg.Cache.Redis) tc, err := cfg.Cache.Redis.TLS.Config() @@ -221,8 +221,9 @@ func TestValidate_CacheRedis(t *testing.T) { {"brace prefix", func(r *CacheRedisConfig) { r.KeyPrefix = "a{b}" }, "hash-tag brace"}, {"zero timeout", func(r *CacheRedisConfig) { r.Timeout = 0 }, "cache.redis.timeout (WH_CACHE_REDIS_TIMEOUT) 0s must be positive"}, {"negative dial timeout", func(r *CacheRedisConfig) { r.DialTimeout = -time.Second }, "cache.redis.dial_timeout"}, - {"dial timeout at the cap", func(r *CacheRedisConfig) { r.DialTimeout = 2 * time.Second }, ""}, - {"dial timeout over the cap", func(r *CacheRedisConfig) { r.DialTimeout = 2*time.Second + time.Millisecond }, "cache.redis.dial_timeout (WH_CACHE_REDIS_DIAL_TIMEOUT) 2.001s is over 2s"}, + {"timeouts at the cap", func(r *CacheRedisConfig) { r.Timeout, r.DialTimeout = time.Second, time.Second }, ""}, + {"timeout over the cap", func(r *CacheRedisConfig) { r.Timeout = time.Second + time.Millisecond }, "cache.redis.timeout (WH_CACHE_REDIS_TIMEOUT) 1.001s is over 1s"}, + {"dial timeout over the cap", func(r *CacheRedisConfig) { r.DialTimeout = time.Second + time.Millisecond }, "cache.redis.dial_timeout (WH_CACHE_REDIS_DIAL_TIMEOUT) 1.001s is over 1s"}, {"short version ttl", func(r *CacheRedisConfig) { r.VersionTTL = time.Second }, "cache.redis.version_ttl (WH_CACHE_REDIS_VERSION_TTL) 1s is under 2s"}, {"zero max value", func(r *CacheRedisConfig) { r.MaxValueBytes = 0 }, "cache.redis.max_value_bytes"}, {"compress never", func(r *CacheRedisConfig) { r.CompressMinBytes = 0 }, ""}, diff --git a/tests/integration/shared_cache_test.go b/tests/integration/shared_cache_test.go index e9857b72..f120365c 100644 --- a/tests/integration/shared_cache_test.go +++ b/tests/integration/shared_cache_test.go @@ -80,7 +80,7 @@ func bootRedisApp(t *testing.T, redisAddr, prefix string, timeout time.Duration) MQ: config.MQ{Backend: config.MQEmbedded}, Cache: config.Cache{Backend: config.CacheRedis, Redis: config.CacheRedisConfig{ Addrs: []string{redisAddr}, Mode: config.RedisStandalone, KeyPrefix: prefix, - Timeout: timeout, DialTimeout: 2 * time.Second, + Timeout: timeout, DialTimeout: time.Second, MaxValueBytes: 1 << 20, CompressMinBytes: 1 << 10, VersionTTL: time.Hour, }}, Dedupe: config.Dedupe{Backend: config.DedupePebble}, @@ -149,8 +149,8 @@ func TestSharedCache_IngestOnOneInstanceInvalidatesAnother(t *testing.T) { table := createTable(t, "user_id String, value Float64", "ORDER BY user_id") _, redisAddr := startRedis(t) prefix := fmt.Sprintf("it%d", cachePrefixes.Add(1)) - a := bootRedisApp(t, redisAddr, prefix, 2*time.Second) - b := bootRedisApp(t, redisAddr, prefix, 2*time.Second) + a := bootRedisApp(t, redisAddr, prefix, time.Second) + b := bootRedisApp(t, redisAddr, prefix, time.Second) rc, err := rueidis.NewClient(rueidis.ClientOption{InitAddress: []string{redisAddr}, DisableCache: true, ForceSingleClient: true}) require.NoError(t, err) From 41ebc664b561d06123093141b738a9b66ff96318 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 06:36:31 -0400 Subject: [PATCH 19/27] =?UTF-8?q?fix(cache):=20review=20round=201=20?= =?UTF-8?q?=E2=80=94=20one=20standalone=20address,=20e2e=20stack=20docs,?= =?UTF-8?q?=20e2e=20gate?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - Standalone mode dials only the first address, so a second one was silently ignored: it now refuses boot. A port must be 1-65535. - The manual e2e boot command needs a Redis now that the fixture sets cache.backend: redis; the e2e stack is described as ClickHouse and Redis wherever it was ClickHouse only. - The e2e gate (59.9%) excludes cache_redis.go's rejection paths and pending.go's outage-only retries, as it excludes internal/settings; unit and integration cover both, and the merged total still counts them. - The integration tests read the TTL floor from cache.QueryTimeToTTL. - A failed InvalidateTenant on reconcile logs at WARN, like the worker. - The overview pages mention the shared cache. Co-Authored-By: Claude Opus 5.5 (1M context) --- .github/workflows/README.md | 2 +- .github/workflows/ci.yml | 4 ++-- .testcoverage.yml | 6 ++++++ AGENTS.md | 6 +++--- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/configuration.mdx | 2 +- docs/src/content/docs/development.md | 11 ++++++----- docs/src/content/docs/index.mdx | 2 +- docs/src/content/docs/sdk/reference.md | 4 ++-- docs/src/content/docs/why-wavehouse.md | 2 +- internal/app/app_test.go | 6 ++++-- internal/app/wire.go | 2 +- internal/config/cache_redis.go | 20 +++++++++++--------- internal/config/cache_redis_test.go | 9 +++++++-- tests/integration/shared_cache_test.go | 3 ++- 16 files changed, 50 insertions(+), 33 deletions(-) diff --git a/.github/workflows/README.md b/.github/workflows/README.md index 575a8389..470f82f8 100644 --- a/.github/workflows/README.md +++ b/.github/workflows/README.md @@ -44,7 +44,7 @@ Break one of these knowingly or not at all. 2. **A dedicated `coverage` job applies the consolidated gate, polling — not `needs`-ing — the suites.** Each suite (`unit`, `integration`, `e2e`) runs with `COV_DEFER=1` and uploads a `coverage-` fragment; the `coverage` job runs `make cov` (merge + every threshold gate) over all three — exactly like local `make ci`'s final step. Keeping it a separate job (not folded into e2e's tail) decouples the gate result from the e2e suite's pass/fail. Crucially it is `needs: changes` **only, not the suites**: a `needs` edge is a *scheduling* barrier — GitHub won't pick up a runner, check out, restore caches, or `pnpm install` until the needed jobs finish — so needing the suites would serialize this job's ~50s of setup onto the critical path after the last suite, for nothing (the setup doesn't depend on their results). Instead it starts at run creation, runs its setup in parallel with the suites, and blocks only at the merge by polling for the three fragments with [`scripts/ci/wait-artifact.sh`](../../scripts/ci/wait-artifact.sh) (fails fast if a producer concluded without producing). Tail on the critical path: ~10s, not ~50s. **The aggregator and `docs-deploy` must keep `coverage` *and* every suite in their `needs`** — the suites directly (a suite failure must red the gate even though `coverage` no longer needs them), and `coverage` (else a coverage-gate failure wouldn't block merge or a prod deploy). -3. **e2e builds its own inputs and mirrors local `make test-e2e`.** It compiles the SDK dist + cover binary itself (`make -j test-e2e`, warm per-suffix cache) rather than waiting on a builder job, and runs the suite exactly as a developer does — one orchestrator, one ClickHouse testcontainer, sequential files. The ClickHouse image pulls in the background while caches restore (also in the integration job). +3. **e2e builds its own inputs and mirrors local `make test-e2e`.** It compiles the SDK dist + cover binary itself (`make -j test-e2e`, warm per-suffix cache) rather than waiting on a builder job, and runs the suite exactly as a developer does — one orchestrator, one ClickHouse and one Redis testcontainer, sequential files. The ClickHouse image pulls in the background while caches restore (also in the integration job). 4. **One change classifier, split into a pure core + a CI wrapper.** The pure allowlist — file list on stdin ⇒ `code`/`docs` — lives in [`scripts/classify-paths.sh`](../../scripts/classify-paths.sh), dependency-free and unit-tested by [`scripts/classify-paths.test.sh`](../../scripts/classify-paths.test.sh) (`make test-classify-paths`, a `verify` leaf) so the allowlists can't silently regress. The `changes` job runs the thin wrapper [`scripts/ci/classify-changes.sh`](../../scripts/ci/classify-changes.sh), which adds the CI-only policy (API file-list fetch + fail-closed: pushes, dispatches, API hiccups ⇒ `code=true`) on top. Keeping the core pure means the local git hooks can share it (`git diff --name-only | scripts/classify-paths.sh`). The `code`/`docs` outputs gate the suites and docs jobs — gate on these, never on workflow-level `paths:` filters, which would orphan the required check (invariant 1). diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index f4761927..fa572e41 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -280,8 +280,8 @@ jobs: with: go-cache-suffix: "-e2e-cov" # `-j` builds the prereqs (build-ts ∥ build-cover) concurrently, - # then runs the orchestrator: ClickHouse testcontainer + the cover - # binary + the SDK vitest suite. + # then runs the orchestrator: ClickHouse and Redis testcontainers + + # the cover binary + the SDK vitest suite. - name: Build SDK dist + cover binary, run E2E suite run: make -j "$(nproc)" test-e2e COV_DEFER=1 - name: Upload coverage fragment diff --git a/.testcoverage.yml b/.testcoverage.yml index 84f028af..451014a3 100644 --- a/.testcoverage.yml +++ b/.testcoverage.yml @@ -84,3 +84,9 @@ exclude: # never reaches them. The unit suite and the integration suite's main # app (cache.backend=local) cover them; the merged total still counts them. - ^internal/cache/(local|version_manager)\.go$ + # What e2e's Redis never makes it run: the cache.redis block's rejection + # paths (boot adopts a valid fixture, as with internal/settings above) + # and the retry of invalidations the server did not take, which needs an + # outage. The unit and integration suites cover both. + - ^internal/config/cache_redis\.go$ + - ^internal/cache/pending\.go$ diff --git a/AGENTS.md b/AGENTS.md index 8feaab73..a04c4a8b 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -152,7 +152,7 @@ If `make ci` passes locally, your commit has crossed the same gates CI will run ### Running `make ci` (for agents) -`make ci` is **self-contained**: the integration suite (`tests/integration/`) and the E2E orchestrator (`scripts/orchestrator/`) each boot ClickHouse via **testcontainers on random host ports**, and the shared cache backend's integration tests (`internal/cache/`) start their own Redis, Valkey, Dragonfly and one-node Redis Cluster containers the same way. The only prerequisite is a running **Docker daemon** — do **not** `make deps-up` or start ClickHouse first (`deps-up` is for `make dev` only). +`make ci` is **self-contained**: the integration suite (`tests/integration/`) and the E2E orchestrator (`scripts/orchestrator/`) each boot ClickHouse and a Redis via **testcontainers on random host ports**, and the shared cache backend's integration tests (`internal/cache/`) start their own Redis, Valkey, Dragonfly and one-node Redis Cluster containers the same way. The only prerequisite is a running **Docker daemon** — do **not** `make deps-up` or start ClickHouse first (`deps-up` is for `make dev` only). Run it via the **background Bash tool** (`run_in_background: true`) and wait for the completion notification; the harness re-invokes you on exit, so polling the log with `tail` only burns context: @@ -449,8 +449,8 @@ internal/stream/ → SSE fan-out (event Hub: project once per role, Subsc internal/tenant/ → Tenant id (type, grammar, reserved default, request header name) internal/testutil/ → Shared test helpers (mocks, JWT + schema helpers; logtest/ captures or silences the default logger; cachetest/ is the conformance suite every cache.Cache backend runs) tests/ → Integration & E2E tests -tests/integration/ → Go integration tests (//go:build integration; ClickHouse testcontainer). A package tested against its own external server keeps them beside it: internal/cache/redis_integration_test.go (Redis, Valkey, Dragonfly, Redis Cluster testcontainers) -tests/e2e/ → E2E test stack (scripts/orchestrator boots a ClickHouse testcontainer + the wavehouse-cov binary) +tests/integration/ → Go integration tests (//go:build integration; ClickHouse testcontainer, and Redis for shared_cache_test.go). A package tested against its own external server keeps them beside it: internal/cache/redis_integration_test.go (Redis, Valkey, Dragonfly, Redis Cluster testcontainers) +tests/e2e/ → E2E test stack (scripts/orchestrator boots ClickHouse and Redis testcontainers + the wavehouse-cov binary) tests/e2e/fixtures/ → Idempotent ClickHouse DDL scripts for test tables tests/e2e/sdk/ → E2E integration tests via TypeScript SDK (Vitest) deployments/compose/ → Docker Compose files (standalone.yaml, dependencies.yaml) diff --git a/CHANGELOG.md b/CHANGELOG.md index 51ed8865..c4e6e293 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `internal/ingest/worker.go`, `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`) and `dial_timeout` (`1s`), each at most `1s` since boot and shutdown each wait out a connection attempt they bound, `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The ingest worker's log of an invalidation that did not land drops from `ERROR` to `WARN`, since the shared backend defers and retries it: an outage would otherwise log an `ERROR` for every batch. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits; a third runs the first one's hit, insert and fresh-miss lifecycle on the suite's own `cache.backend: local` app, since e2e no longer exercises that backend. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. +- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `internal/ingest/worker.go`, `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`) and `dial_timeout` (`1s`), each at most `1s` since boot and shutdown each wait out a connection attempt they bound, `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a valid port, more than one address in `standalone` mode, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The ingest worker's log of an invalidation that did not land drops from `ERROR` to `WARN`, since the shared backend defers and retries it: an outage would otherwise log an `ERROR` for every batch. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits; a third runs the first one's hit, insert and fresh-miss lifecycle on the suite's own `cache.backend: local` app, since e2e no longer exercises that backend. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. - **A Redis-compatible shared cache backend** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/cache.go`, `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`, `CONTRIBUTING.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), open a circuit breaker that skips the server until a probe write succeeds (connections are replaced every minute, so after a failover behind a stable address the process reaches the new primary, and delivers the bumps it owes, within about a minute, provided a fresh connection's dial and handshake fit in the per-operation timeout), and deferred invalidations are retried until they land — the first at once, and at the probe's cadence while the breaker is open — collapsing to one tenant-wide bump per tenant past 100,000 keys; until one lands, the process that owes it bypasses the lookups it would orphan. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. `cache.backend: redis` selects it (the entry above). Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression, server-stops-answering, refused-writes (a demoted primary, a full `noeviction` server), failover-behind-a-stable-address and owed-invalidation cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index d3e47962..7511a2a9 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -126,7 +126,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — the implementation each layer runs on, the process's `roles`, resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `cache.redis.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). - **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load` when `Config.NeedsDataDir` says a selected backend keeps state there, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. - **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` in `Validate`, between `validateRoles` and `validateTopology`), which refuses a value not on the list and names the ones that are. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns what a valid configuration is still likely to get wrong — the combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), a `cache.redis` block that is not read, certificate verification turned off — which `app.New` logs at `WARN`. -- **cache_redis.go** — `CacheRedisConfig`, the `cache.redis` sub-block, and its checks: an address, each `host:port` (a URL or `user:password@` form refused without repeating it, since it may hold a password), a known mode (`sentinel` is refused until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `db` 0 in cluster mode, positive timeouts and sizes, a `timeout` and `dial_timeout` of at most 1 s each (boot and `Close` each wait out a dial: a connect and a handshake bounded by `dial_timeout`, and a cluster's topology read bounded by the larger of the two), a `version_ttl` of at least 2 s, and a `compress_min_bytes` that is not negative (`0` never compresses); its defaults are in `defaults()` with the rest. `CacheRedisTLS.Config` builds the `tls.Config`, reading the files; `Validate` calls it so an unreadable file refuses boot, and `internal/app` calls it again to build the connection. A TLS key set while `tls.enabled` is off is an error rather than a plaintext connection. +- **cache_redis.go** — `CacheRedisConfig`, the `cache.redis` sub-block, and its checks: an address (exactly one in `standalone` mode, which dials only the first), each `host:port` with a port from 1 to 65535 (a URL or `user:password@` form refused without repeating it, since it may hold a password), a known mode (`sentinel` is refused until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `db` 0 in cluster mode, positive timeouts and sizes, a `timeout` and `dial_timeout` of at most 1 s each (boot and `Close` each wait out a dial: a connect and a handshake bounded by `dial_timeout`, and a cluster's topology read bounded by the larger of the two), a `version_ttl` of at least 2 s, and a `compress_min_bytes` that is not negative (`0` never compresses); its defaults are in `defaults()` with the rest. `CacheRedisTLS.Config` builds the `tls.Config`, reading the files; `Validate` calls it so an unreadable file refuses boot, and `internal/app` calls it again to build the connection. A TLS key set while `tls.enabled` is off is an error rather than a plaintext connection. - **config.go**, roles — `roles` (`[]Role`: `api`, `ingest`, `sweeper`; `AllRoles` by default; `Has(Role)`) picks which components `internal/app` wires, and `instance_id` names the process (`-<8 hex>` when empty, resolved in `Load`; today only logged at boot, and a distributed coordinator will record it as a lease's holder). `validateRoles` refuses an empty list, an empty entry, an unknown or a repeated role; `validateTopology` refuses a role set the backends cannot serve: any split over the embedded MQ, and a process with exactly one of `api` and `ingest` over a local cache. `NeedsDataDir` counts Pebble only for a process running `api`, and `Warnings` is empty without `api`, since only that role opens a cache it reads or a dedupe store. - **strict.go** — `rejectUnknownKeys`, the YAML half: re-reads the file as a generic tree and walks it against the struct's `yaml` tags, listing every key the struct doesn't declare. cleanenv itself is lenient by design, which is exactly wrong for boot config once keys have moved to the settings directory. - **persistence.go** — `WarnIfFreshDataDir` logs the startup `WARN` when `data_dir` is missing or empty (on a redeploy, the sign that the volume didn't persist); `LogStorageInitError` attaches the UID-65532 `permissionHint` to a NATS or Pebble open failure that looks like a permission denial — the same hint string `CheckDataDir` uses. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 509088ae..d625b2e4 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -150,7 +150,7 @@ The `redis` backend's settings, read only when `cache.backend` is `redis`. It is | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | -| `cache.redis.addrs` | `WH_CACHE_REDIS_ADDRS` | *(empty)* | **Required** with `backend: redis`. `host:port` of the server; several are a cluster's seed nodes. Comma-separated in the env var. Not a URL: `redis://user:password@host:port` refuses boot, without repeating the value, so set the credentials through `username` and `password`, and `rediss://` through `tls.enabled`. | +| `cache.redis.addrs` | `WH_CACHE_REDIS_ADDRS` | *(empty)* | **Required** with `backend: redis`. `host:port` of the server: exactly one in `standalone` mode, where a second would be ignored and so refuses boot; several are a cluster's seed nodes. Comma-separated in the env var. Not a URL: `redis://user:password@host:port` refuses boot, without repeating the value, so set the credentials through `username` and `password`, and `rediss://` through `tls.enabled`. | | `cache.redis.mode` | `WH_CACHE_REDIS_MODE` | `standalone` | `standalone` or `cluster`. `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656): the cache does not yet authenticate to the sentinels or refresh their topology, so it could not be trusted to follow a failover. | | `cache.redis.username` | `WH_CACHE_REDIS_USERNAME` | *(empty)* | ACL user. Empty uses the server's `default` user. | | `cache.redis.password` | `WH_CACHE_REDIS_PASSWORD` | *(empty)* | A secret: set it through the environment (or your secret store's env injection), not in a tracked `config.yaml`. | diff --git a/docs/src/content/docs/development.md b/docs/src/content/docs/development.md index 7b3ec0e6..92785bda 100644 --- a/docs/src/content/docs/development.md +++ b/docs/src/content/docs/development.md @@ -16,7 +16,7 @@ You need these on your `PATH` before any `make` recipe will work end-to-end: | **Go** | 1.26+ (matches `go.mod`) | Compiles `cmd/wavehouse`; also runs the pinned `tool` deps (`gotestsum`, `gofumpt`, `goimports`, `govulncheck`, `deadcode`, `gsa`, `goda`) via `go tool` | [go.dev/dl](https://go.dev/dl/) | | **GNU Make** | **4.0+** | The Makefile uses `--output-sync=target` (Make 4 only) and bash-pinned recipes. macOS ships with BSD Make 3.81, which **will not work** | macOS: `brew install make` then use `gmake` or put `$(brew --prefix make)/libexec/gnubin` on your PATH. Linux: usually already installed | | **bash** | 4+ recommended | Recipes are pinned to `bash`; the helper scripts under `scripts/` use `set -euo pipefail` and bash arrays | macOS default is bash 3.2 (works for current recipes, but `brew install bash` is safer); Linux distros ship 4+ | -| **Docker** *(or Podman)* | Engine 20.10+ with the Compose **v2** plugin (`docker compose`, no hyphen) | Compose stacks under `deployments/compose/`; the E2E and integration suites boot ClickHouse via testcontainers (no compose file), and the integration suite also starts Redis, Valkey, Dragonfly (pulled from `docker.dragonflydb.io`) and a one-node Redis Cluster for the shared cache backend | [Docker Desktop](https://docs.docker.com/get-docker/), [colima](https://github.com/abiosoft/colima), or [Podman](https://podman.io) with `podman-compose` / the `podman compose` plugin. The testcontainers Go library also honors `DOCKER_HOST` for rootless Podman setups | +| **Docker** *(or Podman)* | Engine 20.10+ with the Compose **v2** plugin (`docker compose`, no hyphen) | Compose stacks under `deployments/compose/`; the E2E and integration suites boot ClickHouse and a Redis via testcontainers (no compose file), and the integration suite also starts Valkey, Valkey, Dragonfly (pulled from `docker.dragonflydb.io`) and a one-node Redis Cluster for the shared cache backend | [Docker Desktop](https://docs.docker.com/get-docker/), [colima](https://github.com/abiosoft/colima), or [Podman](https://podman.io) with `podman-compose` / the `podman compose` plugin. The testcontainers Go library also honors `DOCKER_HOST` for rootless Podman setups | | **Node.js** | 22 LTS — pinned via `.nvmrc` at the repo root | Runtime for pnpm and the Vitest suites. Pinned to match CI (`setup-node` uses 22) and to avoid Node-major surprises; older Vitest versions in this repo were known to crash on Node 26 with a V8 heap-allocation abort | [nodejs.org](https://nodejs.org/) or `nvm use` / `fnm use` / `volta` (all read `.nvmrc`) | | **pnpm** | 11.21+ (pinned via `packageManager` in the root `package.json`) | Package manager for the TypeScript SDK, E2E test harness, and docs site (managed as a single pnpm workspace from the repo root); `make build-ts`, `make test-ts`, `make test-e2e`, `make build-docs`, `make dev-docs`, `make preview-docs` all shell out to `pnpm` | `corepack enable && corepack prepare pnpm@11.21.0 --activate` (recommended), or `npm i -g pnpm` | | **git** + **curl** | any recent | `git` for source + version metadata in builds; `curl` is used by the Makefile to fetch the pinned `golangci-lint` binary into `.bin/` | usually preinstalled | @@ -362,7 +362,7 @@ The primary E2E integration test suite lives in `tests/e2e/sdk/`. It uses the Ty **Architecture**: -- `scripts/orchestrator` — the E2E entrypoint behind `make test-e2e`: it starts a clean ClickHouse **testcontainer** per run, launches the `wavehouse-cov` binary on a random free port, runs the SDK suite against it, then SIGINTs the binary to flush coverage. No Compose file is involved. CI runs the exact same path. +- `scripts/orchestrator` — the E2E entrypoint behind `make test-e2e`: it starts a clean ClickHouse **testcontainer** and a Redis one (the fixture's shared cache, `cache.backend: redis`) per run, launches the `wavehouse-cov` binary on a random free port, runs the SDK suite against it, then SIGINTs the binary to flush coverage. No Compose file is involved. CI runs the exact same path. - `tests/e2e/sdk/setup.ts` — `globalSetup`. Probes the `CLICKHOUSE_URL` / `WAVEHOUSE_URL` the orchestrator injects, creates the per-suite tables, refreshes the schema, and writes the baseline policy into the run's settings directory (adopted via `POST /v1/ops/settings/reload` — files are the only write path). It starts nothing itself and fails fast if either URL isn't up. It also prints the active Node/undici version, warning when the local Node major differs from `.nvmrc` — a runtime-specific transport bug is otherwise indistinguishable from a code failure (see [#440](https://github.com/Wave-RF/WaveHouse/issues/440)). - `tests/e2e/sdk/helpers.ts` — JWT factories, typed client constructors, async wait helpers, direct ClickHouse query helper. @@ -375,10 +375,11 @@ make test-e2e `make test-e2e` builds `bin/wavehouse-cov` (coverage-instrumented) and runs the orchestrator under `scripts/orchestrator/` to wire ClickHouse + the cover binary into the suite. covdata flushes on SIGINT into `tmp/coverage/e2e/data/`. -The orchestrator always provisions its own stack — a fresh ClickHouse testcontainer plus `wavehouse-cov` on a random free port — so a running `make dev` on `:8080` is neither detected nor reused, and the two don't collide. To run vitest against a stack you manage yourself, start the server from the **repo root** with the E2E fixture config: +The orchestrator always provisions its own stack — fresh ClickHouse and Redis testcontainers plus `wavehouse-cov` on a random free port — so a running `make dev` on `:8080` is neither detected nor reused, and the two don't collide. To run vitest against a stack you manage yourself, start a Redis for the fixture's `cache.backend: redis`, then the server from the **repo root** with the E2E fixture config: ```bash -WH_CONFIG=tests/e2e/fixtures/config.yaml go run ./cmd/wavehouse +docker compose -f deployments/compose/dependencies.yaml --profile redis up -d +WH_CONFIG=tests/e2e/fixtures/config.yaml WH_CACHE_REDIS_ADDRS=localhost:6379 go run ./cmd/wavehouse ``` The fixture matters: the suite signs its tokens with its `sdk-dev-secret` and depends on its dedupe, DLQ, and 5s schema-refresh settings. Point the suite at a default `make dev` server (`jwt_secret: change-me-in-production`) and setup's schema calls are rejected, then global setup dies 30s later on a misleading `schema not refreshed within 30s`. The repo root matters too — the fixture's `settings.dir` is relative to the working directory. The fixture's settings directory (policy, pipes, and tunables) points at ClickHouse on `localhost:9000`; if yours isn't there, edit `clickhouse.addr` / `http_port` in `tests/e2e/fixtures/settings/config.json` (the orchestrator patches them itself for its testcontainer). @@ -474,7 +475,7 @@ WaveHouse/ │ └── testutil/ # Shared test helpers and mocks ├── tests/ # Integration & E2E tests │ ├── integration/ # Go integration tests (//go:build integration) -│ └── e2e/ # E2E suite (orchestrator + ClickHouse testcontainer) +│ └── e2e/ # E2E suite (orchestrator + ClickHouse and Redis testcontainers) │ ├── fixtures/ # ClickHouse DDL + config and settings-directory fixtures │ └── sdk/ # E2E specs driven through the TypeScript SDK (Vitest) ├── clients/ # Client SDKs diff --git a/docs/src/content/docs/index.mdx b/docs/src/content/docs/index.mdx index 3a6143ca..205a6f84 100644 --- a/docs/src/content/docs/index.mdx +++ b/docs/src/content/docs/index.mdx @@ -97,7 +97,7 @@ If you're building user-facing analytics, **WaveHouse is like Supabase for Click Every event is broadcast to SSE subscribers **before** it's flushed to ClickHouse. Gap-fill from JetStream history for late-connecting clients. - Ristretto cache plus Go `singleflight` coalesces identical concurrent queries — dashboards survive thundering herds without an extra cache tier to operate. + Ristretto cache plus Go `singleflight` coalesces identical concurrent queries — dashboards survive thundering herds without an extra cache tier to operate. Running several instances? Share one Redis-compatible cache with `cache.backend: redis`. Per-table, per-role column and row-level policies with JWT claim templating, defined in the hot-reloadable settings directory. diff --git a/docs/src/content/docs/sdk/reference.md b/docs/src/content/docs/sdk/reference.md index fa6983da..91df43af 100644 --- a/docs/src/content/docs/sdk/reference.md +++ b/docs/src/content/docs/sdk/reference.md @@ -184,8 +184,8 @@ export interface ClicksRow { The SDK doubles as the E2E integration test harness. Tests in `tests/e2e/sdk/` exercise the full pipeline (ingest → ClickHouse → query) through the SDK, validating both the backend and the client library in one pass. ```bash -# Run all E2E tests: the orchestrator boots a ClickHouse testcontainer + -# the wavehouse-cov binary, then runs the SDK suite +# Run all E2E tests: the orchestrator boots ClickHouse and Redis +# testcontainers + the wavehouse-cov binary, then runs the SDK suite make test-e2e ``` diff --git a/docs/src/content/docs/why-wavehouse.md b/docs/src/content/docs/why-wavehouse.md index 766d6214..883d1928 100644 --- a/docs/src/content/docs/why-wavehouse.md +++ b/docs/src/content/docs/why-wavehouse.md @@ -152,7 +152,7 @@ flowchart TB | ---------- | --------- | --------- | | Durable ingest buffer | Kafka / Redpanda cluster (3+ brokers, Zookeeper/KRaft) | Embedded NATS JetStream | | Batch consumer | Custom Go/Rust/Java service you write and operate | Built in | -| Query cache | Redis + singleflight middleware you write | Built in (Ristretto + singleflight) | +| Query cache | Redis + singleflight middleware you write | Built in (Ristretto + singleflight; or one Redis shared by every instance, `cache.backend: redis`) | | Real-time push | WebSocket service + bridge from Kafka | Built in (`/v1/stream`) | | Schema validation | Custom code in ingest API | Built in (discovers `system.columns`) | | Row/column access control | Custom middleware or a dedicated service | Built in (Hasura-style, JWT-driven) | diff --git a/internal/app/app_test.go b/internal/app/app_test.go index 4b7d400b..7a23330e 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -806,14 +806,14 @@ func TestNew_RedisCacheRefusesAnUnreadableTLSFile(t *testing.T) { func TestRedisConfig_FromLoadedDefaults(t *testing.T) { t.Setenv("WH_SETTINGS_DIR", t.TempDir()) t.Setenv("WH_CACHE_BACKEND", "redis") - t.Setenv("WH_CACHE_REDIS_ADDRS", "a:6379,b:6379") + t.Setenv("WH_CACHE_REDIS_ADDRS", "a:6379") t.Setenv("WH_CACHE_REDIS_PASSWORD", "pw") loaded, err := config.Load(filepath.Join(t.TempDir(), "none.yaml")) require.NoError(t, err) got, err := redisConfig(loaded.Cache.Redis) require.NoError(t, err) assert.Equal(t, cache.RedisConfig{ - Addrs: []string{"a:6379", "b:6379"}, Mode: cache.RedisStandalone, Password: "pw", + Addrs: []string{"a:6379"}, Mode: cache.RedisStandalone, Password: "pw", KeyPrefix: cache.DefaultRedisKeyPrefix, Timeout: cache.DefaultRedisTimeout, DialTimeout: cache.DefaultRedisDialTimeout, MaxValueBytes: cache.DefaultRedisMaxValueBytes, CompressMinBytes: cache.DefaultRedisCompressMinBytes, VersionTTL: cache.DefaultRedisVersionTTL, @@ -821,12 +821,14 @@ func TestRedisConfig_FromLoadedDefaults(t *testing.T) { t.Setenv("WH_CACHE_REDIS_COMPRESS_MIN_BYTES", "0") t.Setenv("WH_CACHE_REDIS_MODE", "cluster") + t.Setenv("WH_CACHE_REDIS_ADDRS", "a:6379,b:6379") loaded, err = config.Load(filepath.Join(t.TempDir(), "none.yaml")) require.NoError(t, err) got, err = redisConfig(loaded.Cache.Redis) require.NoError(t, err) assert.Zero(t, got.CompressMinBytes, "the backend's never, not its default") assert.Equal(t, cache.RedisCluster, got.Mode) + assert.Equal(t, []string{"a:6379", "b:6379"}, got.Addrs, "a cluster's seeds") assert.Equal(t, cache.RedisSentinel, config.RedisSentinel) } diff --git a/internal/app/wire.go b/internal/app/wire.go index a7c82a97..fd1eb2e1 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -322,7 +322,7 @@ func (a *App) wireClickHouse() error { // cached is stale, so all of it is orphaned at once. for _, id := range stale { if err := a.cache.InvalidateTenant(a.stopCtx, id); err != nil { - slog.Error("cache invalidation of a stale tenant failed; it may serve stale rows until they expire", "tenant", id, "error", err) + slog.Warn("cache invalidation of a stale tenant did not land; it may serve stale rows until it does", "tenant", id, "error", err) } } }) diff --git a/internal/config/cache_redis.go b/internal/config/cache_redis.go index df817435..be065266 100644 --- a/internal/config/cache_redis.go +++ b/internal/config/cache_redis.go @@ -7,6 +7,7 @@ import ( "fmt" "net" "os" + "strconv" "strings" "time" ) @@ -78,9 +79,13 @@ func (r CacheRedisConfig) validate() error { if strings.TrimSpace(a) != a { return fmt.Errorf("cache.redis.addrs (WH_CACHE_REDIS_ADDRS) %q: no spaces around an address", a) } - if _, _, err := net.SplitHostPort(a); err != nil { + _, port, err := net.SplitHostPort(a) + if err != nil { return fmt.Errorf("cache.redis.addrs (WH_CACHE_REDIS_ADDRS) %q: want host:port: %w", a, err) } + if n, err := strconv.Atoi(port); err != nil || n < 1 || n > 65535 { + return fmt.Errorf("cache.redis.addrs (WH_CACHE_REDIS_ADDRS) %q: want host:port with a port from 1 to 65535", a) + } } switch r.Mode { case RedisStandalone, RedisCluster: @@ -89,6 +94,11 @@ func (r CacheRedisConfig) validate() error { default: return fmt.Errorf("cache.redis.mode (WH_CACHE_REDIS_MODE) %q: valid: %s, %s", r.Mode, RedisStandalone, RedisCluster) } + // Standalone dials the first address only, so a second one (a replica, + // say) would be silently ignored rather than failed over to. + if r.Mode == RedisStandalone && len(r.Addrs) > 1 { + return fmt.Errorf("cache.redis.addrs (WH_CACHE_REDIS_ADDRS) has %d addresses: mode standalone connects to one server; several are a cluster's seeds (mode cluster)", len(r.Addrs)) + } if r.DB < 0 { return fmt.Errorf("cache.redis.db (WH_CACHE_REDIS_DB) %d is negative", r.DB) } @@ -108,14 +118,6 @@ func (r CacheRedisConfig) validate() error { if d.v <= 0 { return fmt.Errorf("%s %s must be positive", d.key, d.v) } - } - for _, d := range []struct { - key string - v time.Duration - }{ - {"cache.redis.timeout (WH_CACHE_REDIS_TIMEOUT)", r.Timeout}, - {"cache.redis.dial_timeout (WH_CACHE_REDIS_DIAL_TIMEOUT)", r.DialTimeout}, - } { if d.v > maxRedisTimeout { return fmt.Errorf("%s %s is over %s: boot and shutdown each wait out a connection attempt, which both bound", d.key, d.v, maxRedisTimeout) } diff --git a/internal/config/cache_redis_test.go b/internal/config/cache_redis_test.go index 264ed31d..3c1f7c05 100644 --- a/internal/config/cache_redis_test.go +++ b/internal/config/cache_redis_test.go @@ -45,7 +45,7 @@ func TestLoad_CacheRedisFromEnv(t *testing.T) { caFile, certFile, keyFile := writeTestPKI(t, dir) for k, v := range map[string]string{ "WH_CACHE_BACKEND": "redis", - "WH_CACHE_REDIS_ADDRS": "r1:6379,r2:6379", + "WH_CACHE_REDIS_ADDRS": "r1:6379", "WH_CACHE_REDIS_MODE": "standalone", "WH_CACHE_REDIS_USERNAME": "wavehouse", "WH_CACHE_REDIS_PASSWORD": "s3cret", @@ -68,7 +68,7 @@ func TestLoad_CacheRedisFromEnv(t *testing.T) { cfg, err := Load("nonexistent.yaml") require.NoError(t, err) assert.Equal(t, CacheRedisConfig{ - Addrs: []string{"r1:6379", "r2:6379"}, Mode: RedisStandalone, + Addrs: []string{"r1:6379"}, Mode: RedisStandalone, Username: "wavehouse", Password: "s3cret", DB: 2, TLS: CacheRedisTLS{ Enabled: true, CAFile: caFile, CertFile: certFile, KeyFile: keyFile, ServerName: "redis.internal", @@ -211,6 +211,11 @@ func TestValidate_CacheRedis(t *testing.T) { {"no addrs", func(r *CacheRedisConfig) { r.Addrs = nil }, "cache.backend=redis needs cache.redis.addrs (WH_CACHE_REDIS_ADDRS)"}, {"addr with space", func(r *CacheRedisConfig) { r.Addrs = []string{"a:6379", " b:6379"} }, "no spaces around an address"}, {"addr without port", func(r *CacheRedisConfig) { r.Addrs = []string{"redis"} }, `cache.redis.addrs (WH_CACHE_REDIS_ADDRS) "redis": want host:port`}, + {"addr empty port", func(r *CacheRedisConfig) { r.Addrs = []string{"redis:"} }, `"redis:": want host:port with a port from 1 to 65535`}, + {"addr port not a number", func(r *CacheRedisConfig) { r.Addrs = []string{"redis:637x"} }, "a port from 1 to 65535"}, + {"addr port out of range", func(r *CacheRedisConfig) { r.Addrs = []string{"redis:99999"} }, "a port from 1 to 65535"}, + {"standalone with two addrs", func(r *CacheRedisConfig) { r.Addrs = []string{"a:6379", "b:6379"} }, "has 2 addresses: mode standalone connects to one server"}, + {"cluster with two seeds", func(r *CacheRedisConfig) { r.Mode, r.Addrs = RedisCluster, []string{"a:6379", "b:6379"} }, ""}, {"mode", func(r *CacheRedisConfig) { r.Mode = "replica" }, `cache.redis.mode (WH_CACHE_REDIS_MODE) "replica": valid: standalone, cluster`}, {"sentinel refused", func(r *CacheRedisConfig) { r.Mode = RedisSentinel }, `cache.redis.mode (WH_CACHE_REDIS_MODE) "sentinel" is not supported yet: the cache neither authenticates to the sentinels nor refreshes their topology (https://github.com/Wave-RF/WaveHouse/issues/656)`}, {"cluster", func(r *CacheRedisConfig) { r.Mode = RedisCluster }, ""}, diff --git a/tests/integration/shared_cache_test.go b/tests/integration/shared_cache_test.go index f120365c..6ace9a76 100644 --- a/tests/integration/shared_cache_test.go +++ b/tests/integration/shared_cache_test.go @@ -23,6 +23,7 @@ import ( "github.com/testcontainers/testcontainers-go/wait" "github.com/Wave-RF/WaveHouse/internal/app" + "github.com/Wave-RF/WaveHouse/internal/cache" "github.com/Wave-RF/WaveHouse/internal/config" ) @@ -31,7 +32,7 @@ const redisImage = "redis:8.10.2-alpine" // minCacheTTL is cache.QueryTimeToTTL's floor: a fill made less than this // long ago cannot have expired, so a miss inside it is an invalidation. -const minCacheTTL = 10 * time.Second +var minCacheTTL = cache.QueryTimeToTTL(0) var cachePrefixes atomic.Uint64 From 64f09e8968194757cab996583a31a046bb4d6418 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 06:42:03 -0400 Subject: [PATCH 20/27] docs(dev): list Redis among the cache backend's integration servers Co-Authored-By: Claude Opus 5.5 (1M context) --- docs/src/content/docs/development.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/src/content/docs/development.md b/docs/src/content/docs/development.md index 92785bda..581df72d 100644 --- a/docs/src/content/docs/development.md +++ b/docs/src/content/docs/development.md @@ -16,7 +16,7 @@ You need these on your `PATH` before any `make` recipe will work end-to-end: | **Go** | 1.26+ (matches `go.mod`) | Compiles `cmd/wavehouse`; also runs the pinned `tool` deps (`gotestsum`, `gofumpt`, `goimports`, `govulncheck`, `deadcode`, `gsa`, `goda`) via `go tool` | [go.dev/dl](https://go.dev/dl/) | | **GNU Make** | **4.0+** | The Makefile uses `--output-sync=target` (Make 4 only) and bash-pinned recipes. macOS ships with BSD Make 3.81, which **will not work** | macOS: `brew install make` then use `gmake` or put `$(brew --prefix make)/libexec/gnubin` on your PATH. Linux: usually already installed | | **bash** | 4+ recommended | Recipes are pinned to `bash`; the helper scripts under `scripts/` use `set -euo pipefail` and bash arrays | macOS default is bash 3.2 (works for current recipes, but `brew install bash` is safer); Linux distros ship 4+ | -| **Docker** *(or Podman)* | Engine 20.10+ with the Compose **v2** plugin (`docker compose`, no hyphen) | Compose stacks under `deployments/compose/`; the E2E and integration suites boot ClickHouse and a Redis via testcontainers (no compose file), and the integration suite also starts Valkey, Valkey, Dragonfly (pulled from `docker.dragonflydb.io`) and a one-node Redis Cluster for the shared cache backend | [Docker Desktop](https://docs.docker.com/get-docker/), [colima](https://github.com/abiosoft/colima), or [Podman](https://podman.io) with `podman-compose` / the `podman compose` plugin. The testcontainers Go library also honors `DOCKER_HOST` for rootless Podman setups | +| **Docker** *(or Podman)* | Engine 20.10+ with the Compose **v2** plugin (`docker compose`, no hyphen) | Compose stacks under `deployments/compose/`; the E2E and integration suites boot ClickHouse and a Redis via testcontainers (no compose file), and the integration suite also runs the shared cache backend against Redis, Valkey, Dragonfly (pulled from `docker.dragonflydb.io`) and a one-node Redis Cluster | [Docker Desktop](https://docs.docker.com/get-docker/), [colima](https://github.com/abiosoft/colima), or [Podman](https://podman.io) with `podman-compose` / the `podman compose` plugin. The testcontainers Go library also honors `DOCKER_HOST` for rootless Podman setups | | **Node.js** | 22 LTS — pinned via `.nvmrc` at the repo root | Runtime for pnpm and the Vitest suites. Pinned to match CI (`setup-node` uses 22) and to avoid Node-major surprises; older Vitest versions in this repo were known to crash on Node 26 with a V8 heap-allocation abort | [nodejs.org](https://nodejs.org/) or `nvm use` / `fnm use` / `volta` (all read `.nvmrc`) | | **pnpm** | 11.21+ (pinned via `packageManager` in the root `package.json`) | Package manager for the TypeScript SDK, E2E test harness, and docs site (managed as a single pnpm workspace from the repo root); `make build-ts`, `make test-ts`, `make test-e2e`, `make build-docs`, `make dev-docs`, `make preview-docs` all shell out to `pnpm` | `corepack enable && corepack prepare pnpm@11.21.0 --activate` (recommended), or `npm i -g pnpm` | | **git** + **curl** | any recent | `git` for source + version metadata in builds; `curl` is used by the Makefile to fetch the pinned `golangci-lint` binary into `.bin/` | usually preinstalled | From 558bb23835691a2ca43c4b3b7afb85ce3d755972 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:30:43 -0400 Subject: [PATCH 21/27] fix(cache): pin the no-index bump test and fix a "query key" mixup MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit TestVersionManager_BumpWithoutIndex asserted only on vm.size() after its final BumpTenant call, which unconditionally deletes the tenant's index regardless of what the earlier no-index bumps did — so it could not fail against a tableLocked that wrongly creates an index instead of returning nil. Assert after each bump instead, and add the pruned-tenant case a write racing Prune must not revive. Also reword two docs passages that overloaded or misstated a term: architecture.md used "query key" for both the caller's input and the rendered entry key in the same paragraph; AGENTS.md's cache bullet read as if the index maps were keyed by escaped names; CHANGELOG.md's #382 bullet said the version index builds its keys with internal/keyenc, contradicting its own closing sentence that the index's entries are unaffected by the escaping. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- AGENTS.md | 2 +- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- internal/cache/version_manager_test.go | 18 +++++++++++++++++- 4 files changed, 20 insertions(+), 4 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index a504ffd1..a74dc6bb 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -31,7 +31,7 @@ Twenty internal packages under `internal/` (plus `internal/testutil/` for shared - **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers; `ch_errors.go` (`writeCHError`) is the one mapping from a failed ClickHouse query to status, `code` and `retryable` - **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` handing the sweeper each tenant's own gap window (a rejected tenant's as its folder last had it, unbounded for one rejected since boot) and the `mq.max_bytes_gb` reconcile each served tenant's byte budget, and `defaultPolicy` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, the same hook's `Hub.Prune` ends the open streams of a tenant no longer served, and `wireCache`'s hook drops, through `LocalCache.Prune`, the cache version index of a tenant no longer served ([#262](https://github.com/Wave-RF/WaveHouse/issues/262))), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `New` wires only what the process's `roles` need (discovery, dedupe, auth verifiers, the hub bridge and keepalive per API process; the ingest worker per ingest process; the sweeper under its lease through `elected`); a process without `api` serves `api.NewOpsRouter` — probes, `/version`, metrics, and the settings reload behind the operator key alone. `cmd/wavehouse` and `tests/integration` both boot through it - **`auth/`** — JWT auth middleware: HMAC **or** JWKS verification with `alg` pinned to the active verifier, role extraction from a configurable claim path; always runs, never rejects (bad token → empty role + stashed reason). One verifier per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 9): `Authenticator` keys them by `tenant.ID` — the request store's `settings.Store.Tenant()`, through an injected `TenantSource`; `tenant.Default` on the tenant-exempt routes — built from each tenant's `auth` block by `Reconfigure`, dropped by `Prune` once the tenant stops being served (rejected or removed), released by `Close`; the secrets (`Config`) are boot-level and shared. A JWKS key set is fetched off the boot and reload paths: until one has been stored the verifier is pending and a token-bearing request gets `503` + `Retry-After` from `api.refuseUnverifiable` (`auth.ErrVerifierPending`), never a `default_role` evaluation; refresh is library-managed (Eric, 2026-09-22), response capped at 1 MiB; the operator key's admin role is the request tenant's -- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, a version per tenant, per (tenant, table) and per (tenant, table, scope), keyed by name and bumped in place (one entry per live namespace however often it is bumped, [#262](https://github.com/Wave-RF/WaveHouse/issues/262)), each raw table and scope name escaped by `keyenc` where the key is built (a `Namespace` carries them raw, so no caller escapes) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` drops the tenant's index so its next key gets a process-unique generation, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read, taken before the handler chooses any input a bump invalidates — the tenant's connection included — and `Set` files the fill under it, so a write landing mid-query, or a reload moving the tenant to another address or database after the request took its connection, orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) +- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for the caller's query key and its singleflight, escaped whole as the lead field of the stored key `|.||…`, where each raw table and scope name is escaped by `keyenc` (a `Namespace` carries them raw, so no caller escapes); the index holds a version per tenant, per (tenant, table) and per (tenant, table, scope), keyed by raw name and bumped in place (one entry per live namespace however often it is bumped, [#262](https://github.com/Wave-RF/WaveHouse/issues/262)) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` drops the tenant's index so its next key gets a process-unique generation, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read, taken before the handler chooses any input a bump invalidates — the tenant's connection included — and `Set` files the fill under it, so a write landing mid-query, or a reload moving the tenant to another address or database after the request took its connection, orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config. `Classify` (`errclass.go`) says what a failed ClickHouse request means for the request — `Unavailable`, `Denied`, `Rejected` (any unlisted exception code: the server read it and refused it), or `Unknown` (no code, no recognizable transport failure) — over the driver's error types and the HTTP interface's `HTTPError`; the ingest worker and the query handlers (`api/ch_errors.go` `writeCHError`, [#403](https://github.com/Wave-RF/WaveHouse/issues/403), [#271](https://github.com/Wave-RF/WaveHouse/issues/271)) both use it - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) - **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability when a selected backend keeps state there (`NeedsDataDir`); `backends.go` holds each layer's `.backend` (only the in-process value today); `config.go` holds `roles` (`Has(Role)`) and `instance_id`, and `Validate` refuses a role split the backends cannot serve (any split over the embedded MQ; `api` without `ingest`, or the reverse, over a local cache) — boot is the validator, there is no dry run diff --git a/CHANGELOG.md b/CHANGELOG.md index 6631b3c4..14de673f 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -88,7 +88,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **A failed ClickHouse query answers by what went wrong, not a flat `500`/`502`** (`internal/api/ch_errors.go` (new, + tests), `internal/api/{errors,query,structured_query,pipes,schema,ch_settings}.go`, `internal/chconn/errclass.go` (`HTTPStatus` exported), `clients/ts/src/errors.ts` (+ tests), `tests/integration/query_errors_test.go` (new), `tests/integration/query_limits_test.go`, `internal/app/app_test.go`, `tests/e2e/sdk/{admin,query}.test.ts`, `AGENTS.md`, `docs/src/content/docs/{api,architecture}.md`, `docs/src/content/docs/{access-control,configuration}.mdx`, `docs/src/content/docs/sdk/{reference.md,index.mdx}`): fixes [#403](https://github.com/Wave-RF/WaveHouse/issues/403) and [#271](https://github.com/Wave-RF/WaveHouse/issues/271), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). ClickHouse answers a syntax error, a missing grant and an overloaded server alike with HTTP `500`, so `/v1/ops/query` turned a bad statement into a `502` and `/v1/query` and pipes into a `500` the SDK retried. All three now class the failure with `chconn.Classify` through one helper, `writeCHError`: a statement ClickHouse refused is `400 clickhouse.rejected`; a query over a rows/bytes limit, the role's own memory cap, or its time cap where that is no longer than `query_timeout` is `400 clickhouse.limit_exceeded`; `ACCESS_DENIED` is `403 clickhouse.access_denied`; credentials, user or database refused, or a redirect or `4xx` with no exception code from whatever fronts ClickHouse, is `502 clickhouse.misconfigured`; ClickHouse down, unreachable or overloaded is `503 clickhouse.unavailable` with `Retry-After: 5`; a failure with no verdict stays `500` (`502` on the proxy) as `clickhouse.unknown`. The error envelope gains `code` and `retryable` next to `error` on these responses — additive. A role with `max_execution_time` now queries with no context deadline and a cancel two seconds past the cap instead: clickhouse-go overwrote the cap's `max_execution_time` with deadline+5s for any deadline over 1s, so an overrun came back as a bare deadline, indistinguishable from waiting for a pooled connection; ClickHouse now enforces the cap itself and reports `TIMEOUT_EXCEEDED`. `POST /v1/ops/schema/refresh` against an unreachable ClickHouse is a `503` with `Retry-After` instead of a `500`. **SDK:** `WaveHouseError.code` and `retryable` now take the server's `code`/`retryable` when the body has them (`HTTP_` and "5xx retries" otherwise), so a rejected query is `clickhouse.rejected` rather than `HTTP_500`, and is not retried. - **An unavailable ClickHouse is retried with backoff instead of dead-lettering every row** (`internal/chconn/errclass.go` (new, + tests), `internal/ingest/{worker,backoff}.go` (`backoff.go` new, + tests), `internal/mq/{mq,embedded}.go`, `internal/testutil/mocks.go`, `tests/integration/ingest_outage_test.go` (new), `AGENTS.md`, `README.md`, `docs/src/content/docs/{ingest-pipeline,architecture,api,deployment,why-wavehouse}.md`, `docs/src/content/docs/{settings-directory,index,access-control}.mdx`): workstream A of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). A failed batch insert used to go through row-by-row isolation whatever the failure, so a ClickHouse that was down, overloaded or read-only failed every row twice and parked the whole batch on the DLQ. `chconn.Classify` now classes the failure first — `Rejected` (any ClickHouse exception code outside the availability and credential lists: the server read the row and refused it), `Unavailable` (connection refused/reset, timeouts, `TOO_MANY_SIMULTANEOUS_QUERIES`, `SERVER_OVERLOADED`, `MEMORY_LIMIT_EXCEEDED`, `TOO_MANY_PARTS`, `READONLY`, `TABLE_IS_READ_ONLY`, `KEEPER_EXCEPTION`, …), `Denied` (`AUTHENTICATION_FAILED`, `ACCESS_DENIED`, …) or `Unknown` (no code, no recognizable transport failure). Only `Rejected` is isolated and dead-lettered as before, and a multi-row batch refused with `TOO_MANY_PARTS` or `MEMORY_LIMIT_EXCEEDED` is split row by row first (`chconn.Splittable`), because a batch spanning too many partitions or too much memory can fail where each of its rows inserts; every other class hands the batch back to the queue with a delayed nak (`mq.Message.NakWithDelay`, new) under a jittered 1 s → 30 s backoff shared by every table on the same ClickHouse pool (a failure of one table — read-only, too many parts or mutations, a grant missing on it, `chconn.TableScoped` — backs off that table alone), which turns rows away without a request while it runs and probes once per window, and ClickHouse going away mid-isolation stops isolation and retries the rows it had not settled. Counted by the new `wavehouse_ingest_retries_total{table, reason}`; logged at `WARN` when an outage starts and at most every 30 s during it. A long outage now shows as a growing ingest stream and, at `mq.max_bytes_gb`, ingest `503`s — not as a full DLQ; a lasting failure of one table holds back its tenant's other tables once its waiting rows reach `maxAckPending`. Retried rows come back out of arrival order, which matters only to a `ReplacingMergeTree` without a version column or a `CollapsingMergeTree`. - **Schema discovery's retry loop jitters its backoff** (`internal/discovery/discovery.go` (+ tests), `internal/app/wire.go`, `internal/api/errors.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`): `RetryRefresh` slept exactly `2s * 2^n` capped at 60s, so instances retrying against one recovering ClickHouse fired in lockstep, every 60s on the same second. Each sleep is now drawn uniformly from below the backoff (full jitter), spreading the retries over the whole window and halving the mean wait — so a failing tenant's retries, their log lines and `wavehouse_schema_refresh_failures_total` come about twice as often ([#141](https://github.com/Wave-RF/WaveHouse/issues/141)). -- **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/query/ident.go` (removed, + tests), `internal/app/wire.go`, `docs/src/content/docs/{api,architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL (a pipe result's key folded no version, so pipes were unaffected; with the tenant's version in every key they now take the same snapshot). The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. The snapshot is taken before any input a bump invalidates is chosen, the tenant's ClickHouse connection included: both handlers now look up before they resolve the tenant's pool, so a reload that moves the tenant to another address or database after a request took the old pool orphans that request's fill instead of filing the old database's rows as fresh under the new tenant version. A tenant on no pool is still a `503` before a cached result is served or a query runs. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. The cache now escapes table and scope names itself: a `Namespace` carries them raw and the version index builds its keys with `internal/keyenc`, so neither the structured-query read nor the ingest worker's invalidation escapes them (`query.SafeEncodeToken` is gone) and no name reaches a key unescaped. The suite pins that a name holding a dot, a space or a `%` is read and bumped under one key, and that names which would run together unescaped (`a.0.b` against `a` with scope `b.0.`) stay two entries. The version index's own entries are unaffected by the escaping; only the rendered key changes: it now carries the tenant's version and escapes the caller's query key whole. All of them live in the process, so nothing stored is orphaned. +- **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/query/ident.go` (removed, + tests), `internal/app/wire.go`, `docs/src/content/docs/{api,architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL (a pipe result's key folded no version, so pipes were unaffected; with the tenant's version in every key they now take the same snapshot). The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. The snapshot is taken before any input a bump invalidates is chosen, the tenant's ClickHouse connection included: both handlers now look up before they resolve the tenant's pool, so a reload that moves the tenant to another address or database after a request took the old pool orphans that request's fill instead of filing the old database's rows as fresh under the new tenant version. A tenant on no pool is still a `503` before a cached result is served or a query runs. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. The cache now escapes table and scope names itself: a `Namespace` carries them raw and the cache renders each result's key with `internal/keyenc`, so neither the structured-query read nor the ingest worker's invalidation escapes them (`query.SafeEncodeToken` is gone) and no name reaches a key unescaped. The suite pins that a name holding a dot, a space or a `%` is read and bumped under one key, and that names which would run together unescaped (`a.0.b` against `a` with scope `b.0.`) stay two entries. The version index's own entries are unaffected by the escaping; only the rendered key changes: it now carries the tenant's version and escapes the caller's query key whole. All of them live in the process, so nothing stored is orphaned. - **The cache's version index no longer grows with every bump, and forgets a tenant no longer served** (`internal/cache/{local,version_manager}.go` (+ tests), `internal/app/wire.go` (+ tests), `docs/src/content/docs/architecture.md`, `AGENTS.md`): part of [#262](https://github.com/Wave-RF/WaveHouse/issues/262) (growth across bumps and a departed tenant's memory; per-table scope cardinality is left open, see the issue) and of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The index nested each table under its tenant's version and each scope under its table's, and never pruned, so every tenant invalidation left the tenant's whole index behind, and it grew with every tenant ever served. It now holds one version per tenant, per (tenant, table) and per (tenant, table, scope), bumped in place. A tenant invalidation drops the tenant's index and hands its next key a generation unique within the process, so nothing cached before it can match again, and a table bump drops the table's scope versions. After each settings reload the index of every tenant no longer served, removed or rejected, is dropped the same way; its cached results are orphaned with it, as they already were when such a tenant came back on a pool. No change to what is cached or served. The Redis-compatible backend (#613) will bound its versions with a TTL instead. - **An explicit `false`, `0` or `""` in `config.yaml` is no longer replaced by the key's default** (`internal/config/config.go`, `internal/config/defaults_test.go` (new), `docs/src/content/docs/configuration.mdx`, `AGENTS.md`): [#631](https://github.com/Wave-RF/WaveHouse/issues/631). Defaults lived in cleanenv `env-default` tags, which cleanenv applies after the YAML decode to any field still at its zero value, so it could not tell a key the file set to its zero value from one the file left out. `otel.traces.enabled: false`, `otel.metrics.enabled: false` and `otel.logs.enabled: false` came back `true`; `otel.traces.sample_rate: 0` and `otel.logs.sample_rate: 0` came back `1.0`; `server.shutdown_timeout: 0` came back `10`; `cache.l1_max_cost: 0`, `prometheus.path: ""` and `data_dir: ""` came back as their defaults; `server.port: 0` came back `8080`. All of it was silent. Defaults now live in one Go function, `defaults()`, which `Load` starts from before decoding the file and then applying `WH_*` variables, so the order is env > YAML > default and a key the file sets always wins. **Behaviour change if your file relied on the bug:** a zero you wrote now takes effect. A file that says `sample_rate: 0` now exports no traces (or no DEBUG/INFO logs), where it silently exported everything; a signal set `enabled: false` is now off; `shutdown_timeout: 0` now skips the drain. `cache.l1_max_cost: 0`, `server.port: 0`, and `data_dir: ""` now refuse boot (`cache init: MaxCost can't be zero`, `server.port 0 out of range`, `data_dir (WH_DATA_DIR) is required`) instead of running on the default; an empty `prometheus.path` refuses boot when `prometheus.enabled` is true. Delete the key to get the default back. Env vars are unchanged: they already honoured an explicit zero. New tests load through `config.Load` for every affected key (a YAML zero is kept, an absent key gets the default, env wins in both directions), refuse an `env-default` tag on any field, and pin each documented default in `configuration.mdx` to `defaults()`. - **An embedded queue store that cannot be created fails boot at once, naming the cause** (`internal/mq/embedded.go` (+ tests)): part of [#617](https://github.com/Wave-RF/WaveHouse/issues/617). A regular file at `/nats`, or a `nats` directory that could not be created there, failed JetStream in the background, so boot waited out the server's 5s readiness check and reported only `nats server not ready`. `NewEmbedded` now creates the directory first (at `0700`, as the server does) and refuses boot with the mkdir error. An existing but unwritable `nats` directory still takes the old path. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 7aede5ba..678c52a3 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -114,7 +114,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on — one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)) — each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope by their raw names, which the cache escapes where it builds a key. `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The snapshot is taken before any input a bump invalidates is chosen, the tenant's connection included: a reload that moves the tenant to another address or database runs `Pools.Reconcile` and then `InvalidateTenant` (a repoint that keeps both, such as a username or `tls` change, reads the same tables and bumps nothing), so a request that took the old pool files the old database's rows under a version that bump orphans, whether its `Set` lands before the bump or after. The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. -- **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: one version per tenant, per (tenant, table) and per (tenant, table, scope), each keyed by its name alone and bumped in place, so the index holds one entry per live tenant, table and scope however often each is bumped ([#262](https://github.com/Wave-RF/WaveHouse/issues/262)). A query key folds the tenant's version and, for each dependency, its tenant's, table's and scope's, so bumping a table (a scopeless write) orphans every scope of it, and bumping one scope orphans that scope and the whole-table view — scope is reserved and empty today, so every write is the whole-table bump — all without touching the pool. Every field — the caller's query key, the tenant id, and each dependency's table and scope — is escaped and joined by `internal/keyenc` where the key is built, so a dot, a space or a `%` in a name is never read as a separator: each dependency renders as `..
.
..`, and the whole entry key is `|.||…`. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results. A tenant's version is a *generation*, unique within the process and handed out by the first key built for the tenant; `BumpTenant` (behind `InvalidateTenant`) drops the tenant's whole index, so the next key gets a fresh generation no cached entry folds, orphaning every cached result of the tenant in one step — a pipe result with no dependencies, and a table no bump ever keyed, included — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). `LocalCache.Prune` does the same for every tenant no longer served, which `internal/app` runs after each settings reload, so a tenant removed or rejected stops holding its index. A table bump drops the table's scope versions with it, since every key they were folded into also folds the old table version; and any bump (of a table, a scope or the tenant) under a tenant with no index is a no-op that records nothing, since the next key built for it gets a fresh generation no cached entry folds — so neither the `sharedTables` fan-out nor an insert still in flight for a tenant just pruned brings its index back. The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. +- **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: one version per tenant, per (tenant, table) and per (tenant, table, scope), each keyed by its name alone and bumped in place, so the index holds one entry per live tenant, table and scope however often each is bumped ([#262](https://github.com/Wave-RF/WaveHouse/issues/262)). An entry's key (`QueryKey`) folds the tenant's version and, for each dependency, its tenant's, table's and scope's, so bumping a table (a scopeless write) orphans every scope of it, and bumping one scope orphans that scope and the whole-table view — scope is reserved and empty today, so every write is the whole-table bump — all without touching the pool. Every field — the caller's query key, the tenant id, and each dependency's table and scope — is escaped and joined by `internal/keyenc` where the key is built, so a dot, a space or a `%` in a name is never read as a separator: each dependency renders as `..
.
..`, and the whole entry key is `|.||…`. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results. A tenant's version is a *generation*, unique within the process and handed out by the first key built for the tenant; `BumpTenant` (behind `InvalidateTenant`) drops the tenant's whole index, so the next key gets a fresh generation no cached entry folds, orphaning every cached result of the tenant in one step — a pipe result with no dependencies, and a table no bump ever keyed, included — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). `LocalCache.Prune` does the same for every tenant no longer served, which `internal/app` runs after each settings reload, so a tenant removed or rejected stops holding its index. A table bump drops the table's scope versions with it, since every key they were folded into also folds the old table version; and any bump (of a table, a scope or the tenant) under a tenant with no index is a no-op that records nothing, since the next key built for it gets a fresh generation no cached entry folds — so neither the `sharedTables` fan-out nor an insert still in flight for a tenant just pruned brings its index back. The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. ### `config/` — Configuration diff --git a/internal/cache/version_manager_test.go b/internal/cache/version_manager_test.go index 84792054..07945306 100644 --- a/internal/cache/version_manager_test.go +++ b/internal/cache/version_manager_test.go @@ -164,10 +164,26 @@ func TestVersionManager_GenerationsNeverRepeat(t *testing.T) { func TestVersionManager_BumpWithoutIndex(t *testing.T) { t.Parallel() vm := NewVersionManager() + vm.BumpTable("acme", "users") + assert.Zero(t, vm.size(), "a table bump for a tenant with no index creates nothing") + vm.BumpNamespace(Namespace{Tenant: "acme", Table: "users", Scope: "org_1"}) + assert.Zero(t, vm.size(), "a namespace bump for a tenant with no index creates nothing") + vm.BumpTenant("acme") - assert.Zero(t, vm.size()) + assert.Zero(t, vm.size(), "bumping a tenant with no index is a no-op") + + // An insert still in flight for a tenant just pruned must not bring its + // index back: a write racing the prune sees the tenant gone and bumps + // blind, same as above. + vm.QueryKey("acme", "h", nil) + vm.Prune(func(tenant.ID) bool { return false }) + assert.Zero(t, vm.size(), "prune released the tenant's index") + + vm.BumpTable("acme", "users") + vm.BumpNamespace(Namespace{Tenant: "acme", Table: "users", Scope: "org_1"}) + assert.Zero(t, vm.size(), "a bump for a tenant just pruned must not recreate its index") } func TestVersionManager_Prune(t *testing.T) { From 5a72724b83afe4b14bcf9f0f43f66ccd105f9048 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:40:39 -0400 Subject: [PATCH 22/27] =?UTF-8?q?fix(cache):=20review=20round=202=20?= =?UTF-8?q?=E2=80=94=20stale=20docs=20claims,=20redisConfig=20field=20pin?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - configuration.mdx: isAuthError also matches NOPERM (an ACL user missing a connection command); the boot-error-logging line only named WRONGPASS/NOAUTH. - CHANGELOG.md: the E4 entry claimed the e2e coverage exclude for the cache files was gone; .testcoverage.yml still excludes internal/cache/pending.go, internal/config/cache_redis.go and internal/cache/(local|version_manager).go from e2e, for reasons the entry now states. Also added the docs/.github files this PR touched that the entry's file list was missing. - pipes.mdx, api.md: "L1" claims that are wrong once cache.backend is redis (there is no L1) — reworded to "the query cache"/"cache". - deployment.md: "Persistence is not needed" was incomplete — a restart *with* persistence reloads a stale snapshot and can serve invalidated results as hits until their TTL. Now says to run without persistence, and that restoring a snapshot is a rollback. - development.md: note that shared_cache_test.go starts its own Redis per test and boots extra cache.backend=redis instances over the integration suite's ClickHouse. - app_test.go: TestRedisConfig_FromLoadedDefaults left Username, DB and TLS at their zero value on both sides of its assert.Equal, so deleting any of their three mapping lines in redisConfig wouldn't fail it. Added TestRedisConfig_UsernameDBTLSMapped, which drives all three to a non-zero value through config.Load. Mutation-checked: deleting each of the three lines in redisConfig fails this test (the TLS line fails to compile instead, since dropping it leaves the local `t` unused — still a build failure). Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- CHANGELOG.md | 2 +- docs/src/content/docs/api.md | 2 +- docs/src/content/docs/configuration.mdx | 2 +- docs/src/content/docs/deployment.md | 2 +- docs/src/content/docs/development.md | 2 +- docs/src/content/docs/pipes.mdx | 2 +- internal/app/app_test.go | 22 ++++++++++++++++++++++ 7 files changed, 28 insertions(+), 6 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index c4e6e293..472f94af 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `internal/ingest/worker.go`, `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`) and `dial_timeout` (`1s`), each at most `1s` since boot and shutdown each wait out a connection attempt they bound, `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a valid port, more than one address in `standalone` mode, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The ingest worker's log of an invalidation that did not land drops from `ERROR` to `WARN`, since the shared backend defers and retries it: an outage would otherwise log an `ERROR` for every batch. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits; a third runs the first one's hit, insert and fresh-miss lifecycle on the suite's own `cache.backend: local` app, since e2e no longer exercises that backend. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. +- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `internal/ingest/worker.go`, `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `.github/workflows/{ci.yml,README.md}`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md,index.mdx,why-wavehouse.md,sdk/reference.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`) and `dial_timeout` (`1s`), each at most `1s` since boot and shutdown each wait out a connection attempt they bound, `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a valid port, more than one address in `standalone` mode, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The ingest worker's log of an invalidation that did not land drops from `ERROR` to `WARN`, since the shared backend defers and retries it: an outage would otherwise log an `ERROR` for every batch. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end; the e2e per-suite exclude now names what its run still can't reach instead — `internal/cache/pending.go` (the retry of an invalidation the server did not take, which needs an outage), `internal/config/cache_redis.go` (the block's own rejection paths) and `internal/cache/(local|version_manager).go` (the `local` backend, which e2e no longer runs) — and the unit and integration suites keep covering them. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits; a third runs the first one's hit, insert and fresh-miss lifecycle on the suite's own `cache.backend: local` app, since e2e no longer exercises that backend. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. - **A Redis-compatible shared cache backend** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/cache.go`, `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`, `CONTRIBUTING.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), open a circuit breaker that skips the server until a probe write succeeds (connections are replaced every minute, so after a failover behind a stable address the process reaches the new primary, and delivers the bumps it owes, within about a minute, provided a fresh connection's dial and handshake fit in the per-operation timeout), and deferred invalidations are retried until they land — the first at once, and at the probe's cadence while the breaker is open — collapsing to one tenant-wide bump per tenant past 100,000 keys; until one lands, the process that owes it bypasses the lookups it would orphan. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. `cache.backend: redis` selects it (the entry above). Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression, server-stops-answering, refused-writes (a demoted primary, a full `noeviction` server), failover-behind-a-stable-address and owed-invalidation cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index 0ed66f23..77bbe792 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -445,7 +445,7 @@ ClickHouse's inline `FORMAT` clause (e.g. `SELECT 1 FORMAT CSV` or `… FORMAT P The proxy buffers the upstream response in memory before forwarding (no row-streaming yet), so a `SELECT *` from a large table can pin RAM on the API server. To avoid an admin OOMing themselves, responses larger than 64 MiB return 502 with a `clickhouse response exceeded N bytes` error. Narrow the query with `LIMIT`, or use a streaming client outside WaveHouse that talks to ClickHouse directly (the standard escape hatch — the same admin credentials work). ::: -This endpoint **does not cache, does not singleflight, and emits `Cache-Control: no-store`** — every request goes straight to ClickHouse, mutation or read, and downstream HTTP caches are explicitly told not to store the response. Raw SQL is an admin escape hatch with infrequent, ad-hoc traffic, so the L1/singleflight machinery would only add complexity without a real hit-rate win. Use [`POST /v1/query?table={table}`](#post-v1querytabletable--structured-query) or [`GET/POST /v1/pipes/{name}`](#getpost-v1pipesname--execute-named-pipe) for the cached read paths (dashboards, high-QPS clients, etc.) — both go through the query cache ([`cache.backend`](/configuration#backends): in-process, or a Redis shared by every instance) with singleflight coalescing. +This endpoint **does not cache, does not singleflight, and emits `Cache-Control: no-store`** — every request goes straight to ClickHouse, mutation or read, and downstream HTTP caches are explicitly told not to store the response. Raw SQL is an admin escape hatch with infrequent, ad-hoc traffic, so the cache/singleflight machinery would only add complexity without a real hit-rate win. Use [`POST /v1/query?table={table}`](#post-v1querytabletable--structured-query) or [`GET/POST /v1/pipes/{name}`](#getpost-v1pipesname--execute-named-pipe) for the cached read paths (dashboards, high-QPS clients, etc.) — both go through the query cache ([`cache.backend`](/configuration#backends): in-process, or a Redis shared by every instance) with singleflight coalescing. :::note[Admin only] The route is mounted under `/v1/ops/*`, behind the `RequireAdmin` gate: only a caller whose JWT role equals the policy `admin_role` (`"admin"` by default) — or who presents the non-JWT [operator key](#authentication) — may use it. A tokenless request (or a valid token without a role claim) resolves to the `default_role` (not the admin role unless `default_role` is deliberately set to it — a loudly-warned dev-only setting) and is rejected with `403`; a present-but-invalid token — expired, malformed, bad signature — keeps its stashed verification error and fails loud with `401` instead. Raw SQL has no per-statement scope check (a full SQL parser would be needed to authorize predicates), so the role gate is the entire authorization story, shared with the rest of `/v1/ops/*` (see [Admin Endpoints](#admin-endpoints)). The normal surfaces for non-admin callers are `POST /v1/ingest?table={table}` for writes, `POST /v1/query?table={table}` for structured reads, and `GET/POST /v1/pipes/{name}` for pre-defined queries — none of which expose raw SQL. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index d625b2e4..de970487 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -170,7 +170,7 @@ The `redis` backend's settings, read only when `cache.backend` is `redis`. It is **When the server is unreachable or misbehaves, the cache is bypassed; queries are not.** A failure or a timeout makes the lookup a miss and the fill a no-op. Five in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), open a circuit breaker that skips the server entirely until a probe write, every 5 s, succeeds. Queries then go straight to ClickHouse, still coalesced per instance by `singleflight`. An invalidation the server did not take is kept and retried until it lands, and until then the instance that owes it bypasses the lookups it would orphan. `/readyz` does not depend on the cache. -**Boot does not wait for the server.** A malformed block (an address without a port, `mode: cluster` with `db` other than `0`, an unreadable or unparsable TLS file) refuses boot. A server that cannot be reached, or that refuses the credentials, does not: the process boots with the cache bypassed and keeps reconnecting, with backoff up to 30 s. A rejected password (`WRONGPASS`, `NOAUTH`) is logged at `ERROR` on every attempt; any other failure at `WARN`. This is deliberate: a rotated Redis password must not crash-loop every instance at once. Watch `wavehouse_cache_breaker_open`, which reads `1` while the cache is bypassed, including before the first connection. +**Boot does not wait for the server.** A malformed block (an address without a port, `mode: cluster` with `db` other than `0`, an unreadable or unparsable TLS file) refuses boot. A server that cannot be reached, or that refuses the credentials, does not: the process boots with the cache bypassed and keeps reconnecting, with backoff up to 30 s. A rejected credential (`WRONGPASS`, `NOAUTH`, or `NOPERM` for an ACL user missing a connection command) is logged at `ERROR` on every attempt; any other failure at `WARN`. This is deliberate: a rotated Redis password must not crash-loop every instance at once. Watch `wavehouse_cache_breaker_open`, which reads `1` while the cache is bypassed, including before the first connection. ### Authentication diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 1e14399b..aa29bb6b 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -437,7 +437,7 @@ The query-result cache is the layer that can be shared today. With the default ` - **A pipe that writes** (an `INSERT` in `pipes.json`) has its result cached like a read, so a repeated identical call is answered from the cache and the write does not run again ([#386](https://github.com/Wave-RF/WaveHouse/issues/386)). With a shared cache that holds on every instance, until the entry's TTL. - **Admin writes through `POST /v1/ops/query`** do not invalidate the cache ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). With a shared cache, the stale results they leave are served by every instance, not only one. -**Sizing the server.** Every key WaveHouse writes has a TTL, and a version token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses, never bring back an entry it had invalidated. So set `maxmemory` and let the server evict: `maxmemory-policy allkeys-lru` (or `allkeys-lfu`, `volatile-lru`, `volatile-lfu`). Under `noeviction`, a full server refuses the writes: each refusal (a fill's is counted by `wavehouse_cache_set_failures_total{reason="oom"}`) bypasses the cache of the instance that got it, and invalidations are kept and retried, so the pre-insert results above stay served by the others: avoid `noeviction`. A stored result is capped at `cache.redis.max_value_bytes` (1 MiB compressed). A tenant's version tokens share one hash tag, so each lookup reads them in one `MGET` in cluster mode as well. The results themselves carry no hash tag and spread across shards. Persistence is not needed: an empty server after a restart is a cold cache, not a wrong one. +**Sizing the server.** Every key WaveHouse writes has a TTL, and a version token lost to eviction, expiry or `FLUSHALL` can only cause misses, never bring back an entry it had invalidated. So set `maxmemory` and let the server evict: `maxmemory-policy allkeys-lru` (or `allkeys-lfu`, `volatile-lru`, `volatile-lfu`). Under `noeviction`, a full server refuses the writes: each refusal (a fill's is counted by `wavehouse_cache_set_failures_total{reason="oom"}`) bypasses the cache of the instance that got it, and invalidations are kept and retried, so the pre-insert results above stay served by the others: avoid `noeviction`. A stored result is capped at `cache.redis.max_value_bytes` (1 MiB compressed). A tenant's version tokens share one hash tag, so each lookup reads them in one `MGET` in cluster mode as well. The results themselves carry no hash tag and spread across shards. **Run it without persistence** (`save ""` and `appendonly no`): a restart without persistence can only cause misses, the same as any other token loss. With persistence on, a restart is not that — it reloads whatever snapshot or AOF it last wrote, tokens and values it had already invalidated included, so a fresh instance can serve the pre-write rows filed under them as hits until their TTL (up to 1 h) expires. Treat restoring a snapshot as a rollback, not a resume. **The server is inside the trust boundary.** A cached result is served after the access policy has filtered it, so whoever can write to the server can change what any caller reads. Keep it on a private network, require a password or ACL user (`WH_CACHE_REDIS_PASSWORD`), use TLS across links you do not trust, and share it only with deployments you trust as much as this one. diff --git a/docs/src/content/docs/development.md b/docs/src/content/docs/development.md index 581df72d..11fcb82d 100644 --- a/docs/src/content/docs/development.md +++ b/docs/src/content/docs/development.md @@ -345,7 +345,7 @@ Each test target writes `covdata` to `tmp/coverage//data/`, renders a tex | E2E tests (SDK) | `tests/e2e/sdk/*.test.ts` | Yes | `make test-e2e` | - **Unit tests** live beside the code they test (e.g., `internal/discovery/discovery_test.go`). They use mocks or embedded NATS (in-process, no Docker needed). -- **Integration tests** use the `//go:build integration` build tag. In `tests/integration`, `TestMain` starts one ClickHouse testcontainer and boots the production wiring against it through `app.New` (embedded NATS, ingest worker, sweeper, hub, the API server on a random loopback port); tests reach it via `env(t)` and create their own tables. DLQ tests use `assert.Eventually` with a 30-second timeout for the 5-second ingest worker batch window. `internal/cache`'s integration tests start their own containers instead — Redis, Valkey, Dragonfly and a one-node Redis Cluster — for the shared backend. +- **Integration tests** use the `//go:build integration` build tag. In `tests/integration`, `TestMain` starts one ClickHouse testcontainer and boots the production wiring against it through `app.New` (embedded NATS, ingest worker, sweeper, hub, the API server on a random loopback port); tests reach it via `env(t)` and create their own tables. DLQ tests use `assert.Eventually` with a 30-second timeout for the 5-second ingest worker batch window. `internal/cache`'s integration tests start their own containers instead — Redis, Valkey, Dragonfly and a one-node Redis Cluster — for the shared backend. `shared_cache_test.go` starts its own Redis testcontainer per test (`startRedis`) and boots extra, independent `cache.backend: redis` instances over that same ClickHouse (`bootRedisApp`), to exercise the cache shared across processes rather than one package in isolation. Shared test utilities live in `internal/testutil/`. The packages log through `slog.Default()`, so tests reach log output through `internal/testutil/logtest`: `logtest.Silence()` in a package's `TestMain` discards it, and `logtest.Capture(t, level)` routes it to a buffer for a test that asserts on log lines — such a test must not call `t.Parallel()`, because the default logger is process-wide. diff --git a/docs/src/content/docs/pipes.mdx b/docs/src/content/docs/pipes.mdx index b6412922..b6ded200 100644 --- a/docs/src/content/docs/pipes.mdx +++ b/docs/src/content/docs/pipes.mdx @@ -7,7 +7,7 @@ sidebar: A **named pipe** is a saved SQL query, registered under a name, that callers run by name with parameters — without ever sending raw SQL. They turn an ad-hoc query into a stable, cached, access-controlled endpoint: you write the SQL once as an operator in the settings directory's [`pipes.json`](/settings-directory#pipesjson), expose it at `GET/POST /v1/pipes/{name}`, and clients supply only the declared parameters. -Pipes are the right tool when a query is reusable and shouldn't live in client code — dashboards, reports, public APIs over curated slices of data. They sit on the **cached read path** (shared L1 + singleflight, same as structured queries), and authorize through a simple per-pipe allowlist rather than the full [policy engine](/access-control). +Pipes are the right tool when a query is reusable and shouldn't live in client code — dashboards, reports, public APIs over curated slices of data. They sit on the **cached read path** (the query cache + singleflight, same as structured queries), and authorize through a simple per-pipe allowlist rather than the full [policy engine](/access-control). ## Anatomy of a pipe diff --git a/internal/app/app_test.go b/internal/app/app_test.go index 7a23330e..945f7eb9 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -832,6 +832,28 @@ func TestRedisConfig_FromLoadedDefaults(t *testing.T) { assert.Equal(t, cache.RedisSentinel, config.RedisSentinel) } +// Username, DB and TLS are zero on both sides of TestRedisConfig_FromLoadedDefaults' +// assert.Equal, so deleting any of their three mapping lines in redisConfig +// would pass it anyway. Drive all three through config.Load to a non-zero +// value and assert on them directly. +func TestRedisConfig_UsernameDBTLSMapped(t *testing.T) { + t.Setenv("WH_SETTINGS_DIR", t.TempDir()) + t.Setenv("WH_CACHE_BACKEND", "redis") + t.Setenv("WH_CACHE_REDIS_ADDRS", "a:6379") + t.Setenv("WH_CACHE_REDIS_USERNAME", "u") + t.Setenv("WH_CACHE_REDIS_DB", "2") + t.Setenv("WH_CACHE_REDIS_TLS_ENABLED", "true") + t.Setenv("WH_CACHE_REDIS_TLS_SERVER_NAME", "r.internal") + loaded, err := config.Load(filepath.Join(t.TempDir(), "none.yaml")) + require.NoError(t, err) + got, err := redisConfig(loaded.Cache.Redis) + require.NoError(t, err) + assert.Equal(t, "u", got.Username) + assert.Equal(t, 2, got.DB) + require.NotNil(t, got.TLS) + assert.Equal(t, "r.internal", got.TLS.ServerName) +} + // keepalive is a config.json patch setting the stream block's keepalive pair. func keepalive(interval, buckets int) map[string]any { return map[string]any{"stream": map[string]any{"keepalive_interval": interval, "keepalive_buckets": buckets, "gap_window_minutes": 15}} From 568f2ddd20a826be68fd7bc8800e2fd2d7524df5 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:46:46 -0400 Subject: [PATCH 23/27] docs(cache): version_manager.go's own doc carries the query-key/entry-key mixup too Same overload as architecture.md's version_manager.go bullet: the type doc's "A query key folds all three versions of each dependency" means the rendered entry key QueryKey returns, not the caller's input. Reword to match. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- internal/cache/version_manager.go | 7 ++++--- 1 file changed, 4 insertions(+), 3 deletions(-) diff --git a/internal/cache/version_manager.go b/internal/cache/version_manager.go index d67a6077..d6fd7496 100644 --- a/internal/cache/version_manager.go +++ b/internal/cache/version_manager.go @@ -16,9 +16,10 @@ import ( // place, so the index holds one entry per live tenant, table and scope // however often each is bumped, and forgetting a tenant releases all of it. // -// A query key folds all three versions of each dependency, which gives the -// lattice: a table bump orphans every scope, a scope bump that scope and the -// whole-table view, and a tenant bump everything of the tenant's. +// An entry's key (`QueryKey`) folds all three versions of each dependency, +// which gives the lattice: a table bump orphans every scope, a scope bump +// that scope and the whole-table view, and a tenant bump everything of the +// tenant's. type VersionManager struct { mu sync.RWMutex From c42b6db2603dc8d517441bba981dc6685833c135 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 09:04:44 -0400 Subject: [PATCH 24/27] docs(cache): fix the one query-key/entry-key spot the terminology pass missed MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit version_manager.go's tenants field comment still said "the first query key built for it" — the same overload the type doc three lines above and the other doc fixes in this round eliminated (query key = the caller's input, entry key = what QueryKey renders). Both confirmation reviewers caught this independently. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- internal/cache/version_manager.go | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/internal/cache/version_manager.go b/internal/cache/version_manager.go index d6fd7496..6b1de3b5 100644 --- a/internal/cache/version_manager.go +++ b/internal/cache/version_manager.go @@ -23,8 +23,8 @@ import ( type VersionManager struct { mu sync.RWMutex - // tenants holds each tenant's index from the first query key built for - // it until the tenant is bumped or pruned. + // tenants holds each tenant's index from the first entry key (`QueryKey`) + // built for it until the tenant is bumped or pruned. tenants map[tenant.ID]*tenantVersions // lastGen is the last generation handed to a tenant; see tenantVersions.gen. From ce0da484740928e07cd721831ed5afd54159c401 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 09:32:52 -0400 Subject: [PATCH 25/27] docs(cache): match configuration.mdx and deployment.md to #626's fixes Verified against the merged internal/cache/{redis,breaker}.go: - rejectsCredentials (WRONGPASS, NOAUTH) now trips the breaker at once at runtime too, logged at ERROR once per opening and once per refused probe (breaker.trip()); NOPERM deliberately does not (it names one key/command, not every operation, so an ACL denial on one tenant's traffic shouldn't bypass the cache for all of them). configuration.mdx's circuit-breaker paragraph didn't mention this at all; added it, distinct from the dial-time isAuthError logging (a separate function, used only in dial()/dialLoop(), where WRONGPASS, NOAUTH and NOPERM are all ERROR since any of them blocks the connection outright before a client exists to send a scoped command on). - deployment.md's failover bullet gets the probe's DialTimeout+Timeout budget and the caveat that every other connection still redials under Timeout alone, matching architecture.md's redis.go bullet. - deployment.md's metrics list presented invalidations_total's `ok`/ `deferred` as if they partition the count; they overlap (a retried landing counts `ok` again), per metrics.go's own doc comment. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- docs/src/content/docs/configuration.mdx | 2 +- docs/src/content/docs/deployment.md | 4 ++-- 2 files changed, 3 insertions(+), 3 deletions(-) diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index de970487..c8878afe 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -168,7 +168,7 @@ The `redis` backend's settings, read only when `cache.backend` is `redis`. It is | `cache.redis.compress_min_bytes` | `WH_CACHE_REDIS_COMPRESS_MIN_BYTES` | `1024` | Results at least this large are zstd-compressed when that makes them smaller. `0` never compresses. | | `cache.redis.version_ttl` | `WH_CACHE_REDIS_VERSION_TTL` | `168h` | How long a table's or tenant's version token outlives its last write, so the tokens of dropped tables and removed tenants eventually expire. At least `2s`. An expired token only causes misses. | -**When the server is unreachable or misbehaves, the cache is bypassed; queries are not.** A failure or a timeout makes the lookup a miss and the fill a no-op. Five in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), open a circuit breaker that skips the server entirely until a probe write, every 5 s, succeeds. Queries then go straight to ClickHouse, still coalesced per instance by `singleflight`. An invalidation the server did not take is kept and retried until it lands, and until then the instance that owes it bypasses the lookups it would orphan. `/readyz` does not depend on the cache. +**When the server is unreachable or misbehaves, the cache is bypassed; queries are not.** A failure or a timeout makes the lookup a miss and the fill a no-op. Five in a row, one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), or one refusing the credentials (`WRONGPASS`, `NOAUTH` — what an already-open connection meets once the password is rotated; logged at `ERROR`, once per opening and once per refused probe) open a circuit breaker that skips the server entirely until a probe write, every 5 s, succeeds. `NOPERM` does not open it: it names one key or command an ACL user cannot use, not every operation, so it should not bypass the cache for every tenant. Queries then go straight to ClickHouse, still coalesced per instance by `singleflight`. An invalidation the server did not take is kept and retried until it lands, and until then the instance that owes it bypasses the lookups it would orphan. `/readyz` does not depend on the cache. **Boot does not wait for the server.** A malformed block (an address without a port, `mode: cluster` with `db` other than `0`, an unreadable or unparsable TLS file) refuses boot. A server that cannot be reached, or that refuses the credentials, does not: the process boots with the cache bypassed and keeps reconnecting, with backoff up to 30 s. A rejected credential (`WRONGPASS`, `NOAUTH`, or `NOPERM` for an ACL user missing a connection command) is logged at `ERROR` on every attempt; any other failure at `WARN`. This is deliberate: a rotated Redis password must not crash-loop every instance at once. Watch `wavehouse_cache_breaker_open`, which reads `1` while the cache is bypassed, including before the first connection. diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index fcbcc252..b3fdcf80 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -432,7 +432,7 @@ The query-result cache is the layer that can be shared today. With the default ` **What another instance can see.** Ingest is already asynchronous: `/v1/ingest` answers before the batch is inserted. Once the inserting instance's worker has written the batch to ClickHouse, it replaces the table's version token in Redis, and from then on a lookup on any instance misses and reads the new rows. The cache adds no delay of its own beyond that single write. The exceptions: - **The server is unreachable from the inserting instance.** The invalidation is kept and retried until it lands (`wavehouse_cache_invalidations_pending` counts what is owed). Meanwhile other instances that can still reach the server keep serving the older results, for as long as the outage lasts and at most until each entry's TTL. An instance that stops while invalidations are still owed loses them, with the same bound. The same thing happens today when a process stops between an insert and its invalidation. -- **A failover to a replica that had not yet received the latest token writes** can bring back entries filed under the older tokens, bounded by the replication lag at the moment of failover and those entries' TTL. WaveHouse never reads from replicas. Behind a stable address (a managed primary endpoint), an instance still connected to the demoted node has its writes refused, which bypasses its cache; connections are replaced every minute, so it reaches the new primary and delivers the invalidations it owes within about that long. +- **A failover to a replica that had not yet received the latest token writes** can bring back entries filed under the older tokens, bounded by the replication lag at the moment of failover and those entries' TTL. WaveHouse never reads from replicas. Behind a stable address (a managed primary endpoint), an instance still connected to the demoted node has its writes refused, which bypasses its cache; connections are replaced every minute, so it reaches the new primary and delivers the invalidations it owes within about that long. The breaker's own probe gets a longer reconnect budget (`dial_timeout` plus `timeout`) so a slow reconnect still closes it, but every other connection redials under `timeout` alone: size it above how long a reconnect actually takes, or operations that land on one of those keep failing after the probe has already succeeded. - **The server is full and `maxmemory-policy` is `noeviction`.** It refuses the token writes. The inserting instance keeps its invalidations and retries them, bypassing its cache meanwhile, but every other instance serves the results from before the insert until one lands, up to their TTL. - **A pipe that writes** (an `INSERT` in `pipes.json`) has its result cached like a read, so a repeated identical call is answered from the cache and the write does not run again ([#386](https://github.com/Wave-RF/WaveHouse/issues/386)). With a shared cache that holds on every instance, until the entry's TTL. - **Admin writes through `POST /v1/ops/query`** do not invalidate the cache ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). With a shared cache, the stale results they leave are served by every instance, not only one. @@ -443,7 +443,7 @@ The query-result cache is the layer that can be shared today. With the default ` **Coalescing stays per instance.** `singleflight` collapses identical concurrent queries within each instance, so a cold hot query costs at most one ClickHouse query per instance, not one per request. -**Metrics** (meter `wavehouse-cache`, every series labeled `backend="redis"`, no tenant label): `wavehouse_cache_lookups_total{result}` (`hit`, `miss`, `stale`, `bypass`, `error`), `wavehouse_cache_op_duration_seconds{op}` (`lookup`, `set`, `invalidate`), `wavehouse_cache_breaker_open` (1 while the cache is bypassed), `wavehouse_cache_invalidations_total{result}` (`ok`, `deferred`), `wavehouse_cache_invalidations_pending`, `wavehouse_cache_value_bytes`, `wavehouse_cache_oversize_total` and `wavehouse_cache_set_failures_total{reason}` (`oom`, `timeout`, `other`). Two signals are worth alerting on: `wavehouse_cache_breaker_open` at 1, or `wavehouse_cache_invalidations_pending` above 0, for more than a few minutes. +**Metrics** (meter `wavehouse-cache`, every series labeled `backend="redis"`, no tenant label): `wavehouse_cache_lookups_total{result}` (`hit`, `miss`, `stale`, `bypass`, `error`), `wavehouse_cache_op_duration_seconds{op}` (`lookup`, `set`, `invalidate`), `wavehouse_cache_breaker_open` (1 while the cache is bypassed), `wavehouse_cache_invalidations_total{result}` (`ok` counts every bump that lands, retried ones included, and `deferred` each time one is put off, a repeat of one already owed included — the two overlap, not a split), `wavehouse_cache_invalidations_pending`, `wavehouse_cache_value_bytes`, `wavehouse_cache_oversize_total` and `wavehouse_cache_set_failures_total{reason}` (`oom`, `timeout`, `other`). Two signals are worth alerting on: `wavehouse_cache_breaker_open` at 1, or `wavehouse_cache_invalidations_pending` above 0, for more than a few minutes. For local development, `docker compose -f deployments/compose/dependencies.yaml --profile redis up -d` starts a Redis on `localhost:6379` with no persistence. From ecd21c0890b5e84e5e830547a32e1e7a6c98233f Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 13:10:19 -0400 Subject: [PATCH 26/27] docs(cache): match deployment.md to #626's probe-fix (0eea1cbc) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Verified against the merged internal/cache/redis.go: - The failover bullet said the probe gets DialTimeout+Timeout; it's now 2×DialTimeout+Timeout for the first write (rueidis bounds the dial and the handshake by DialTimeout in turn), and a write slower than Timeout is repeated under Timeout alone — only the repeat decides whether the breaker closes, so a server that merely answers slowly stays bypassed instead of flapping open and shut every cycle. - Confirmed the probe is unaffected by, and doesn't affect, the boot/ Close 1s-cap arithmetic in configuration.mdx's dial_timeout row and cache_redis.go's maxRedisTimeout comment: probe() runs in an untracked goroutine (no r.wg.Add before `go r.probe(...)`), so Close's r.wg.Wait() never waits on it. No doc change needed there. - "Run it without persistence" didn't say why: stock Redis and Valkey persist by default (periodic RDB save points), so an ordinary crash- restart reloads the last save on its own — the same rollback as restoring a snapshot by hand. Reworded to match #626's CHANGELOG/ architecture.md wording (grepped for other copies; found none). Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- docs/src/content/docs/deployment.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index b3fdcf80..0ba84aa5 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -432,12 +432,12 @@ The query-result cache is the layer that can be shared today. With the default ` **What another instance can see.** Ingest is already asynchronous: `/v1/ingest` answers before the batch is inserted. Once the inserting instance's worker has written the batch to ClickHouse, it replaces the table's version token in Redis, and from then on a lookup on any instance misses and reads the new rows. The cache adds no delay of its own beyond that single write. The exceptions: - **The server is unreachable from the inserting instance.** The invalidation is kept and retried until it lands (`wavehouse_cache_invalidations_pending` counts what is owed). Meanwhile other instances that can still reach the server keep serving the older results, for as long as the outage lasts and at most until each entry's TTL. An instance that stops while invalidations are still owed loses them, with the same bound. The same thing happens today when a process stops between an insert and its invalidation. -- **A failover to a replica that had not yet received the latest token writes** can bring back entries filed under the older tokens, bounded by the replication lag at the moment of failover and those entries' TTL. WaveHouse never reads from replicas. Behind a stable address (a managed primary endpoint), an instance still connected to the demoted node has its writes refused, which bypasses its cache; connections are replaced every minute, so it reaches the new primary and delivers the invalidations it owes within about that long. The breaker's own probe gets a longer reconnect budget (`dial_timeout` plus `timeout`) so a slow reconnect still closes it, but every other connection redials under `timeout` alone: size it above how long a reconnect actually takes, or operations that land on one of those keep failing after the probe has already succeeded. +- **A failover to a replica that had not yet received the latest token writes** can bring back entries filed under the older tokens, bounded by the replication lag at the moment of failover and those entries' TTL. WaveHouse never reads from replicas. Behind a stable address (a managed primary endpoint), an instance still connected to the demoted node has its writes refused, which bypasses its cache; connections are replaced every minute, so it reaches the new primary and delivers the invalidations it owes within about that long. The breaker's own probe write gets a longer budget for a reconnect — twice `dial_timeout` (the client bounds the dial and the handshake by it in turn) plus `timeout` for the write itself — but only closes the breaker when a write actually lands within `timeout`: one slower than that is repeated under `timeout` alone, and the repeat decides, so a server that merely answers slowly stays bypassed instead of flapping open and shut. Every other connection redials under `timeout` alone: size it above how long a reconnect actually takes, or operations that land on one of those keep failing after the probe has already succeeded. - **The server is full and `maxmemory-policy` is `noeviction`.** It refuses the token writes. The inserting instance keeps its invalidations and retries them, bypassing its cache meanwhile, but every other instance serves the results from before the insert until one lands, up to their TTL. - **A pipe that writes** (an `INSERT` in `pipes.json`) has its result cached like a read, so a repeated identical call is answered from the cache and the write does not run again ([#386](https://github.com/Wave-RF/WaveHouse/issues/386)). With a shared cache that holds on every instance, until the entry's TTL. - **Admin writes through `POST /v1/ops/query`** do not invalidate the cache ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). With a shared cache, the stale results they leave are served by every instance, not only one. -**Sizing the server.** Every key WaveHouse writes has a TTL, and a version token lost to eviction, expiry or `FLUSHALL` can only cause misses, never bring back an entry it had invalidated. So set `maxmemory` and let the server evict: `maxmemory-policy allkeys-lru` (or `allkeys-lfu`, `volatile-lru`, `volatile-lfu`). Under `noeviction`, a full server refuses the writes: each refusal (a fill's is counted by `wavehouse_cache_set_failures_total{reason="oom"}`) bypasses the cache of the instance that got it, and invalidations are kept and retried, so the pre-insert results above stay served by the others: avoid `noeviction`. A stored result is capped at `cache.redis.max_value_bytes` (1 MiB compressed). A tenant's version tokens share one hash tag, so each lookup reads them in one `MGET` in cluster mode as well. The results themselves carry no hash tag and spread across shards. **Run it without persistence** (`save ""` and `appendonly no`): a restart without persistence can only cause misses, the same as any other token loss. With persistence on, a restart is not that — it reloads whatever snapshot or AOF it last wrote, tokens and values it had already invalidated included, so a fresh instance can serve the pre-write rows filed under them as hits until their TTL (up to 1 h) expires. Treat restoring a snapshot as a rollback, not a resume. +**Sizing the server.** Every key WaveHouse writes has a TTL, and a version token lost to eviction, expiry or `FLUSHALL` can only cause misses, never bring back an entry it had invalidated. So set `maxmemory` and let the server evict: `maxmemory-policy allkeys-lru` (or `allkeys-lfu`, `volatile-lru`, `volatile-lfu`). Under `noeviction`, a full server refuses the writes: each refusal (a fill's is counted by `wavehouse_cache_set_failures_total{reason="oom"}`) bypasses the cache of the instance that got it, and invalidations are kept and retried, so the pre-insert results above stay served by the others: avoid `noeviction`. A stored result is capped at `cache.redis.max_value_bytes` (1 MiB compressed). A tenant's version tokens share one hash tag, so each lookup reads them in one `MGET` in cluster mode as well. The results themselves carry no hash tag and spread across shards. **Run it without persistence** (`save ""` and `appendonly no`): stock Redis and Valkey persist by default (periodic RDB save points), so a crash that is followed by a restart reloads the last save on its own — the same rollback as restoring a snapshot by hand, not a loss. Without persistence, a restart can only cause misses, the same as any other token loss. With it on (the default), a restart reloads whatever snapshot or AOF it last wrote, tokens and values it had already invalidated included, so a fresh instance can serve the pre-write rows filed under them as hits until their TTL (up to 1 h) expires. Treat restoring a snapshot, or a crash-restart on a server that still has its defaults, as a rollback, not a resume. **The server is inside the trust boundary.** A cached result is served after the access policy has filtered it, so whoever can write to the server can change what any caller reads. Keep it on a private network, require a password or ACL user (`WH_CACHE_REDIS_PASSWORD`), use TLS across links you do not trust, and share it only with deployments you trust as much as this one. From 0281ca31efe0327b9428b8087f68a4abef2df103 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 14:36:05 -0400 Subject: [PATCH 27/27] docs(cache): correct the Close bound, deferred count and probe wording Close can wait up to 4s, not 3s: a dial in flight, then the 1s final drain. deferred counts each deferral once, not failed retries. The probe write closes the breaker only within timeout. The breaker's WARN on opening is documented alongside the credentials ERROR. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01FyrXjhR7iDg33paioLQHFq --- docs/src/content/docs/configuration.mdx | 2 +- docs/src/content/docs/deployment.md | 2 +- internal/config/cache_redis.go | 5 +++-- 3 files changed, 5 insertions(+), 4 deletions(-) diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index c8878afe..dab012d3 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -168,7 +168,7 @@ The `redis` backend's settings, read only when `cache.backend` is `redis`. It is | `cache.redis.compress_min_bytes` | `WH_CACHE_REDIS_COMPRESS_MIN_BYTES` | `1024` | Results at least this large are zstd-compressed when that makes them smaller. `0` never compresses. | | `cache.redis.version_ttl` | `WH_CACHE_REDIS_VERSION_TTL` | `168h` | How long a table's or tenant's version token outlives its last write, so the tokens of dropped tables and removed tenants eventually expire. At least `2s`. An expired token only causes misses. | -**When the server is unreachable or misbehaves, the cache is bypassed; queries are not.** A failure or a timeout makes the lookup a miss and the fill a no-op. Five in a row, one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), or one refusing the credentials (`WRONGPASS`, `NOAUTH` — what an already-open connection meets once the password is rotated; logged at `ERROR`, once per opening and once per refused probe) open a circuit breaker that skips the server entirely until a probe write, every 5 s, succeeds. `NOPERM` does not open it: it names one key or command an ACL user cannot use, not every operation, so it should not bypass the cache for every tenant. Queries then go straight to ClickHouse, still coalesced per instance by `singleflight`. An invalidation the server did not take is kept and retried until it lands, and until then the instance that owes it bypasses the lookups it would orphan. `/readyz` does not depend on the cache. +**When the server is unreachable or misbehaves, the cache is bypassed; queries are not.** A failure or a timeout makes the lookup a miss and the fill a no-op. Five in a row, one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), or one refusing the credentials (`WRONGPASS`, `NOAUTH` — what an already-open connection meets once the password is rotated) open a circuit breaker that skips the server entirely until a probe write, every 5 s, succeeds within `timeout`. Each opening is logged once, at `ERROR` for the credentials and `WARN` otherwise; a failed probe opens it again, so a server that stays down logs every 5 s. `NOPERM` does not open it: it names one key or command an ACL user cannot use, not every operation, so it should not bypass the cache for every tenant. Queries then go straight to ClickHouse, still coalesced per instance by `singleflight`. An invalidation the server did not take is kept and retried until it lands, and until then the instance that owes it bypasses the lookups it would orphan. `/readyz` does not depend on the cache. **Boot does not wait for the server.** A malformed block (an address without a port, `mode: cluster` with `db` other than `0`, an unreadable or unparsable TLS file) refuses boot. A server that cannot be reached, or that refuses the credentials, does not: the process boots with the cache bypassed and keeps reconnecting, with backoff up to 30 s. A rejected credential (`WRONGPASS`, `NOAUTH`, or `NOPERM` for an ACL user missing a connection command) is logged at `ERROR` on every attempt; any other failure at `WARN`. This is deliberate: a rotated Redis password must not crash-loop every instance at once. Watch `wavehouse_cache_breaker_open`, which reads `1` while the cache is bypassed, including before the first connection. diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 0ba84aa5..1d0c5ff1 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -443,7 +443,7 @@ The query-result cache is the layer that can be shared today. With the default ` **Coalescing stays per instance.** `singleflight` collapses identical concurrent queries within each instance, so a cold hot query costs at most one ClickHouse query per instance, not one per request. -**Metrics** (meter `wavehouse-cache`, every series labeled `backend="redis"`, no tenant label): `wavehouse_cache_lookups_total{result}` (`hit`, `miss`, `stale`, `bypass`, `error`), `wavehouse_cache_op_duration_seconds{op}` (`lookup`, `set`, `invalidate`), `wavehouse_cache_breaker_open` (1 while the cache is bypassed), `wavehouse_cache_invalidations_total{result}` (`ok` counts every bump that lands, retried ones included, and `deferred` each time one is put off, a repeat of one already owed included — the two overlap, not a split), `wavehouse_cache_invalidations_pending`, `wavehouse_cache_value_bytes`, `wavehouse_cache_oversize_total` and `wavehouse_cache_set_failures_total{reason}` (`oom`, `timeout`, `other`). Two signals are worth alerting on: `wavehouse_cache_breaker_open` at 1, or `wavehouse_cache_invalidations_pending` above 0, for more than a few minutes. +**Metrics** (meter `wavehouse-cache`, every series labeled `backend="redis"`, no tenant label): `wavehouse_cache_lookups_total{result}` (`hit`, `miss`, `stale`, `bypass`, `error`), `wavehouse_cache_op_duration_seconds{op}` (`lookup`, `set`, `invalidate`), `wavehouse_cache_breaker_open` (1 while the cache is bypassed), `wavehouse_cache_invalidations_total{result}` (`ok` counts every bump that lands, retried ones included, and `deferred` each bump an invalidation could not deliver when made, a repeat of one already owed included; a failed retry is not counted again — the two overlap, not a split), `wavehouse_cache_invalidations_pending`, `wavehouse_cache_value_bytes`, `wavehouse_cache_oversize_total` and `wavehouse_cache_set_failures_total{reason}` (`oom`, `timeout`, `other`). Two signals are worth alerting on: `wavehouse_cache_breaker_open` at 1, or `wavehouse_cache_invalidations_pending` above 0, for more than a few minutes. For local development, `docker compose -f deployments/compose/dependencies.yaml --profile redis up -d` starts a Redis on `localhost:6379` with no persistence. diff --git a/internal/config/cache_redis.go b/internal/config/cache_redis.go index be065266..45600667 100644 --- a/internal/config/cache_redis.go +++ b/internal/config/cache_redis.go @@ -23,8 +23,9 @@ const ( // maxRedisTimeout caps cache.redis.timeout and cache.redis.dial_timeout. // Boot and the cache's Close each wait out a dial in flight: a connect and // a handshake, each bounded by dial_timeout, then for a cluster a topology -// read bounded by the larger of the two. At the caps that is at most 3s, -// inside the 5s budget Close shares with the stores released after it. +// read bounded by the larger of the two. At the caps that is at most 3s; +// Close then spends up to 1s delivering owed invalidations, so at most 4s, +// inside the 5s budget it shares with the stores released after it. const maxRedisTimeout = time.Second // CacheRedisConfig configures cache.backend=redis: one Redis-compatible server