From ef93154b6528054466744df0eec2f26bf1575002 Mon Sep 17 00:00:00 2001 From: taitelee Date: Thu, 24 Sep 2026 17:49:08 -0400 Subject: [PATCH 01/79] feat(mq): give every tenant a queue of its own --- AGENTS.md | 8 +- CHANGELOG.md | 8 +- clients/ts/src/dlq.ts | 18 +- clients/ts/src/namespaces.test.ts | 19 + clients/ts/src/types.ts | 4 +- docs/src/content/docs/api.md | 9 +- docs/src/content/docs/architecture.md | 16 +- docs/src/content/docs/configuration.mdx | 2 +- docs/src/content/docs/deployment.md | 18 +- docs/src/content/docs/durability.md | 6 +- docs/src/content/docs/ingest-pipeline.md | 26 +- docs/src/content/docs/sdk/admin.md | 7 + docs/src/content/docs/sdk/reference.md | 4 +- docs/src/content/docs/settings-directory.mdx | 16 +- docs/src/content/docs/why-wavehouse.md | 8 +- internal/api/dlq.go | 29 +- internal/api/dlq_test.go | 138 ++-- internal/api/ingest.go | 2 +- internal/api/router_test.go | 5 +- internal/app/app.go | 9 +- internal/app/app_test.go | 56 +- internal/app/wire.go | 117 +-- internal/ingest/sweeper.go | 41 +- internal/ingest/sweeper_test.go | 31 +- internal/ingest/worker.go | 19 +- internal/ingest/worker_test.go | 28 +- internal/mq/embedded.go | 756 ++++++++++++++----- internal/mq/embedded_test.go | 649 ++++++++++++---- internal/mq/mq.go | 130 ++-- internal/mq/subject.go | 76 +- internal/mq/subject_test.go | 52 +- internal/settings/settings.go | 23 +- internal/settings/store.go | 2 +- internal/stream/subscriber.go | 20 +- internal/testutil/mocks.go | 16 +- internal/testutil/testutil.go | 19 + 36 files changed, 1679 insertions(+), 708 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 59f08ef36..165957219 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -29,7 +29,7 @@ One binary: Eighteen internal packages under `internal/` (plus `internal/testutil/` for shared test helpers): - **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers -- **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive`/`longestGapWindow` for the two settings folded over every tenant served, and `defaultSetting`/`onDefaultAdopt` for the one resource a process still has one of, the MQ, which follows tenant `0`; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, and the same hook's `Hub.Prune` ends the open streams of a tenant no longer served), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it +- **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` and the `mq.max_bytes_gb` reconcile handing the MQ each served tenant's own gap window and byte budget, and `defaultSetting`/`onDefaultAdopt` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, and the same hook's `Hub.Prune` ends the open streams of a tenant no longer served), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it - **`auth/`** — JWT auth middleware: HMAC **or** JWKS verification with `alg` pinned to the active verifier, role extraction from a configurable claim path; always runs, never rejects (bad token → empty role + stashed reason). One verifier per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 9): `Authenticator` keys them by `tenant.ID` — the request store's `settings.Store.Tenant()`, through an injected `TenantSource`; `tenant.Default` on the tenant-exempt routes — built from each tenant's `auth` block by `Reconfigure`, dropped by `Prune` once the tenant stops being served (rejected or removed), released by `Close`; the secrets (`Config`) are boot-level and shared. A JWKS key set is fetched off the boot and reload paths: until one has been stored the verifier is pending and a token-bearing request gets `503` + `Retry-After` from `api.refuseUnverifiable` (`auth.ErrVerifierPending`), never a `default_role` evaluation; refresh is library-managed (Eric, 2026-09-22), response capped at 1 MiB; the operator key's admin role is the request tenant's - **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, `...
.` for a namespace — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version that leads every namespace key of one tenant, orphaning every cached query keyed by its tables in one step (a pipe result names no table and keeps its TTL, [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the table-keyed cache (the structured-query results) of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config @@ -38,14 +38,14 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`dedupe/`** — `Deduplicator` interface → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) -- **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, a topic without one is refused, and a pre-tenant subject reads as tenant `0`'s) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds the byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload) and `Stats` (the system gauges' source). Subjects, prefixes, wildcards, token encoding, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`); `internal/app` constructs it and hands everything else a `mq.Broker` +- **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, token encoding, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`); `internal/app` constructs it and hands everything else a `mq.Broker` - **`observability/`** — OpenTelemetry pipeline: `InitProvider` wires trace/metric/log providers via OTLP gRPC (each signal independently gated). A top-level `Prometheus` config block drives an optional `/metrics` scrape endpoint that runs independently of OTLP push — standalone (Alloy/Mimir scrape, no collector), alongside OTLP, or off. `NewLogger` produces a slog handler that fans out to stdout AND OTLP (stdout always 100%, OTLP sample-rate-aware). `TraceHandler` injects trace_id/span_id from active spans. `tracer.go` provides W3C trace context propagation over message headers (`InjectHeaders`/`ExtractHeaders` on a plain header map; `internal/mq` injects on every publish and extracts on the `Subscribe` path, so this package never sees a NATS type). - **`pipes/`** — Named query pipes: `NamedQuery` type + `BindParams` + `Source` (read per request; `settings.Store` in production, `Static(q...)` in tests) - **`policy/`** — Hasura-style access control, **role-first**: `TablePolicy` is `map[string]RolePermissions`, and a role's grant splits by operation into `SelectPermissions` (columns, row `filter`, aggregations, the `max_*` limits) and `InsertPermissions` (columns, `check`) — so a field only one side honors does not exist on the other. `Evaluate()` resolves ONE operation and leaves the other side **nil** (`Select *ResolvedSelect` / `Insert *ResolvedInsert`), which every accessor fails closed on — nil is "not resolved", distinct from an empty side, which is "unrestricted" (what the admin return builds). Claim templating (`{{ jwt.claim.path }}`) resolves during that call. Policies come from `Source`, a `func() *Policy` read per call (`settings.Store.Policy` in production, `Static(p)` in tests) - **`query/`** — Structured query AST types + SQL builder with schema validation, structural policy predicate/limit emission, timestamp bucketing - **`settings/`** — the settings directory, in either shape ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)): flat (the four files: tenant `0` alone) or nested (one folder per tenant, never mixed). `Validate` detects the shape and checks it — `ValidateDir` per directory (strict JSON, per-file rules, cross-file role references), folder names against `tenant.Parse`, a nested finding's `File` led by its folder; `Store` is a passive holder (one tenant's adopted snapshot, typed accessors read per call); `Registry` (tenant id → `Store`) owns `Open`, the serialized `Reload`/`ReloadTenant`, the `AfterAdopt` hooks, and the fsnotify `Watch` (flat only). Flat refuses an invalid directory at boot and keeps the previous snapshot on a rejected reload; nested fails closed per tenant (a rejected folder stops being served, the rest carry on, a whole-tree reload mirrors the folders, down to none, and a finding about the root itself rejects the reload whole). Plus the embedded (`go:embed`) seed `wavehouse bootstrap` writes - **`stream/`** — SSE fan-out: rows travel POSITIONALLY, so each connection is told its projected column list in an `event: schema` frame before its first row and again on drift — **not** guaranteed after a gap-fill across a column change, which can leave a connection reading live rows against a stale list until it reconnects ([#543](https://github.com/Wave-RF/WaveHouse/issues/543)) — (tracked per connection; replay tracks its own). The event `Hub` (registers subscribers by `(mq.Topic, role)` — one tenant's table — and evaluates each event under its own tenant's policy and schema registry; `Prune` evicts the subscribers of every tenant a reload stopped serving; `Broadcast` projects + serializes each event once per role, the #294 delivery hot path — a role carrying a row-level `filter` keeps the shared projection but delivers per subscriber, each subscriber's claims evaluated against the row, #319), `Subscriber` (per-connection outbound `Frame` queue, `Send`/`Frames`; claims fixed at construction, immutable; `Evict` asks its handler to end the stream), the `Bucket` fan-out set (`subscriberSet`, one per `(topic, role)`), the `Heartbeater` keepalive wheel, and `Metrics` (the `wavehouse_sse_*` stream instruments) -- **`tenant/`** — the tenant identifier ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)): `ID` (a validated string), `Parse` (letters, digits, `_`, `-`; ≤ 64 bytes — safe as a folder name and as an MQ subject token), `Default` (`"0"`), and `Header` (`X-Tenant-ID`). Imports nothing from the rest of the repo. `api.TenantMW` resolves the header against `settings.Registry` before auth on every `/v1` route outside `/v1/ops/*` (`400` malformed, `404` unknown, a bare `503` for a nested tenant whose folder was rejected) and puts the resolved `*settings.Store` in the request context; the ops routes that address one tenant (`GET /v1/ops/pipes[/{name}]`, `POST /v1/ops/settings/reload`, `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh`, `POST /v1/ops/query`) take a strictly parsed `?tenant=` instead; handlers read it once (`api.StoreFromContext`) and pass it down as an argument, and nothing below a handler reads context. The stream hub and the ingest worker read each message's tenant off its `mq.Topic` and their getters take it; the sweeper folds over the tenants served (`longestGapWindow`); each served tenant has a schema registry of its own (story 6) +- **`tenant/`** — the tenant identifier ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)): `ID` (a validated string), `Parse` (letters, digits, `_`, `-`; ≤ 64 bytes — safe as a folder name and as an MQ subject token), `Default` (`"0"`), and `Header` (`X-Tenant-ID`). Imports nothing from the rest of the repo. `api.TenantMW` resolves the header against `settings.Registry` before auth on every `/v1` route outside `/v1/ops/*` (`400` malformed, `404` unknown, a bare `503` for a nested tenant whose folder was rejected) and puts the resolved `*settings.Store` in the request context; the ops routes that address one tenant (`GET /v1/ops/pipes[/{name}]`, `POST /v1/ops/settings/reload`, `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh`, `POST /v1/ops/query`, `GET /v1/ops/dlq/stats`) take a strictly parsed `?tenant=` instead; handlers read it once (`api.StoreFromContext`) and pass it down as an argument, and nothing below a handler reads context. The stream hub and the ingest worker read each message's tenant off its `mq.Topic` and their getters take it; the sweeper hands the MQ each served tenant's own gap window (`gapWindows`); each served tenant has a schema registry of its own (story 6) ## Key Design Decisions @@ -56,7 +56,7 @@ The invariant index — what must stay true. Full narrative and rationale live i 3. **Schema-driven ingest** — `POST /v1/ingest?table={table}` takes flat JSON, validated against the discovered schema (unknown fields rejected, types/nullability enforced). No envelope. The **declared `Content-Type` chooses the format and the bytes never do** (arity within the JSON family is still the body's): no declaration, one whose **media type** is unsupported or unparseable, a comma-bearing value that, as a whole, does not parse as one media type, or repeated lines that **disagree**, is a `415` decided *before* the body is read. A malformed *parameter* on a comma-free line never costs the request (`; charset=a; charset=b` still reads as its media type), and repeated lines are accepted only when they all resolve to the same **supported** format — two agreeing `text/csv` lines are still a `415`. A body declared NDJSON stays NDJSON whatever its bytes, so a bad line is a per-record error rather than a silent re-framing; the reverse (NDJSON sent as `application/json`) is deliberately **not** caught — record one, `200`, the rest ignored ([#561](https://github.com/Wave-RF/WaveHouse/issues/561)). Fail-closed — preserve it when touching `internal/api`. 4. **Async ingestion** — ingest returns 200 after optional dedup + MQ publish; ClickHouse writes happen later via `StartIngestWorker`. NATS full → 503 + Retry-After. 5. **Per-tenant-table batching** — the worker groups events by tenant table (the tenant read off each message's `mq.Topic`), so one INSERT never mixes tenants and a batch invalidates its own tenant's cache namespaces; then it splits each batch by column list (`groupByColumns`), emitting one `INSERT INTO … (cols) FORMAT JSONCompactEachRow` per distinct list so a schema change mid-stream can't corrupt a statement. Each tenant table's batch is independent. -6. **Dead Letter Queue** — failed batch inserts publish to `WAVEHOUSE_DLQ` (`dlq..
`), gated per table by the tenant's `dlq.enabled` in the settings directory's `config.json` (hot-reloadable; off = leave the row unacked for redelivery). No silent data loss on the insert path. The one drop is an envelope the worker cannot READ (malformed JSON, an unknown **or absent** `format` — a pre-v2 envelope carries none — or columns and row that don't pair): it is poison by construction, so with the DLQ off it is acked-and-dropped rather than redelivered forever — logged at `ERROR` and counted by `wavehouse_ingest_poison_total` under `disposition="dropped"`. With the DLQ on it is parked like any other failure, and counted under `disposition="parked"`. +6. **Dead Letter Queue** — failed batch inserts publish to the tenant's own dead-letter queue (`dlq..
`), gated per table by the tenant's `dlq.enabled` in the settings directory's `config.json` (hot-reloadable; off = leave the row unacked for redelivery). No silent data loss on the insert path. The one drop is an envelope the worker cannot READ (malformed JSON, an unknown **or absent** `format` — a pre-v2 envelope carries none — or columns and row that don't pair): it is poison by construction, so with the DLQ off it is acked-and-dropped rather than redelivered forever — logged at `ERROR` and counted by `wavehouse_ingest_poison_total` under `disposition="dropped"`. With the DLQ on it is parked like any other failure, and counted under `disposition="parked"`. 7. **Auth: always on, fail-loud, decoupled from authz (security)** — the JWT middleware always runs (no `auth.enabled`/`dev_mode` flag); it verifies with HMAC **or** JWKS (not both), with accepted `alg` pinned to the active verifier and checked before any key is used (rejects `alg:none` and cross-family confusion). No/invalid/expired token → empty role → policy `default_role`, with the bad-token reason stashed so a denying gate returns a loud `401`, not a bare `403`; the one token outcome that never reaches `default_role` is a verifier still fetching its JWKS (`auth.ErrVerifierPending` → `503` + `Retry-After`, `api.refuseUnverifiable`). Elevated access needs a valid granted role. **Sanctioned exception:** a configured non-JWT operator key (`auth.operator_key`; presented via `Authorization: Operator ` or the `X-Operator-Key` alias) deliberately couples authN+authZ — a constant-time match authorizes a full-access platform operator (stamps the admin role plus an operator bit) independent of the verifier (see #11). Detail: architecture.md § `api/` + `internal/auth`; see also #11, §Security Considerations. 8. **Optional dedup, per tenant** — opt-in via `dedupe.enabled` in the settings directory's `config.json` (hot-reloadable: a reload opens or closes that tenant's store via `dedupe.Managed`, one per tenant in `dedupe.Stores`, each a share of the one Pebble instance whose keys lead with the tenant; the ingest handler picks the store off the request's `settings.Store`, so one tenant's seen ids are never another's); `dedupe.id_field` there selects the JSON key, overridable per table. 9. **Singleflight** — the cached read handlers coalesce concurrent misses (`x/sync/singleflight`) under the tenant-led cache key to prevent cache stampede, per tenant. diff --git a/CHANGELOG.md b/CHANGELOG.md index 5bf580196..23c0c7152 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked to hold the ack floor that the one shared stream's purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. +- **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. @@ -24,7 +24,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **Schema discovery captures each table's DDL, its columns' ordinals and default expressions, and the server version** (`internal/discovery/discovery.go`, `internal/testutil/testutil.go`): `Column` gains `DefaultExpression` and `Position` (both from a widened `system.columns` select), `TableSchema` gains `DDL` from `system.tables.create_table_query`, and `SchemaRegistry` gains `ServerVersion()` from a `SELECT version()` probe next to the existing `SELECT timezone()`. Groundwork for the native type layer, captured on the same refresh as the columns so a stale version cannot outlive the schemas it describes. That is a publication guarantee, not a same-server one: `chconn.Manager` resolves the connection per call, so a reload changing `clickhouse.addr` mid-refresh can still pair a version from one server with schemas from another — narrow, and self-correcting on the next refresh. `DDL` is `json:"-"` and does **not** appear in `/v1/ops/schema`: that endpoint marshals `TableSchema` straight to the client, and an external-engine table (S3, MySQL, PostgreSQL, Kafka) renders its wiring there unconditionally — endpoint, bucket or host, database, username, S3 access key id. ClickHouse masks the password itself as `[HIDDEN]` from ~23.9 (verified on 26.7.3), so the exposure is the topology rather than the secret — except on an older server, or one with `display_secrets_in_show_and_select` enabled. `position` and `default_expression` are additive fields in the response. A table listed in `system.tables` with no `system.columns` rows is skipped rather than published column-less, and both new queries fail the refresh on error exactly as `timezone()` and `system.columns` do — callers keep the prior cache and retry. -- **Settings-directory hot reload — boot loading, three reload triggers, and the config-key migration** (`internal/settings/` (new: `store.go`, `watch.go`, + tests), `internal/api/settings.go` (new, + tests), `internal/api/{router,ingest,structured_query}.go`, `internal/discovery/discovery.go`, `internal/config/config.go`, `cmd/wavehouse/main.go`, `config.yaml`, `deployments/compose/standalone.yaml`, `docs/src/content/docs/settings-directory.mdx` (new — the hot-reloadable half of configuration gets its own page; `configuration.mdx` is boot config only); closes the loop [#500](https://github.com/Wave-RF/WaveHouse/pull/500) opened, tracked by [#48](https://github.com/Wave-RF/WaveHouse/issues/48)): the server now *consumes* the settings directory instead of only validating it. `settings.Store` owns the adopted snapshot: `settings.dir` / `WH_SETTINGS_DIR` is now **required**, boot validates and adopts the directory (missing or invalid refuses to start); a running instance then re-validates and re-adopts on any of three triggers — a **directory watch** (fsnotify on the directory, not the files, so atomic-writer replaces and Kubernetes ConfigMap symlink swaps aren't lost; bursts debounce into one reload), **`SIGHUP`**, and **`POST /v1/ops/settings/reload`** (admin-gated; returns `{"adopted", "findings"}`, `200` adopted / `422` rejected) — all funneling through one serialized reload path. A reload that fails validation keeps the previous good snapshot (an operator mid-edit degrades to a log line, never a broken server); warnings don't block adoption, matching `wavehouse validate`. The tenant tunables **migrate out of boot config** into the directory's `config.json`: `dedupe.id_field` / `dedupe.require_id` (now with the per-table overrides under `dedupe.tables` that [#222](https://github.com/Wave-RF/WaveHouse/issues/222) asked for, resolved per record through the table → global cascade in one atomic snapshot read, so a reload lands at a record boundary and never mixes documents within one record), `query.default_max_rows` and `query.timestamp_bucket_seconds` (read per query), `schema.refresh_interval` (re-read after each tick, so a change applies from the next cycle), `stream.keepalive_interval` / `stream.keepalive_buckets` (a reload calls the new `Heartbeater.Reconfigure`, which rebuilds the keepalive wheel in place with every live subscriber carried over and re-times the running ticker) and `stream.gap_window_minutes` (the sweeper re-reads it every sweep), `mq.max_bytes_gb` (an after-adopt hook updates the `WAVEHOUSE` and `WAVEHOUSE_DLQ` stream limits in place via `EmbeddedNATS.Resize` — shrinking below the buffered size backpressures until the worker drains, nothing is dropped), `dlq.enabled` with per-table overrides under `dlq.tables` (resolved by the ingest worker at the moment a poison row is isolated: on → park it on `WAVEHOUSE_DLQ` and ack; off → leave it unacked for redelivery, never dropped; the DLQ stream and `GET /v1/ops/dlq/stats` now always exist, so the switch is purely behavioral), the **ClickHouse wiring** (`clickhouse.addr` / `http_port` / `http_scheme` / `database` / `username` / `query_timeout`: the new `chconn.Manager` is the one `driver.Conn` every consumer holds and swaps the connection behind it on reload — unconditionally, since the adopted settings are the authority and reachability already surfaces through schema discovery and `/readyz`; the replaced one closes after a `query_timeout` grace; the ingest worker, raw-SQL proxy, and schema registry read the HTTP target, timeout, and database per call), the **auth verifier wiring** (`auth.jwks_url` / `auth.role_claim`: the new `auth.Authenticator` swaps a whole verifier — key source plus its pinned algorithm allowlist — atomically per reload, unconditionally, so an unreachable JWKS fails closed until it can be fetched; `auth.Middleware` is gone — `Authenticator` is the one constructor), and the CORS allowlist (`cors.allowed_origins`, resolved per request). The corresponding YAML/env keys are **removed**: `server.cors_allowed_origins`, `query.default_max_rows`, `schema.refresh_interval`, `dedupe.enabled`, `dedupe.id_field`, `dedupe.require_id`, `stream.keepalive_interval`, `stream.keepalive_buckets`, `mq.gap_window_minutes`, `cache.timestamp_bucket_seconds`, `mq.max_bytes_gb`, `dlq.enabled`, `clickhouse.addr`, `clickhouse.http_port`, `clickhouse.http_scheme`, `clickhouse.database`, `clickhouse.username`, `clickhouse.query_timeout`, `auth.jwks_url`, `auth.role_claim` (and `WH_SERVER_CORS_ALLOWED_ORIGINS`, `WH_QUERY_DEFAULT_MAX_ROWS`, `WH_SCHEMA_REFRESH_INTERVAL`, `WH_DEDUPE_ENABLED`, `WH_DEDUPE_ID_FIELD`, `WH_DEDUPE_REQUIRE_ID`, `WH_STREAM_KEEPALIVE_INTERVAL`, `WH_STREAM_KEEPALIVE_BUCKETS`, `WH_MQ_GAP_WINDOW_MINUTES`, `WH_CACHE_TIMESTAMP_BUCKET_SECONDS`, `WH_MQ_MAX_BYTES_GB`, `WH_DLQ_ENABLED`, `WH_CH_ADDR`, `WH_CH_HTTP_PORT`, `WH_CH_HTTP_SCHEME`, `WH_CH_DATABASE`, `WH_CH_USERNAME`, `WH_CH_QUERY_TIMEOUT`, `WH_AUTH_JWKS_URL`, `WH_AUTH_ROLE_CLAIM`); the secrets — `clickhouse.password`, `auth.jwt_secret`, `auth.operator_key` — stay boot config on purpose (never in a tracked JSON file; combined with the adopted wiring on every reconnect, rotating one is a restart), and boot config is now **strict**: `config.Load` re-reads the YAML against the struct's tags and refuses to start naming every undeclared key, so a `dlq:` or `clickhouse: addr:` left behind can't be read, ignored, and believed; the binary carries **no compiled defaults** — every `config.json` key is required (validation names each missing one), so the adopted snapshot is what the files say, and once adopted it outlives its files (a deleted file or vanished directory is just a rejected reload). Defaults live in one checked-in seed directory (`internal/settings/seed/`, `go:embed`ded): the new **`wavehouse bootstrap [dir]`** writes it (refusing a non-empty directory, the `initdb` contract; the directory resolves exactly as it does for `validate` — the argument, else `WH_SETTINGS_DIR`, usage error with neither — so the two commands are interchangeable on one path and a bare `bootstrap` inside the container images seeds `/app/settings`), the dev `config.yaml` points at a gitignored `./settings` that `make dev` seeds from it, and the e2e fixture ships a copy. The container images ship **no** settings directory: `WH_SETTINGS_DIR` is preset to `/app/settings`, the operator mounts a directory there (`standalone.yaml` bind-mounts the checked-in `deployments/compose/settings/`), and a missing mount refuses to boot rather than running on defaults nobody chose. `dedupe.enabled` moves too: the new `dedupe.Managed` wraps the Pebble store and a `Store.AfterAdopt` hook opens or closes it after every adoption, so flipping the switch is a reload, not a restart (seen ids persist across an off/on cycle; a failed open on reload is logged and ingest fails closed with `500` until the next reload, since the files asked for dedupe — at boot it still refuses to start; a record caught in the instant of the flip is published un-deduped and counted by `wavehouse_ingest_dedupe_disabled_total` rather than failed, and the hook is registered before the boot apply so a reload can never leave the settings and the store out of step). The watcher reloads once as soon as its watch exists, closing the gap between the boot read and the watch — an edit landing in between (a ConfigMap update during a rolling restart) is adopted, not silently missed. `dedupe.enabled` / `WH_DEDUPE_ENABLED` are removed from boot config alongside the other keys. What stays in boot config is only what cannot change under a running process — resource sizing (`data_dir`, `cache.l1_max_cost`), the listeners, the observability exporters — and the secrets. The compose stack now bind-mounts a checked-in `deployments/compose/settings/` (the seed with `clickhouse.addr` pointed at the `clickhouse` service) instead of a volume seeded with `bootstrap`, so the quickstart is `up -d` again; the e2e orchestrator copies the fixture settings per run and patches the testcontainer's ClickHouse ports into `config.json`, since that wiring no longer has an env override. Every after-adopt hook (dedupe, keepalive wheel) is registered before the reload triggers start, so the watcher's first reload can never be missed by a hook. Consumers take functions, not values (`IngestHandler.DedupeSettings`, the structured-query handler's `defaultMaxRows` / `bucketSecs func() int`, the ingest worker's `dlqEnabled func(table) bool`, the sweeper's `gapWindow func() time.Duration`, `corsMiddleware`'s origins getter, `SchemaRegistry`'s database and refresh-interval sources, the query handlers' timeout sources), so `internal/api` stays testable without materializing settings directories. The settings directory is also the **runtime authority for access control and named pipes** (`internal/settings/store.go`, `internal/policy/source.go` (new), `internal/pipes/pipes.go`, `internal/api/{policy,pipes,router}.go`, `internal/stream/hub.go`, `internal/auth/auth.go`, `cmd/wavehouse/main.go`, `Makefile`, `deployments/compose/settings/{policies,roles}.json`, `clients/ts/src/settings.ts` (new); closes [#229](https://github.com/Wave-RF/WaveHouse/issues/229), [#33](https://github.com/Wave-RF/WaveHouse/issues/33), [#461](https://github.com/Wave-RF/WaveHouse/issues/461), [#514](https://github.com/Wave-RF/WaveHouse/issues/514), [#460](https://github.com/Wave-RF/WaveHouse/issues/460), [#363](https://github.com/Wave-RF/WaveHouse/issues/363); advances [#48](https://github.com/Wave-RF/WaveHouse/issues/48) and [#214](https://github.com/Wave-RF/WaveHouse/issues/214)): `roles.json`, `policies.json`, and `pipes.json` are adopted with `config.json` as one snapshot and re-adopted on the same three triggers, and **files are the only write path** — standalone, the operator edits them on the host; on WaveHouse Cloud the control plane writes them — so there is no stored copy that can skip validation: every adoption runs the current rules (strict decode rejecting unknown and duplicate keys, the full policy validation including the claim-template grammar, pipe name/SQL/parameter-type rules, and the cross-file check that every role a grant or `allowed_roles` names is declared in `roles.json`), and a rejected edit keeps the previous good policy and pipes in effect. `policies.json` is one policy document (`{}` = no policy, adopted fail-closed with a warning); `pipes.json` carries full definitions (`allowed_roles`, `parameters`, `description`), so a file-defined pipe is no longer admin-only by construction. Consumers read the adopted snapshot per request through `policy.Source` (a `func() *policy.Policy`; `settings.Store.Policy` in production, `policy.Static(p)` in tests) and `pipes.Source` (`settings.Store`; `pipes.Static(q...)` in tests), so a reload applies to the very next request, including the SSE hub's per-event policy read. `GET /v1/ops/policy`, `POST /v1/ops/policy/validate`, `GET /v1/ops/pipes`, `GET /v1/ops/pipes/{name}`, and pipe execution are unchanged; the operator key still passes the `/v1/ops/*` gate under no policy, now as the break-glass that inspects the policy and triggers `POST /v1/ops/settings/reload` after `policies.json` is fixed. The SDK gains `wh.settings.reload()` (`POST /v1/ops/settings/reload`, returning `{ adopted, findings }`). The compose stack's trial `public` policy moves into the bind-mounted `deployments/compose/settings/policies.json` + `roles.json`, and `make dev` copies the same two files into its seeded `./settings` so a fresh dev server works tokenless. **Removed** — the write endpoints `PUT /v1/ops/policy`, `PUT /v1/ops/pipes/{name}`, and `DELETE /v1/ops/pipes/{name}`; the NATS KV buckets `WAVEHOUSE_POLICY` and `WAVEHOUSE_PIPES` and their KV Watch sync (`internal/policy/store.go`, the pipes KV store); the boot-config keys `policy.file_path` / `WH_POLICY_FILE_PATH` and `pipes.dir` / `WH_PIPES_DIR` (a leftover `policy:` or `pipes:` YAML block now refuses boot by name, like the other moved keys) and the `.sql`-directory pipes bootstrap; `deployments/compose/dev-policy.yaml`; the SDK methods `wh.policy.set`, `wh.pipes.set`, and `wh.pipes.delete`; and the test helpers `policy.NewMemoryStore`, `pipes.NewMemoryStore`, and `testutil/natsjs.go`. +- **Settings-directory hot reload — boot loading, three reload triggers, and the config-key migration** (`internal/settings/` (new: `store.go`, `watch.go`, + tests), `internal/api/settings.go` (new, + tests), `internal/api/{router,ingest,structured_query}.go`, `internal/discovery/discovery.go`, `internal/config/config.go`, `cmd/wavehouse/main.go`, `config.yaml`, `deployments/compose/standalone.yaml`, `docs/src/content/docs/settings-directory.mdx` (new — the hot-reloadable half of configuration gets its own page; `configuration.mdx` is boot config only); closes the loop [#500](https://github.com/Wave-RF/WaveHouse/pull/500) opened, tracked by [#48](https://github.com/Wave-RF/WaveHouse/issues/48)): the server now *consumes* the settings directory instead of only validating it. `settings.Store` owns the adopted snapshot: `settings.dir` / `WH_SETTINGS_DIR` is now **required**, boot validates and adopts the directory (missing or invalid refuses to start); a running instance then re-validates and re-adopts on any of three triggers — a **directory watch** (fsnotify on the directory, not the files, so atomic-writer replaces and Kubernetes ConfigMap symlink swaps aren't lost; bursts debounce into one reload), **`SIGHUP`**, and **`POST /v1/ops/settings/reload`** (admin-gated; returns `{"adopted", "findings"}`, `200` adopted / `422` rejected) — all funneling through one serialized reload path. A reload that fails validation keeps the previous good snapshot (an operator mid-edit degrades to a log line, never a broken server); warnings don't block adoption, matching `wavehouse validate`. The tenant tunables **migrate out of boot config** into the directory's `config.json`: `dedupe.id_field` / `dedupe.require_id` (now with the per-table overrides under `dedupe.tables` that [#222](https://github.com/Wave-RF/WaveHouse/issues/222) asked for, resolved per record through the table → global cascade in one atomic snapshot read, so a reload lands at a record boundary and never mixes documents within one record), `query.default_max_rows` and `query.timestamp_bucket_seconds` (read per query), `schema.refresh_interval` (re-read after each tick, so a change applies from the next cycle), `stream.keepalive_interval` / `stream.keepalive_buckets` (a reload calls the new `Heartbeater.Reconfigure`, which rebuilds the keepalive wheel in place with every live subscriber carried over and re-times the running ticker) and `stream.gap_window_minutes` (the sweeper re-reads it every sweep), `mq.max_bytes_gb` (an after-adopt hook updates the tenant's ingest and dead-letter stream limits in place via `mq.Broker.SetMaxBytes` — shrinking below the buffered size backpressures until the worker drains, nothing is dropped), `dlq.enabled` with per-table overrides under `dlq.tables` (resolved by the ingest worker at the moment a poison row is isolated: on → park it on the tenant's dead-letter stream and ack; off → leave it unacked for redelivery, never dropped; a served tenant's DLQ stream and `GET /v1/ops/dlq/stats` always exist, so the switch is purely behavioral), the **ClickHouse wiring** (`clickhouse.addr` / `http_port` / `http_scheme` / `database` / `username` / `query_timeout`: the new `chconn.Manager` is the one `driver.Conn` every consumer holds and swaps the connection behind it on reload — unconditionally, since the adopted settings are the authority and reachability already surfaces through schema discovery and `/readyz`; the replaced one closes after a `query_timeout` grace; the ingest worker, raw-SQL proxy, and schema registry read the HTTP target, timeout, and database per call), the **auth verifier wiring** (`auth.jwks_url` / `auth.role_claim`: the new `auth.Authenticator` swaps a whole verifier — key source plus its pinned algorithm allowlist — atomically per reload, unconditionally, so an unreachable JWKS fails closed until it can be fetched; `auth.Middleware` is gone — `Authenticator` is the one constructor), and the CORS allowlist (`cors.allowed_origins`, resolved per request). The corresponding YAML/env keys are **removed**: `server.cors_allowed_origins`, `query.default_max_rows`, `schema.refresh_interval`, `dedupe.enabled`, `dedupe.id_field`, `dedupe.require_id`, `stream.keepalive_interval`, `stream.keepalive_buckets`, `mq.gap_window_minutes`, `cache.timestamp_bucket_seconds`, `mq.max_bytes_gb`, `dlq.enabled`, `clickhouse.addr`, `clickhouse.http_port`, `clickhouse.http_scheme`, `clickhouse.database`, `clickhouse.username`, `clickhouse.query_timeout`, `auth.jwks_url`, `auth.role_claim` (and `WH_SERVER_CORS_ALLOWED_ORIGINS`, `WH_QUERY_DEFAULT_MAX_ROWS`, `WH_SCHEMA_REFRESH_INTERVAL`, `WH_DEDUPE_ENABLED`, `WH_DEDUPE_ID_FIELD`, `WH_DEDUPE_REQUIRE_ID`, `WH_STREAM_KEEPALIVE_INTERVAL`, `WH_STREAM_KEEPALIVE_BUCKETS`, `WH_MQ_GAP_WINDOW_MINUTES`, `WH_CACHE_TIMESTAMP_BUCKET_SECONDS`, `WH_MQ_MAX_BYTES_GB`, `WH_DLQ_ENABLED`, `WH_CH_ADDR`, `WH_CH_HTTP_PORT`, `WH_CH_HTTP_SCHEME`, `WH_CH_DATABASE`, `WH_CH_USERNAME`, `WH_CH_QUERY_TIMEOUT`, `WH_AUTH_JWKS_URL`, `WH_AUTH_ROLE_CLAIM`); the secrets — `clickhouse.password`, `auth.jwt_secret`, `auth.operator_key` — stay boot config on purpose (never in a tracked JSON file; combined with the adopted wiring on every reconnect, rotating one is a restart), and boot config is now **strict**: `config.Load` re-reads the YAML against the struct's tags and refuses to start naming every undeclared key, so a `dlq:` or `clickhouse: addr:` left behind can't be read, ignored, and believed; the binary carries **no compiled defaults** — every `config.json` key is required (validation names each missing one), so the adopted snapshot is what the files say, and once adopted it outlives its files (a deleted file or vanished directory is just a rejected reload). Defaults live in one checked-in seed directory (`internal/settings/seed/`, `go:embed`ded): the new **`wavehouse bootstrap [dir]`** writes it (refusing a non-empty directory, the `initdb` contract; the directory resolves exactly as it does for `validate` — the argument, else `WH_SETTINGS_DIR`, usage error with neither — so the two commands are interchangeable on one path and a bare `bootstrap` inside the container images seeds `/app/settings`), the dev `config.yaml` points at a gitignored `./settings` that `make dev` seeds from it, and the e2e fixture ships a copy. The container images ship **no** settings directory: `WH_SETTINGS_DIR` is preset to `/app/settings`, the operator mounts a directory there (`standalone.yaml` bind-mounts the checked-in `deployments/compose/settings/`), and a missing mount refuses to boot rather than running on defaults nobody chose. `dedupe.enabled` moves too: the new `dedupe.Managed` wraps the Pebble store and a `Store.AfterAdopt` hook opens or closes it after every adoption, so flipping the switch is a reload, not a restart (seen ids persist across an off/on cycle; a failed open on reload is logged and ingest fails closed with `500` until the next reload, since the files asked for dedupe — at boot it still refuses to start; a record caught in the instant of the flip is published un-deduped and counted by `wavehouse_ingest_dedupe_disabled_total` rather than failed, and the hook is registered before the boot apply so a reload can never leave the settings and the store out of step). The watcher reloads once as soon as its watch exists, closing the gap between the boot read and the watch — an edit landing in between (a ConfigMap update during a rolling restart) is adopted, not silently missed. `dedupe.enabled` / `WH_DEDUPE_ENABLED` are removed from boot config alongside the other keys. What stays in boot config is only what cannot change under a running process — resource sizing (`data_dir`, `cache.l1_max_cost`), the listeners, the observability exporters — and the secrets. The compose stack now bind-mounts a checked-in `deployments/compose/settings/` (the seed with `clickhouse.addr` pointed at the `clickhouse` service) instead of a volume seeded with `bootstrap`, so the quickstart is `up -d` again; the e2e orchestrator copies the fixture settings per run and patches the testcontainer's ClickHouse ports into `config.json`, since that wiring no longer has an env override. Every after-adopt hook (dedupe, keepalive wheel) is registered before the reload triggers start, so the watcher's first reload can never be missed by a hook. Consumers take functions, not values (`IngestHandler.DedupeSettings`, the structured-query handler's `defaultMaxRows` / `bucketSecs func() int`, the ingest worker's `dlqEnabled func(table) bool`, the sweeper's `gapWindow func() time.Duration`, `corsMiddleware`'s origins getter, `SchemaRegistry`'s database and refresh-interval sources, the query handlers' timeout sources), so `internal/api` stays testable without materializing settings directories. The settings directory is also the **runtime authority for access control and named pipes** (`internal/settings/store.go`, `internal/policy/source.go` (new), `internal/pipes/pipes.go`, `internal/api/{policy,pipes,router}.go`, `internal/stream/hub.go`, `internal/auth/auth.go`, `cmd/wavehouse/main.go`, `Makefile`, `deployments/compose/settings/{policies,roles}.json`, `clients/ts/src/settings.ts` (new); closes [#229](https://github.com/Wave-RF/WaveHouse/issues/229), [#33](https://github.com/Wave-RF/WaveHouse/issues/33), [#461](https://github.com/Wave-RF/WaveHouse/issues/461), [#514](https://github.com/Wave-RF/WaveHouse/issues/514), [#460](https://github.com/Wave-RF/WaveHouse/issues/460), [#363](https://github.com/Wave-RF/WaveHouse/issues/363); advances [#48](https://github.com/Wave-RF/WaveHouse/issues/48) and [#214](https://github.com/Wave-RF/WaveHouse/issues/214)): `roles.json`, `policies.json`, and `pipes.json` are adopted with `config.json` as one snapshot and re-adopted on the same three triggers, and **files are the only write path** — standalone, the operator edits them on the host; on WaveHouse Cloud the control plane writes them — so there is no stored copy that can skip validation: every adoption runs the current rules (strict decode rejecting unknown and duplicate keys, the full policy validation including the claim-template grammar, pipe name/SQL/parameter-type rules, and the cross-file check that every role a grant or `allowed_roles` names is declared in `roles.json`), and a rejected edit keeps the previous good policy and pipes in effect. `policies.json` is one policy document (`{}` = no policy, adopted fail-closed with a warning); `pipes.json` carries full definitions (`allowed_roles`, `parameters`, `description`), so a file-defined pipe is no longer admin-only by construction. Consumers read the adopted snapshot per request through `policy.Source` (a `func() *policy.Policy`; `settings.Store.Policy` in production, `policy.Static(p)` in tests) and `pipes.Source` (`settings.Store`; `pipes.Static(q...)` in tests), so a reload applies to the very next request, including the SSE hub's per-event policy read. `GET /v1/ops/policy`, `POST /v1/ops/policy/validate`, `GET /v1/ops/pipes`, `GET /v1/ops/pipes/{name}`, and pipe execution are unchanged; the operator key still passes the `/v1/ops/*` gate under no policy, now as the break-glass that inspects the policy and triggers `POST /v1/ops/settings/reload` after `policies.json` is fixed. The SDK gains `wh.settings.reload()` (`POST /v1/ops/settings/reload`, returning `{ adopted, findings }`). The compose stack's trial `public` policy moves into the bind-mounted `deployments/compose/settings/policies.json` + `roles.json`, and `make dev` copies the same two files into its seeded `./settings` so a fresh dev server works tokenless. **Removed** — the write endpoints `PUT /v1/ops/policy`, `PUT /v1/ops/pipes/{name}`, and `DELETE /v1/ops/pipes/{name}`; the NATS KV buckets `WAVEHOUSE_POLICY` and `WAVEHOUSE_PIPES` and their KV Watch sync (`internal/policy/store.go`, the pipes KV store); the boot-config keys `policy.file_path` / `WH_POLICY_FILE_PATH` and `pipes.dir` / `WH_PIPES_DIR` (a leftover `policy:` or `pipes:` YAML block now refuses boot by name, like the other moved keys) and the `.sql`-directory pipes bootstrap; `deployments/compose/dev-policy.yaml`; the SDK methods `wh.policy.set`, `wh.pipes.set`, and `wh.pipes.delete`; and the test helpers `policy.NewMemoryStore`, `pipes.NewMemoryStore`, and `testutil/natsjs.go`. - **"Was this page helpful?" feedback widget on every docs page** (`docs/src/components/PageFeedback.astro` (new), `docs/src/components/Footer.astro`): a thumbs-up / thumbs-down vote below the page content, captured to PostHog as `docs_feedback` with `{ helpful, page }`. It renders from `Footer.astro`'s sidebar branch — the same indirection the Cloud CTA uses — rather than a per-page import or frontmatter flag, so every content page gets it automatically, including ones not written yet; it sits *below* the Cloud CTA on the pages that carry one, and splash pages (the homepage and 404) take the other footer branch and never render it. One vote per page per visitor: the choice is remembered in `localStorage` keyed by pathname, and a revisit renders the thanks message instead of re-prompting (storage is a nicety, not the record — a browser with storage disabled still votes). - **Settings-directory validation — `wavehouse validate [dir]`** (`internal/settings/` (new: `settings.go`, `validate.go`, `decode.go`, `finding.go`, + tests), `cmd/wavehouse/validate.go` (new, + tests), `cmd/wavehouse/main.go`): first piece of the file-based control plane (settings live in a directory of JSON documents — `roles.json`, `policies.json`, `pipes.json`, `config.json` — that a running instance will hot-reload; this change is validation-only — boot loading and reload wiring land separately). `settings.Validate(dir)` is the single gate every consumer of the directory runs: deliberately pure (no network, no ClickHouse — table/column existence stays with schema discovery, per Bring-Your-Own-Schema), and it collects **all** findings in one pass instead of failing on the first. Checks, layered: the directory holds exactly the four files (a missing file is an error — an empty document is `{}`, so absence always means deletion or a wrong path; any unexpected entry — file or directory — is an error so a typoed `polices.json` or a stray backup can't be silently ignored; dot-prefixed entries are the one carve-out, since erroring on vim swap files or the `..data` machinery Kubernetes ConfigMap mounts publish through would break hand editing and the cloud fan-out's mount pattern alike); strict JSON syntax (unknown fields rejected — the JSON form of the retired-config-key trap; empty/truncated files rejected, never read as an empty document; a leading UTF-8 byte order mark named as such instead of surfacing as a cryptic invalid-character error; a directory, unreadable file, or non-regular file (a FIFO would hang the read forever waiting for a writer; a stat gate rejects it — following symlinks, so Kubernetes ConfigMap mounts' symlink layout still passes) squatting on a settings filename named as the one real problem, not double-reported as "missing"; a top-level `null` rejected — the one well-formed document that decodes into a zero value without error, so it would silently read as "no settings"; trailing content rejected; duplicated object keys detected by a token-level pass, since `encoding/json` silently keeps the last copy); per-file shape rules (role names non-empty/unique, pipe names/SQL/param types, `config.json` bounds mirroring boot-config validation — its sections are the *tenant-owned* behavioral tunables (dedupe id_field/require_id plus per-table overrides under `dedupe.tables` — each entry overrides only the fields it names, resolving table → global → compiled default per field, so the effective id_field can never be empty — an explicit empty, whitespace-only, or whitespace-padded id_field is rejected at both levels, since an exact-match JSON key lookup would silently miss every row ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)'s shape, unblocked by the file design since table names are runtime-resolved like policy grants); query default_max_rows, schema refresh_interval, CORS origins); platform-owned knobs like the SSE keepalives deliberately stay boot config); and cross-file referential integrity (every role a policy grant, `default_role`/`admin_role`, or pipe allowlist references must be declared in `roles.json`; an empty role string in a grant or allowlist is named as such — it matches no request and authorizes nobody). Warnings don't invalidate: a grant scoping the admin role (an unconditional bypass — dead config), `default_role` = admin, and a `default` on a required pipe parameter are flagged but legal. An empty `policies.json` means no policy — fail closed, matching deleted-policy semantics — and draws a warning naming the total lockout, so it announces itself at validation time instead of one 403 at a time. The CLI (`cmd/wavehouse/validate.go`, following the `health` subcommand pattern) takes the directory as an argument or from `WH_SETTINGS_DIR`, prints findings, and exits 0/1/2 (valid/invalid/usage) so CI and operators can gate config changes before they reach a running instance. The dispatch in `main.go` also grows `help` and `version` subcommands, and an unknown command is now a usage error instead of silently falling through and starting the server (`wavehouse validat` booting a listener is not a typo anyone wants); each subcommand parses its arguments with a stdlib `flag.FlagSet`, so `wavehouse -h` prints command-specific help and a stray flag or argument is a usage error rather than being silently swallowed. `WH_SETTINGS_DIR` has a single authority: `config.EnvSettingsDir`, with a reflection test pinning the `settings.dir` struct tag to it. The directory's location joins boot config as `settings.dir` (`WH_SETTINGS_DIR`; `internal/config/config.go`, `config.yaml`, `docs/src/content/docs/configuration.mdx`) — boot-tier by necessity, since it's the pointer the reload machinery follows; no default, same silent-misconfiguration reasoning as `policy.file_path`. @@ -32,7 +32,9 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Changed -- **Message-queue subjects lead with the tenant, and the async paths read it off each message** (`internal/mq/{mq,subject,embedded}.go`, `internal/api/{ingest,stream}.go`, `internal/stream/hub.go`, `internal/ingest/{worker,sweeper}.go`, `internal/app/wire.go`, `docs/src/content/docs/{architecture,ingest-pipeline,deployment,api}.md`, `docs/src/content/docs/{settings-directory,access-control}.mdx`, `AGENTS.md`): story 5a of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). `mq.Topic` gains a `Tenant`, the leading token of every subject — `ingest..
[.]`, `dlq..
[.]` — placed verbatim, since the tenant-id grammar makes it one token, so one wildcard selects a tenant's traffic (`ingest.acme.>`); a settings directory that holds the four files produces the same subjects with `0` as the token, and nothing else about them changes. The ingest and stream handlers address the request's tenant, which the resolved store carries (`settings.Store.Tenant`, from story 8). The stream hub indexes subscribers by the full topic and evaluates each event under its own tenant's policy, so a subscriber on one tenant's table never receives another tenant's rows for a table of the same name; a gap-fill and the opening schema frame read the connection's tenant. The ingest worker reads each message's tenant off its topic, batches per tenant table, and resolves the dead-letter switch under the row's own tenant — an envelope it cannot read is parked or dropped under the topic's tenant too — and bumps that tenant's cache namespaces (story 8's `invalidate` now receives the message's tenant rather than tenant `0`; the wiring's `sharedTables` still repeats each bump under every tenant the directory holds, since every tenant reads the same ClickHouse table until story 6). The sweeper keeps the longest `stream.gap_window_minutes` among the tenants being served: the ingest queue is one stream and a purge is one bound over it, so purging less is the safe direction until the streams are per tenant. `Publish` and gap-fill refuse a topic whose tenant is empty or outside the grammar, so nothing lands on tenant `0` by omission. The subject change needs no drain of its own (the v2 envelope's drain, below, still applies): the durable consumers filter `ingest.>`, which the previous two-token subjects match, and a subject with no tenant token reads as tenant `0`'s, so the subject an event in flight arrived on changes nothing about how it is inserted, streamed, or parked, and rows already parked keep counting in `GET /v1/ops/dlq/stats` — which sums a table across tenants until the queue is per tenant; a gap-fill spanning the upgrade omits the pre-upgrade events for one gap window. Per-tenant JetStream streams, the DLQ shrink guard, and `?tenant=` on the DLQ stats route are story 5b, after a research spike. +- **Each tenant has a message queue of its own** (`internal/mq/{mq,subject,embedded}.go` (+ tests), `internal/ingest/{sweeper,worker}.go` (+ tests), `internal/api/{dlq,ingest}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `internal/stream/subscriber.go`, `internal/settings/{settings,store}.go`, `internal/testutil/{mocks,testutil}.go`, `clients/ts/src/{dlq,types}.ts` (+ tests), `docs/src/content/docs/{deployment,api,architecture,ingest-pipeline,durability,why-wavehouse}.md`, `docs/src/content/docs/sdk/{admin,reference}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `AGENTS.md`): story 5b of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). The embedded NATS server keeps each tenant's events on a pair of JetStream streams of its own — `INGEST_` (`ingest..>`, `DiscardNew`) at the tenant's own `mq.max_bytes_gb`, and `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed; subjects are unchanged, and nothing outside `internal/mq` names a stream. A tenant at its budget gets `503` while the others keep publishing, and the ingest worker's and the hub bridge's durables are held on every tenant's stream, each with its own ack floor and `MaxAckPending`, so one tenant's backlog holds back neither another's delivery nor its purge; the worker's prefetch is shared across the tenants' streams. A removed or rejected tenant's stream is still consumed, so its queued rows reach the worker and are parked on its own dead-letter queue. The sweeper purges each tenant's stream at that tenant's own `stream.gap_window_minutes`, keeping no acknowledged history for a tenant no longer served, and every served tenant's `mq.max_bytes_gb` is applied after each reload rather than tenant `0`'s alone (the boot warning about a nested directory with no tenant `0` is gone with it). A reload that shrinks a budget no longer deletes dead letters: a dead-letter stream holding more than a tenth of the new budget keeps what it holds, and that is logged — the interim guard of [#532](https://github.com/Wave-RF/WaveHouse/issues/532). JetStream's own check of the streams' caps against the disk, which made them fit 75% of the free disk at boot together, is lifted: a budget is a cap and never a reservation, what the budgets add up to against the disk is [#138](https://github.com/Wave-RF/WaveHouse/issues/138)'s, and a flat directory whose budget exceeds three quarters of the free disk now boots where it used to be refused. A queue that cannot be opened refuses a flat boot like any other store and, over a nested directory, costs its tenant alone — its ingest answers `503`, each publish and reload trying again. `GET /v1/ops/dlq/stats` reads one tenant's dead-letter queue — the one `?tenant=` names, parsed strictly like the other admin reads, and tenant `0`'s without it, no longer the sum across tenants — answering for a rejected or removed tenant too and `404` for a tenant with no queue; the SDK's `wh.dlq.list()` and `.table()` take a `tenant` option. The streams an earlier build kept for every tenant together (`WAVEHOUSE`, `WAVEHOUSE_DLQ`) overlap every tenant's subjects and are deleted at boot, what they held with them, and a subject with no tenant token no longer reads as tenant `0`'s. + +- **Message-queue subjects lead with the tenant, and the async paths read it off each message** (`internal/mq/{mq,subject,embedded}.go`, `internal/api/{ingest,stream}.go`, `internal/stream/hub.go`, `internal/ingest/{worker,sweeper}.go`, `internal/app/wire.go`, `docs/src/content/docs/{architecture,ingest-pipeline,deployment,api}.md`, `docs/src/content/docs/{settings-directory,access-control}.mdx`, `AGENTS.md`): story 5a of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). `mq.Topic` gains a `Tenant`, the leading token of every subject — `ingest..
[.]`, `dlq..
[.]` — placed verbatim, since the tenant-id grammar makes it one token, so one wildcard selects a tenant's traffic (`ingest.acme.>`); a settings directory that holds the four files produces the same subjects with `0` as the token, and nothing else about them changes. The ingest and stream handlers address the request's tenant, which the resolved store carries (`settings.Store.Tenant`, from story 8). The stream hub indexes subscribers by the full topic and evaluates each event under its own tenant's policy, so a subscriber on one tenant's table never receives another tenant's rows for a table of the same name; a gap-fill and the opening schema frame read the connection's tenant. The ingest worker reads each message's tenant off its topic, batches per tenant table, and resolves the dead-letter switch under the row's own tenant — an envelope it cannot read is parked or dropped under the topic's tenant too — and bumps that tenant's cache namespaces (story 8's `invalidate` now receives the message's tenant rather than tenant `0`; the wiring's `sharedTables` still repeats each bump under every tenant the directory holds, since every tenant reads the same ClickHouse table until story 6). `Publish` and gap-fill refuse a topic whose tenant is empty or outside the grammar, so nothing lands on tenant `0` by omission. - **The query cache and its singleflight are keyed by tenant** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/settings/{store,registry}.go` (+ tests), `internal/api/{cache_key,pipes,structured_query}.go` (+ tests; `cache_tenant_test.go` new), `internal/ingest/worker.go` (+ tests), `docs/src/content/docs/{architecture,deployment,api,ingest-pipeline}.md`, `AGENTS.md`): story 8 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files — every key simply gains tenant `0`'s prefix. The tenant leads every key the cached read paths build: the key `POST /v1/query?table={table}` and `GET/POST /v1/pipes/{name}` cache a result under is `:query:` and is their singleflight key too, and a version namespace is `.
.
.` (`cache.Namespace` gains `Tenant`; scope stays where it was, and inert). So identical requests from two tenants are two entries and two flights to ClickHouse, a tenant is never served another tenant's cached rows, and a batch the ingest worker inserts bumps the namespaces of the tenant it was inserted for and no other's (`IngestWorker.invalidate` takes the tenant as a parameter — the worker's own until story 5 reads it off the message, so over a nested directory it is still tenant `0`'s namespaces every batch bumps, and another tenant's cached query results expire on their TTL alone until then). The pool stays one Ristretto instance sized by `cache.l1_max_cost`. The handlers read the tenant off the request's store — `settings.Store.Tenant`, stamped by the registry when it creates the store (the commit is shared with story 7) — so nothing new rides the request context. This lands ahead of story 6 on purpose: once each tenant has its own ClickHouse connection, a tenant-blind key would be a silent cross-tenant read. diff --git a/clients/ts/src/dlq.ts b/clients/ts/src/dlq.ts index a783e8eb8..c12762250 100644 --- a/clients/ts/src/dlq.ts +++ b/clients/ts/src/dlq.ts @@ -1,7 +1,7 @@ import { err, ok } from "./errors.js"; -import { request } from "./http.js"; +import { request, tenantParam } from "./http.js"; import type { StreamController } from "./stream/controller.js"; -import type { DLQStats, HttpContext, Result, StreamOptions } from "./types.js"; +import type { DLQStats, HttpContext, OpsRequestOptions, Result, StreamOptions } from "./types.js"; type CreateStreamFn = (table: string, opts?: StreamOptions) => StreamController; @@ -15,23 +15,27 @@ export class DLQNamespace { this._createStream = createStream; } - /** Get DLQ statistics (message counts per table). */ - async list(opts?: { signal?: AbortSignal }): Promise> { + /** + * Get DLQ statistics (message counts per table) — of `opts.tenant`, the + * default tenant without it. A tenant with no dead-letter queue is a `404`. + */ + async list(opts?: OpsRequestOptions): Promise> { const { data, error } = await request(this._ctx, { method: "GET", path: "/v1/ops/dlq/stats", + params: tenantParam(opts), signal: opts?.signal, }); if (error) return err(error); return ok(data!); } - /** Get DLQ stats filtered by table name. */ - async table(name: string, opts?: { signal?: AbortSignal }): Promise> { + /** Get DLQ stats filtered by table name — of `opts.tenant`, the default tenant without it. */ + async table(name: string, opts?: OpsRequestOptions): Promise> { const { data, error } = await request(this._ctx, { method: "GET", path: "/v1/ops/dlq/stats", - params: { table: name }, + params: { table: name, ...tenantParam(opts) }, signal: opts?.signal, }); if (error) return err(error); diff --git a/clients/ts/src/namespaces.test.ts b/clients/ts/src/namespaces.test.ts index d1c0db5d6..a32a9dfec 100644 --- a/clients/ts/src/namespaces.test.ts +++ b/clients/ts/src/namespaces.test.ts @@ -146,6 +146,25 @@ describe("DLQNamespace", () => { expect(fetchSpy.mock.calls[0][0]).toContain("table=clicks"); }); + it("list() and table() send opts.tenant as ?tenant=, and nothing without it", async () => { + fetchSpy.mockImplementation( + async () => new Response(JSON.stringify({ tables: {}, total: 0 }), { status: 200 }), + ); + const ns = new DLQNamespace(makeCtx(), mockStream); + + await ns.list({ tenant: "acme" }); + await ns.table("clicks", { tenant: "acme" }); + await ns.list(); + await ns.table("clicks"); + + const urls = fetchSpy.mock.calls.map((call) => new URL(call[0])); + expect(urls[0].pathname + urls[0].search).toBe("/v1/ops/dlq/stats?tenant=acme"); + expect(urls[1].searchParams.get("table")).toBe("clicks"); + expect(urls[1].searchParams.get("tenant")).toBe("acme"); + expect(urls[2].search).toBe(""); + expect(urls[3].search).toBe("?table=clicks"); + }); + it("stream() delegates to createStream", () => { const ctrl = {} as any; mockStream.mockReturnValue(ctrl); diff --git a/clients/ts/src/types.ts b/clients/ts/src/types.ts index 5feae56d9..158d5c449 100644 --- a/clients/ts/src/types.ts +++ b/clients/ts/src/types.ts @@ -471,8 +471,8 @@ export interface PipeRequestOptions { /** * Options for a call to one of the admin routes that address a tenant: * `wh.pipes.list()`, `wh.pipes.get()`, `wh.settings.reload()`, - * `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and - * `wh.sql()`. + * `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()`, + * `wh.dlq.list()`, `wh.dlq.table()` and `wh.sql()`. */ export interface OpsRequestOptions { signal?: AbortSignal; diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index 75311090c..1634aaabc 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -745,14 +745,16 @@ Triggers an immediate re-discovery of the `?tenant=`'s ClickHouse table schemas #### `GET /v1/ops/dlq/stats` — DLQ Statistics -Returns per-table message counts in the Dead Letter Queue — a table's count summed across tenants, since one queue serves every tenant until each has its own. Admin-only, like the rest of this section. Whether a poison row lands here is the settings directory's [`dlq.enabled`](/settings-directory#dead-letter-queue) switch (global or per table); the stream and this endpoint always exist. Before any failure has ever occurred, the endpoint returns `200` with `{"tables":{},"total":0}`. +Returns per-table message counts in one tenant's Dead Letter Queue: the [tenant](/deployment#the-nested-settings-directory) an optional `?tenant=` names, the default tenant `0` without it, which is the whole settings directory unless it is nested. The queue is read from the message queue rather than the settings, so a tenant whose folder was rejected or removed is read like one being served, for as long as its queue is kept. The query string is parsed strictly, as on the other admin reads. Admin-only, like the rest of this section. Whether a poison row lands here is the settings directory's [`dlq.enabled`](/settings-directory#dead-letter-queue) switch (global or per table); a tenant's dead-letter stream exists from the moment the tenant is first served, and this endpoint always exists. Before any failure has ever occurred, the endpoint returns `200` with `{"tables":{},"total":0}`. **Error responses:** | Status | Body | Cause | | ------ | ---- | ----- | | 401 | `{"error":"invalid token"}` / `{"error":"token expired"}` | A present-but-invalid/expired token was supplied and denied (the gate surfaces the token reason) | +| 400 | `{"error":"invalid query string: …"}` / `{"error":"invalid ?tenant: …"}` | The query string does not parse (`?tenant=acme;x=1`, a bad `%` escape), or `tenant` is empty, repeated, or not a tenant id | | 403 | `{"error":"forbidden"}` | Caller's role is not the policy `admin_role` (`"admin"` by default) | +| 404 | `{"error":"no dead-letter queue for tenant: "}` | The tenant has no dead-letter queue: it has never been served on this data directory, or the id names no tenant | | 500 | `{"error":"stream info failed"}` | NATS JetStream stream-info lookup failed | | 503 | `{"error":"token verifier not ready: the tenant's JWKS has not been fetched yet"}` | A token was supplied, with no valid operator key, while tenant `0`'s JWKS has not been fetched yet (the ops tree verifies as tenant `0`); refused before any policy runs, with a `Retry-After: 30` header — see [Authentication](#authentication) | @@ -760,6 +762,7 @@ Returns per-table message counts in the Dead Letter Queue — a table's count su | Param | Type | Default | Description | | ----- | ---- | ------- | ----------- | +| `tenant` | string | `0` | The tenant whose dead-letter queue is read. | | `table` | string | — | Filter stats to a specific table name (e.g., `?table=clicks` returns only the `clicks` count). | **Response:** @@ -873,9 +876,9 @@ Three values, where the envelope above has four: this is the frame a role restri ## Dead Letter Queue (DLQ) -When a batch insert to ClickHouse fails (e.g., type errors, connection issues), the worker re-inserts the batch row by row: rows that succeed are acked, and only the rows that fail again are published to the DLQ NATS stream (`WAVEHOUSE_DLQ`) under subjects `dlq.{tenant}.{table}` (the tenant the row was ingested under; `0` for a settings directory that holds the four files). This prevents infinite retry loops — those messages are ACKed from the main stream and moved to the DLQ for inspection. A batch whose tenant has no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips that retry, which no row of it could pass, and is parked whole; only a served tenant whose DLQ is off for the table leaves it for redelivery, since a tenant no longer served has no switch to read. A second class lands here too: an envelope the worker cannot *read* at all — malformed JSON, an unknown **or absent** `format` (a pre-v2 message has no `format` field at all, which is how it presents here), or `columns` and `row` that do not pair — is parked without ever reaching a table batch, which is what an operator sees after upgrading across the wire change without draining first. **Two different body shapes land here, and a consumer must not assume one decoder.** A row that failed its INSERT is parked as the `EventMessage` envelope above. An envelope the worker could not *read* is parked as **its original bytes, verbatim** — `parkOnDLQ` republishes what arrived — so it is whatever the producer sent: a pre-v2 `data` object, malformed JSON, or a v2 envelope whose `columns` and `row` do not pair. Being undecodable as an `EventMessage` is precisely why it was parked, so decode defensively and fall back on the `X-DLQ-Error` header, which names the reason. For the first shape the body is the published `EventMessage` envelope (`{"table_name":…,"scope":"","received_timestamp":…,"format":…,"columns":[…],"row":[…]}` — the failed row is the `row` array, read against `columns`, its `DateTime`/`DateTime64` values as published: canonicalized where WaveHouse could parse them, otherwise the producer's original spelling — see [timestamp canonicalization](#timestamp-canonicalization)); the failure reason, table, and time travel in the `X-DLQ-Table` / `X-DLQ-Error` / `X-DLQ-Timestamp` message headers. +When a batch insert to ClickHouse fails (e.g., type errors, connection issues), the worker re-inserts the batch row by row: rows that succeed are acked, and only the rows that fail again are published to the tenant's own DLQ NATS stream (`DLQ_{tenant}`) under subjects `dlq.{tenant}.{table}` (the tenant the row was ingested under; `0` for a settings directory that holds the four files). This prevents infinite retry loops — those messages are ACKed from the main stream and moved to the DLQ for inspection. A batch whose tenant has no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips that retry, which no row of it could pass, and is parked whole; only a served tenant whose DLQ is off for the table leaves it for redelivery, since a tenant no longer served has no switch to read. A second class lands here too: an envelope the worker cannot *read* at all — malformed JSON, an unknown **or absent** `format` (a pre-v2 message has no `format` field at all, which is how it presents here), or `columns` and `row` that do not pair — is parked without ever reaching a table batch, which is what an operator sees after upgrading across the wire change without draining first. **Two different body shapes land here, and a consumer must not assume one decoder.** A row that failed its INSERT is parked as the `EventMessage` envelope above. An envelope the worker could not *read* is parked as **its original bytes, verbatim** — `parkOnDLQ` republishes what arrived — so it is whatever the producer sent: a pre-v2 `data` object, malformed JSON, or a v2 envelope whose `columns` and `row` do not pair. Being undecodable as an `EventMessage` is precisely why it was parked, so decode defensively and fall back on the `X-DLQ-Error` header, which names the reason. For the first shape the body is the published `EventMessage` envelope (`{"table_name":…,"scope":"","received_timestamp":…,"format":…,"columns":[…],"row":[…]}` — the failed row is the `row` array, read against `columns`, its `DateTime`/`DateTime64` values as published: canonicalized where WaveHouse could parse them, otherwise the producer's original spelling — see [timestamp canonicalization](#timestamp-canonicalization)); the failure reason, table, and time travel in the `X-DLQ-Table` / `X-DLQ-Error` / `X-DLQ-Timestamp` message headers. -Use `GET /v1/ops/dlq/stats` to monitor DLQ depth. +Use `GET /v1/ops/dlq/stats` to monitor DLQ depth, per tenant (`?tenant=`). ## Generating a JWT for Testing diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 9d1dc6363..6eaf3d54e 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -77,20 +77,20 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request - **router.go** — Route definitions. Public: `/livez`, `/readyz`, and the content-free `/v1/health` SDK ping (plus the permanent `/healthz` alias and the deprecated `/health`, `/ready` aliases). Policy-gated: `/v1/ingest?table={table}`, `/v1/query?table={table}` (structured), `/v1/pipes/{name}` (named pipes), `/v1/stream`. Admin-only (`RequireAdmin` — role == `policy.admin_role`, or a request bearing the operator key's operator bit, which passes even under a nil policy; over a nested settings directory `NewRouter` mounts the gate with no policy at all, whatever `Dependencies.PolicySource` was wired, so the operator key alone passes): `/v1/ops/schema/*`, `/v1/ops/dlq/stats`, `GET /v1/ops/pipes[/{name}]`, `/v1/ops/settings/reload`, `/v1/ops/query` (raw SQL — same gate as the rest of `/v1/ops/*`). - **auth middleware** — the JWT/JWKS authentication middleware is its own package, [`auth/`](#auth--authentication); the router runs it on every `/v1/*` route. -- **tenant.go** — `TenantMW` resolves the request's tenant ahead of the auth middleware on every `/v1` route outside `/v1/ops/*`: the [`X-Tenant-ID`](/deployment#multi-tenant-deployments) header (absent means `tenant.Default`), validated by `tenant.Parse` (`400`), looked up in the `settings.Registry` (`resolveStore`: `404` for an id it does not hold, a bare `503` for a tenant whose folder was rejected — the findings stay out of a body answered before authentication), and the resolved `*settings.Store` stored in the request context (`WithStore` / `StoreFromContext` — here rather than in `tenant/`, because `settings` names `tenant.ID`). A handler reads the store once and passes it down as an argument — the per-tenant getters it holds take it as a parameter (`(*settings.Store).Policy`, `.DedupeFor`, `.DefaultMaxRows`, … in production) — and nothing below a handler reads the context; a tenant route reached without a resolved store answers `500` rather than fall back to a tenant. The probes, `/version`, the metrics path, and `/v1/ops/*` are tenant-exempt; an ops route that addresses one tenant — the admin pipe reads, the schema routes, the raw-SQL proxy and the settings reload — names it in `?tenant=` (`opsTenant`, and `opsStore` over it for the routes that need the tenant's store), parsed strictly so that a query `url.ParseQuery` would half-read is a `400` rather than a read of the default tenant, which is what it means when absent on the reads (on the reload, absent is the whole directory). +- **tenant.go** — `TenantMW` resolves the request's tenant ahead of the auth middleware on every `/v1` route outside `/v1/ops/*`: the [`X-Tenant-ID`](/deployment#multi-tenant-deployments) header (absent means `tenant.Default`), validated by `tenant.Parse` (`400`), looked up in the `settings.Registry` (`resolveStore`: `404` for an id it does not hold, a bare `503` for a tenant whose folder was rejected — the findings stay out of a body answered before authentication), and the resolved `*settings.Store` stored in the request context (`WithStore` / `StoreFromContext` — here rather than in `tenant/`, because `settings` names `tenant.ID`). A handler reads the store once and passes it down as an argument — the per-tenant getters it holds take it as a parameter (`(*settings.Store).Policy`, `.DedupeFor`, `.DefaultMaxRows`, … in production) — and nothing below a handler reads the context; a tenant route reached without a resolved store answers `500` rather than fall back to a tenant. The probes, `/version`, the metrics path, and `/v1/ops/*` are tenant-exempt; an ops route that addresses one tenant — the admin pipe reads, the schema routes, the raw-SQL proxy, the settings reload and the DLQ stats — names it in `?tenant=` (`opsTenant`, and `opsStore` over it for the routes that need the tenant's store; the DLQ stats need none, since the MQ holds the queue), parsed strictly so that a query `url.ParseQuery` would half-read is a `400` rather than a read of the default tenant, which is what it means when absent on the reads (on the reload, absent is the whole directory). - **pipes.go** — Named query pipe handlers: admin listing (`GET /v1/ops/pipes[/{name}]`, read per request from its `pipes.Source`) and execution with parameter binding. `pipes.json` is the only write path. - **structured_query.go** — Handler for `POST /v1/query?table={table}`: validates query AST, enforces permissions, builds and executes SQL. - **ingest.go** — Accepts `POST /v1/ingest?table={table}` in three body shapes: one flat JSON object, a JSON array of them, or NDJSON. The **required** `Content-Type` chooses the format *family* — `application/json` versus the four NDJSON spellings — and within the JSON family the body's first non-whitespace byte picks array versus single object; the bytes never choose the family. Anything that is not exactly one readable media type is a `415`, decided before the body is read: the header is parsed per RFC 9110 §8.3, and because `Content-Type` is a singleton field, repeated header lines must all resolve to the same format and a value carrying a comma is refused unless the value as a whole parses as one media type — a comma inside a *quoted* parameter value is data, so `application/json; a=", application/x-ndjson; b="` is accepted. It then reads the whole (`MaxBytesReader`-capped) body into a pooled buffer and runs the per-format record readers over those bytes, so the `413` lands before any record is processed and peak memory per request is O(body) rather than O(record). Then it validates each record against the discovered schema, optional dedup, and publishes each row through `mq.Publisher` on `mq.Topic{Tenant, Table, Scope}` (the request's tenant, read off its resolved store — `store.Tenant()` — and raw names; the subject it becomes is `internal/mq`'s; a full queue comes back as `mq.ErrQueueFull`, which is the `503` + `Retry-After`). When dedup is on, a row missing the configured `id_field` can't be deduped: it is logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total` (labeled by `table`), then published un-deduped — or rejected when `dedupe.require_id` is set ([#219](https://github.com/Wave-RF/WaveHouse/issues/219)). - **query.go** — Proxies raw SQL for `POST /v1/ops/query` straight to the `?tenant=`'s ClickHouse HTTP interface (`chconn.Pools.Target` by the resolved store's tenant; the zero target — no pool — is a `503` with `Retry-After`). **Not cached** — sets `Cache-Control: no-store` so every request hits ClickHouse; DateTime is rendered ISO-8601 via `date_time_output_format=iso` (the Go-side type conversion lives in the structured-query / pipes path, not here). - **stream.go** — Real-time streaming via SSE. Callers select a table with the `?table=` query parameter. Each connection registers one `Subscriber` (the `stream/` package) with both the event `Hub` (under its `(topic, role)`) and the shared keepalive wheel, then drains both from a single byte-pump — so idle streams keep emitting `:` keepalive comments (surviving reverse-proxy idle timeouts) while live events arrive already projected and serialized. Per-event projection/serialization happens **once per role** in the `Hub`, not once per subscriber ([#294](https://github.com/Wave-RF/WaveHouse/issues/294)); the handler also snapshots the connection's JWT claims onto the `Subscriber`, which the `Hub` evaluates per subscriber when the role carries a row-level `filter` ([#319](https://github.com/Wave-RF/WaveHouse/issues/319)). Gap-fill replay (`mq.Replayer.ReplaySince` on the connection's `mq.Topic` — a `DeliverByStartTime` consumer inside `internal/mq`) stays per-connection (low-volume, one-time on connect). A stream ends, a gap-fill in progress included, when the server begins shutting down (`Closing`) or its `Subscriber` is evicted because its tenant is no longer served (`Hub.Prune`); one admitted just before the reload that stopped serving its tenant, and registered just after the prune, is ended right after it registers (`Served`). - **schema.go** — Schema discovery API of one tenant, the `?tenant=` (`opsStore`): list all schemas, get one table, trigger refresh. `lookupSchema`, shared with the ingest and structured-query handlers, is the one reading of a `SchemaRegistry.Lookup` miss: `503` with `Retry-After` before the tenant's first discovery (`ErrNotLoaded`, or no registry built yet), `404` for a table the discovered schema lacks; the list answers the same `503` rather than `[]`. A refresh of a tenant on no pool (`discovery.ErrNoConnection`) is a `503` with `Retry-After` too. The handlers hold `RegistrySource`, `func(*settings.Store) *discovery.SchemaRegistry`, and the query paths a `func(*settings.Store) driver.Conn` beside it — each resolves the request's tenant per call, and a nil connection (a tenant no pool could be opened for, such as by the connection ceiling) is a `503` ahead of the cache, so nothing cached before is served. -- **dlq.go** — DLQ stats endpoint (`GET /v1/ops/dlq/stats`): asks `mq.DeadLetterStats.DeadLetterCounts` for the per-table parked counts (optionally one table) and the total. A dead-letter queue that does not exist (`mq.ErrNoDeadLetterQueue`) reads as empty; any other failure to read it is a 500. The queue itself is `internal/mq`'s. +- **dlq.go** — DLQ stats endpoint (`GET /v1/ops/dlq/stats`): asks `mq.DeadLetterStats.DeadLetterCounts` for one tenant's per-table parked counts (optionally one table) and its total — the tenant `?tenant=` names, read strictly by `opsTenant`, tenant `0` without it. The tenant is looked up in the MQ, not the settings registry, so a rejected or removed tenant's parked rows are read like a served one's; a tenant with no dead-letter queue (`mq.ErrNoDeadLetterQueue`) is a 404, and any other failure to read it a 500. The queue itself is `internal/mq`'s. - **health.go** — Liveness (`/livez`), readiness (`/readyz`), and a content-free `Online` ping (`/v1/health`, the SDK's public liveness check); `/healthz` is a permanent alias of `/livez`, and `/health`/`/ready` are deprecated aliases. All three consult an optional `BootState` so they can return 503 while boot-time schema discovery is still failing in the retry loop (see `internal/app`; over a nested directory, while no tenant's has succeeded); once `BootState.Set(nil)` fires, `/livez` returns 200 and stays there. `/readyz` additionally runs a `Ping` each call — `chconn.Pools.Ping` in production: every open pool at once, ready at the first answer, every pool's error joined when none answers; `/v1/health` deliberately does not. ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability, ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one resource a process still has one of, the MQ byte budget, follows the default tenant: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`; the zero value when a nested directory has never served a tenant `0`, warned about once at boot), and `onDefaultAdopt` runs its hook only after a reload that adopted it, so another tenant's reload never moves it and a `0` folder that a reload rejects or removes leaves it as it was — like the one setting read per request that follows tenant `0`, the ops gate's admin role. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. Two settings are shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)); and the sweeper keeps the longest `stream.gap_window_minutes` (`longestGapWindow`, read every sweep), since the ingest queue is one stream and a purge is one bound over it — a stream per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 5b) gives each its own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe` builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. The `mq.max_bytes_gb` hook only hands the adopted budget to `mq.Broker.SetMaxBytes` under the App's stop context; how it is split across the streams, the time bounds, and the rollback are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe` builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ` hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out @@ -139,16 +139,16 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **worker.go** — `StartIngestWorker` launches an ingest pipeline: a durable `buffer-consumer` consumer of the ingest queue (created through `mq.ConsumerManager`) reads events, batches them per tenant table — the tenant read off each message's `mq.Topic` — and performs bulk INSERTs to ClickHouse. The pipeline is **insert-only**. The wire format `EventMessage` carries `{table_name, scope, received_timestamp, format, columns, row}` — the row positionally as one `JSONCompactEachRow` line, with `columns` naming its positions (the table's insertable columns — a computed one cannot be named in an `INSERT`); the worker batches per (tenant, table, column list) and writes `INSERT INTO … (cols) FORMAT JSONCompactEachRow`. It accepts any table name (events are addressed by `mq.Topic{Tenant, Table, Scope}` with raw names; `internal/mq` encodes them into subject tokens), then bulk-INSERTs. The embedded NATS server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only publishers that can reach the ingest queue are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (`policy.admin_role`) — see the Query Path section below; the `/v1/ops/*` `RequireAdmin` middleware enforces the check at the API layer, so a no/invalid-token request (resolved to `default_role`, not admin in a production config) never reaches the proxy. On a bulk-insert failure the batch is re-inserted row by row — except a batch whose tenant has no ClickHouse connection (no longer served, or no pool could be opened for it, such as by the connection ceiling), which no row could pass and `parkBatch` takes to the DLQ switch whole, logging once per batch rather than twice per row; rows that succeed are acked, and only the rows that fail again are routed to the DLQ (`sendToDLQ` → `mq.DeadLetterer.DeadLetter`), which parks the as-published `EventMessage` envelope under the topic it arrived on (`dlq.{tenant}.{table}` subjects inside `internal/mq`) with the failure context in `X-DLQ-*` headers when the tenant's `dlq.enabled` is on for the table — see [Ingest Pipeline](/ingest-pipeline) for the worker internals. - **types.go** — `EventMessage` struct (TableName, Scope — reserved, always empty today, ReceivedTimestamp, Format, Columns, Row; `Format` is `FormatJSONCompactEachRow` and `Row` is one positional line whose slots `Columns` names) and `BufferConsumerName` constant, shared across API handlers and the ingest pipeline. - **compact.go** — `EncodeCompactRow`, the positional row encoder every published row goes through, rendering one record over the table's **insertable** columns in declaration order. Serialization only: it validates nothing and judges no value. -- **sweeper.go** — `Sweeper` implements the Active Sweeper pattern. It runs every minute and asks the MQ (`mq.Purger.PurgeAcked`) to drop the ingest events that are **both** ACKed by the buffer consumer (written to ClickHouse) **and** older than the gap window (re-read every sweep: the longest `stream.gap_window_minutes` among the tenants being served — `internal/app`'s `longestGapWindow`). Finding the purge point is `internal/mq`'s (`purge.go`). +- **sweeper.go** — `Sweeper` implements the Active Sweeper pattern. It runs every minute and asks the MQ (`mq.Purger.PurgeAcked`) to drop the ingest events that are **both** ACKed by the buffer consumer (written to ClickHouse) **and** older than the gap window (re-read every sweep: each served tenant's own `stream.gap_window_minutes` — `internal/app`'s `gapWindows` — and none for a tenant no longer served). Finding the purge point is `internal/mq`'s (`purge.go`). ### `mq/` — Message Queue The **only** package that imports NATS/JetStream — a `depguard` rule in `.golangci.yml` fails `make lint` on any `github.com/nats-io` import in every package golangci-lint builds; the `integration`-tagged files under `tests/` sit outside its default build context, so the boundary there rests on convention (AGENTS.md Key Design Decision #20). Every other package talks to the broker through the types below, so a subject, stream, or broker change lands here once. -- **mq.go** — The owned surface, stated as intent rather than broker mechanics. `Topic{Tenant, Table, Scope}` is the only address the rest of the process handles (a validated tenant id and raw names; comparable, so the SSE hub keys its index by the value). `Message` carries `Data`, its topic (`Topic()` decodes the delivered key on demand — the tenant included, which is how the hub bridge and the worker learn whose event it is; `TopicKey()` is the delivered form, for log lines), and the ack family (`DoubleAck(ctx)`, `Ack()`, `Nak()`); `Headers` is the message header map (`Add`/`Set`/`Get`, exact-key) that `PublishOpt`s such as `WithHeader` shape. Interfaces: `Publisher` (`ErrQueueFull` when the ingest queue is at its byte budget — the API's 503 + `Retry-After`), `Subscriber` (every ingest event, under a named durable consumer — the hub bridge), `ConsumerManager` → `Consumer` (a durable explicit-ack consumer from a `ConsumerConfig`; `Consume` delivers on the client goroutine so a blocking handler is backpressure, and returns a `stop` plus a `failed` channel that reports delivery ending on its own — `ErrDeliveryEnded`, e.g. a deleted consumer or a closed connection — since no message would ever say so) for the ingest worker, `DeadLetterer.DeadLetter` (park a message under its own topic; the caller acks), `DeadLetterStats.DeadLetterCounts` (`ErrNoDeadLetterQueue` when there is none), `Purger.PurgeAcked` (drop what is both acked by a consumer and stored before a cutoff; `ErrConsumerNotFound` when the consumer has not been created yet) for the sweeper, and `Replayer.ReplaySince` for SSE gap-fill. `Broker` composes them with the byte budget (`SetMaxBytes`/`MaxBytes`), `Stats`, and `Close`; it is what `internal/app` holds. -- **subject.go** — The embedded broker's naming, private to the package: the stream names (`WAVEHOUSE`, `WAVEHOUSE_DLQ`), the `ingest.`/`dlq.` prefixes and `>` wildcards, the subject-token encoder (alphanumerics and `_` pass, everything else is percent-encoded, so a name can never split or wildcard a subject), and `Topic` ↔ subject conversion. A subject is `.
[.]`: the tenant verbatim — its grammar (`tenant.Parse`) makes it one token, and it is checked on the way to the wire, so a topic without one has no subject — then the table and scope as encoded tokens; tenant first so one wildcard selects a tenant's traffic (`ingest.acme.>`). A topic has the same tail on both streams, so parking on the DLQ is a prefix swap on the delivered subject — nothing is decoded or re-encoded. A tail of one token is the form written before the tenant led it and reads as tenant `0`'s table, which is how the events in flight across that upgrade keep inserting, streaming, and counting. -- **purge.go** — The Active Sweeper's arithmetic over JetStream sequences: purge target = `MIN(consumer ack floor + 1, first sequence stored at or after the cutoff)`, the latter found by binary search over message timestamps (~15 lookups). Every uncertainty resolves toward purging less: a sequence that holds no message is kept as a candidate bound rather than discarding the half below it, and a lookup that fails outright aborts the sweep. Healthy state keeps exactly the gap window; ClickHouse down freezes purging; a catastrophic outage fills the stream to `MaxBytes` and `DiscardNew` pushes back. -- **embedded.go** — `EmbeddedNATS`, the one `Broker`: an in-process NATS server with JetStream. Creates stream `WAVEHOUSE` with subjects `ingest.>`, capped at the settings directory's `mq.max_bytes_gb`, and stream `WAVEHOUSE_DLQ` (`dlq.>`, `DiscardOld`) at a tenth of it — always present, since an empty stream costs nothing. `SetMaxBytes` applies a reloaded budget to both live streams as a pair: if the DLQ update fails after the ingest one succeeded, the ingest resize is undone so both stay on the previous budget — best effort, since if that undo also fails the ingest stream stays at the new limit and the DLQ at the previous, and the error says so. Its JetStream calls are bounded to ten seconds, plus five more for the rollback (a budget of its own, not the one that just expired), since a reload holds the settings store's lock while its hooks run; `MaxBytes` reports the budget last applied in full, so a failed resize is retried by the next reload. `Stats` reports connection and inbound-message counters for `observability.RegisterSystemMetrics`. Trace context rides in the message headers: `Publish` applies `observability.InjectHeaders`, and a message delivered through `Subscribe` (the hub bridge) carries `observability.ExtractHeaders` on its `Ctx`; the worker's `Consumer` path skips the extraction, since it batches across messages and reads no per-message context. +- **mq.go** — The owned surface, stated as intent rather than broker mechanics. `Topic{Tenant, Table, Scope}` is the only address the rest of the process handles (a validated tenant id and raw names; comparable, so the SSE hub keys its index by the value). `Message` carries `Data`, its topic (`Topic()` decodes the delivered key on demand — the tenant included, which is how the hub bridge and the worker learn whose event it is; `TopicKey()` is the delivered form, for log lines), and the ack family (`DoubleAck(ctx)`, `Ack()`, `Nak()`); `Headers` is the message header map (`Add`/`Set`/`Get`, exact-key) that `PublishOpt`s such as `WithHeader` shape. Interfaces, each speaking per tenant and never per stream: `Publisher` (`ErrQueueFull` when the tenant's ingest queue is at its byte budget or not open yet — the API's 503 + `Retry-After`), `Subscriber` (every ingest event of every tenant, under a named durable consumer — the hub bridge), `ConsumerManager` → `Consumer` (a durable explicit-ack consumer from a `ConsumerConfig`, whose `MaxAckPending` holds per tenant; `Consume` delivers each tenant's messages on a goroutine of that tenant's, in order, so a blocking handler is backpressure on its own tenant alone, spreads the prefetch across the tenants, and returns a `stop` plus a `failed` channel that reports delivery ending on its own — `ErrDeliveryEnded`, e.g. a deleted consumer, a closed connection, or a tenant's queue that could not be joined — since no message would ever say so) for the ingest worker, `DeadLetterer.DeadLetter` (park a message under its own topic, in its tenant's dead-letter queue; the caller acks), `DeadLetterStats.DeadLetterCounts` (one tenant's; `ErrNoDeadLetterQueue` when it has none), `Purger.PurgeAcked` (drop what is both acked by a consumer and stored before its tenant's cutoff, everything acked for a tenant given none; `ErrConsumerNotFound` when the consumer has not been created yet) for the sweeper, and `Replayer.ReplaySince` for SSE gap-fill. `Broker` composes them with each tenant's byte budget (`SetMaxBytes`/`MaxBytes`), `Stats`, and `Close`; it is what `internal/app` holds. +- **subject.go** — The embedded broker's naming, private to the package: the stream names (`INGEST_` and `DLQ_` — prefixes that differ in their first letter, so no tenant id makes one kind's name the other's — and the one pair an earlier build kept for every tenant together, `WAVEHOUSE`/`WAVEHOUSE_DLQ`, which boot deletes), the `ingest.`/`dlq.` prefixes and `>` wildcards, the subject-token encoder (alphanumerics and `_` pass, everything else is percent-encoded, so a name can never split or wildcard a subject), and `Topic` ↔ subject conversion. A subject is `.
[.]`: the tenant verbatim — its grammar (`tenant.Parse`) makes it one token, and it is checked on the way to the wire, so a topic without one has no subject — then the table and scope as encoded tokens; tenant first so one wildcard selects a tenant's traffic (`ingest.acme.>`). A topic has the same tail on both streams, so parking on the DLQ is a prefix swap on the delivered subject — nothing is decoded or re-encoded — and the tail's first token picks the tenant's stream. +- **purge.go** — The Active Sweeper's arithmetic over JetStream sequences: purge target = `MIN(consumer ack floor + 1, first sequence stored at or after the cutoff)`, the latter found by binary search over message timestamps (~15 lookups). Every uncertainty resolves toward purging less: a sequence that holds no message is kept as a candidate bound rather than discarding the half below it, and a lookup that fails outright aborts the sweep. It runs on each tenant's stream at that tenant's cutoff. Healthy state keeps exactly the gap window; ClickHouse down freezes purging; a catastrophic outage fills the stream to `MaxBytes` and `DiscardNew` pushes back. +- **embedded.go** — `EmbeddedNATS`, the one `Broker`: an in-process NATS server with JetStream, giving each tenant a queue of its own — stream `INGEST_` with subjects `ingest..>`, capped at the tenant's `mq.max_bytes_gb` (`DiscardNew`), and stream `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — with the durable consumers on the ingest one; nothing outside the package sees that layout. JetStream's own check of the streams' caps against the disk (75% of the free disk by default) is set out of reach, so a budget is a cap and never a reservation. Boot deletes the pair an earlier build kept for every tenant together — its subjects overlap every tenant's — and takes stock of the tenants' streams on disk with their budgets, so a consumer created later is held on every one, a tenant no longer served included. `SetMaxBytes` opens a tenant's queue the first time — its dead-letter stream first, so no row is queued that could not be parked — and every registered consumer joins it; a publish or park that finds a stream missing reopens it at the budget last asked for the tenant, or is refused as a full queue with none asked yet. After that `SetMaxBytes` applies a reloaded budget to the tenant's two live streams as a pair: if the DLQ update fails after the ingest one succeeded, the ingest resize is undone so both stay on the previous budget — best effort, since if that undo also fails the ingest stream stays at the new limit and the DLQ at the previous, and the error says so. A dead-letter stream is never capped below the bytes it holds, which `DiscardOld` would delete to fit ([#532](https://github.com/Wave-RF/WaveHouse/issues/532)): it keeps what it holds, and that is logged. Its JetStream calls are bounded to ten seconds, plus five more for the rollback (a budget of its own, not the one that just expired), since a reload holds the settings store's lock while its hooks run; `MaxBytes` reports the budget last applied in full, so a failed resize is retried by the next reload. The consumers `CreateConsumer` and `Subscribe` build hold one durable on each tenant's stream, looked up before anything is written so a boot over many queues writes nothing it need not, each delivering on a goroutine of its own into the one handler. `PurgeAcked` and `DeadLetterCounts` run per tenant stream. `Stats` reports connection and inbound-message counters for `observability.RegisterSystemMetrics`. Trace context rides in the message headers: `Publish` applies `observability.InjectHeaders`, and a message delivered through `Subscribe` (the hub bridge) carries `observability.ExtractHeaders` on its `Ctx`; the worker's `Consumer` path skips the extraction, since it batches across messages and reads no per-message context. ### `observability/` — OpenTelemetry Pipeline diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index a9c9de1db..8a709426f 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -97,7 +97,7 @@ WaveHouse's per-role caps are sent as per-query `SETTINGS` on its connection, so ### Message Queue (NATS) -The stream's disk budget, `mq.max_bytes_gb`, is a hot-reloadable key in the [Settings Directory](/settings-directory#message-queue) — there is no boot-config knob for it. +Each tenant's queue has its own disk budget, `mq.max_bytes_gb`, a hot-reloadable key in the [Settings Directory](/settings-directory#message-queue) — there is no boot-config knob for it. **Durability.** The embedded server runs with JetStream `SyncAlways`, so every event is `fsync`'d to disk before `POST /v1/ingest` returns `200`. This makes your storage's `fsync` latency your ingest latency floor — see [Durability & Storage](/durability) to check whether your substrate can sustain it. There is no knob to relax this today ([#139](https://github.com/Wave-RF/WaveHouse/issues/139) tracks a configurable group-commit interval). diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 020c676c6..095b49909 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -379,19 +379,19 @@ settings/ └── roles.json ``` -That is the layout a control plane writes. Each folder's `clickhouse` block is its tenant's own ClickHouse, so a tenant answers queries once its first schema discovery against that ClickHouse succeeds (until then its schema-aware routes answer `503`, `schema not loaded yet`); what tenant `0`'s folder still supplies for the whole process — the message queue's budget, the token verifier of the routes that name no tenant, their CORS list — is listed under "What a tenant's folder decides", below. +That is the layout a control plane writes. Each folder's `clickhouse` block is its tenant's own ClickHouse, so a tenant answers queries once its first schema discovery against that ClickHouse succeeds (until then its schema-aware routes answer `503`, `schema not loaded yet`); what tenant `0`'s folder still supplies for the whole process — the token verifier of the routes that name no tenant, their CORS list — is listed under "What a tenant's folder decides", below. The folder name is the tenant id, and each folder is a complete settings directory: everything on the [Settings Directory](/settings-directory) page applies to it as written, except where the rules below say otherwise. The two shapes don't mix — a folder beside the four files, or a loose file beside the folders, is a validation error — and a running server keeps the shape it booted with, so switching is stop, restructure, start. The dedupe store needs no restructuring: it keys every tenant's seen ids by tenant, and the four files are tenant `0`, as a `0` folder is. Dot-prefixed entries are ignored in either shape. `wavehouse validate` checks either shape with the same exit codes; a finding in a nested directory names its folder (`acme/policies.json`), and a folder whose name is not a tenant id is a finding of its own — that folder is skipped, and the rest of the directory still loads. -**A rejected folder fails closed, for that tenant alone — tenant `0`'s excepted.** A folder that fails validation stops its tenant being served — its requests answer `503` — while every other tenant carries on, at boot and on a reload alike. Tenant `0` is the exception: the process still draws some shared wiring from that folder, so rejecting it costs every tenant something ("What a lost tenant `0` costs", below, says what). There is no fall back to the tenant's previous settings, unlike [the single-tenant directory](/settings-directory#loading-and-hot-reload): the recovery is fixing the folder and reloading it. A request already in flight finishes on the settings it started with, except an open `GET /v1/stream`, which is ended at once: its reconnect gets the `503` until the folder is fixed — the SDK keeps retrying and then resumes from `Last-Event-ID`, while a browser `EventSource` gives up on the `503` and has to be reopened. The rows the tenant had already accepted but not yet inserted, those of an ingest request in flight included, which still answers `200`, are parked on the DLQ under the tenant's own subject rather than held for the fix, as a removed tenant's are (see [Dead Letter Queue](#dead-letter-queue-dlq)). The findings go to the log and to the reload response, never into the `503`. A finding about the directory itself — a loose file, an entry or a directory that can't be read, a changed shape — is another matter: it refuses boot, and on a reload it rejects the reload whole and leaves every tenant as it was. +**A rejected folder fails closed, for that tenant alone — tenant `0`'s excepted.** A folder that fails validation stops its tenant being served — its requests answer `503` — while every other tenant carries on, at boot and on a reload alike. Tenant `0` is the exception: the process still draws some shared wiring from that folder, so rejecting it costs every tenant something ("What a lost tenant `0` costs", below, says what). There is no fall back to the tenant's previous settings, unlike [the single-tenant directory](/settings-directory#loading-and-hot-reload): the recovery is fixing the folder and reloading it. A request already in flight finishes on the settings it started with, except an open `GET /v1/stream`, which is ended at once: its reconnect gets the `503` until the folder is fixed — the SDK keeps retrying and then resumes from `Last-Event-ID`, while a browser `EventSource` gives up on the `503` and has to be reopened. The rows the tenant had already accepted but not yet inserted, those of an ingest request in flight included, which still answers `200`, are parked on the DLQ under the tenant's own subject rather than held for the fix, as a removed tenant's are (see [Dead Letter Queue](#dead-letter-queue-dlq)). Its message queue is kept, at the budget it last had, but the history gap-fill replays is purged from it at the next sweep, as a removed tenant's is, so a stream resumed after the fix has a hole where that history was. The findings go to the log and to the reload response, never into the `503`. A finding about the directory itself — a loose file, an entry or a directory that can't be read, a changed shape — is another matter: it refuses boot, and on a reload it rejects the reload whole and leaves every tenant as it was. -**Reloading is the writer's call.** A nested directory is not watched, because a watcher would validate a folder halfway through being written and drop its tenant. Whoever writes a tenant's folder reloads it once it is complete: `POST /v1/ops/settings/reload?tenant=acme` re-validates that folder and reads nothing else. It must name a tenant the server already holds (`404` otherwise), so a folder the server does not hold yet — one added since the last whole-directory reload — is picked up by a whole-directory reload, not by naming it; a tenant it holds but rejected is reloaded by name like any other. Without the parameter — and on `SIGHUP` — the whole directory is reloaded and mirrors its folders: a new folder becomes a tenant, and a removed one becomes unknown. That is how a tenant is removed: delete its folder, then reload the whole directory. Its open streams end, its routes answer `404`, and its queued rows are parked on the DLQ under its own subject; nothing it stored is deleted, so restoring the folder restores the tenant, seen ids included. Reloading a deleted folder by name instead leaves its tenant rejected, answering `503`. The last folder can be removed the same way, with two catches, since `wavehouse validate` and boot both read an emptied directory as the four files missing: `validate` exits `1`, so a writer that gates each reload on it has to skip the check for that one reload, and a server restarted before a folder is written back refuses to boot. A whole-directory reload re-validates every folder, so it carries the exposure the watcher would: a folder caught halfway through being written can fail validation, and its tenant then stops being served until a later reload adopts it. The response is the [single-tenant one](/api#post-v1opssettingsreload--reload-settings-directory). After a whole-directory reload, `adopted: false` with a `422` can mean adopted in part: the folders with an error among their `findings` were rejected and the rest were adopted — warnings included, since `findings` carries every folder's. +**Reloading is the writer's call.** A nested directory is not watched, because a watcher would validate a folder halfway through being written and drop its tenant. Whoever writes a tenant's folder reloads it once it is complete: `POST /v1/ops/settings/reload?tenant=acme` re-validates that folder and reads nothing else. It must name a tenant the server already holds (`404` otherwise), so a folder the server does not hold yet — one added since the last whole-directory reload — is picked up by a whole-directory reload, not by naming it; a tenant it holds but rejected is reloaded by name like any other. Without the parameter — and on `SIGHUP` — the whole directory is reloaded and mirrors its folders: a new folder becomes a tenant, and a removed one becomes unknown. That is how a tenant is removed: delete its folder, then reload the whole directory. Its open streams end, its routes answer `404`, and its queued rows are parked on the DLQ under its own subject; nothing it stored is deleted — its message queue is kept at the budget it last had, and only the history gap-fill replays goes from it, at the next sweep — so restoring the folder restores the tenant, seen ids and parked rows included. Reloading a deleted folder by name instead leaves its tenant rejected, answering `503`. The last folder can be removed the same way, with two catches, since `wavehouse validate` and boot both read an emptied directory as the four files missing: `validate` exits `1`, so a writer that gates each reload on it has to skip the check for that one reload, and a server restarted before a folder is written back refuses to boot. A whole-directory reload re-validates every folder, so it carries the exposure the watcher would: a folder caught halfway through being written can fail validation, and its tenant then stops being served until a later reload adopts it. The response is the [single-tenant one](/api#post-v1opssettingsreload--reload-settings-directory). After a whole-directory reload, `adopted: false` with a `422` can mean adopted in part: the folders with an error among their `findings` were rejected and the rest were adopted — warnings included, since `findings` carries every folder's. -**The admin routes take the operator key only.** `/v1/ops/*` reaches every tenant, so over a nested directory no tenant's admin role opens it: the [operator key](/api#authentication) alone does, and a token carrying an admin role gets `403`. Boot a nested directory without `auth.operator_key` and no caller can reach these routes at all, which leaves `SIGHUP` as the only reload; the server warns about it at boot. `GET /v1/ops/pipes`, `GET /v1/ops/pipes/{name}`, `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the same `?tenant=`, and address tenant `0` without it; `GET /v1/ops/dlq/stats` reads the queue the whole process shares and ignores the parameter. On the routes that take it the parameter is parsed strictly — a query string that does not parse, an empty or repeated `tenant`, or a malformed id is a `400`, never a silent read of the default tenant or, on the reload route, a reload of every tenant. The SDK sends it as the [`tenant` option](/sdk/admin#settings--whsettings). +**The admin routes take the operator key only.** `/v1/ops/*` reaches every tenant, so over a nested directory no tenant's admin role opens it: the [operator key](/api#authentication) alone does, and a token carrying an admin role gets `403`. Boot a nested directory without `auth.operator_key` and no caller can reach these routes at all, which leaves `SIGHUP` as the only reload; the server warns about it at boot. `GET /v1/ops/pipes`, `GET /v1/ops/pipes/{name}`, `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the same `?tenant=`, and address tenant `0` without it; `GET /v1/ops/dlq/stats` takes it too, and reads a rejected or removed tenant's dead-letter queue like a served one's, since the queue is kept; a tenant that has none is a `404`. On the routes that take it the parameter is parsed strictly — a query string that does not parse, an empty or repeated `tenant`, or a malformed id is a `400`, never a silent read of the default tenant or, on the reload route, a reload of every tenant. The SDK sends it as the [`tenant` option](/sdk/admin#settings--whsettings). -**What a tenant's folder decides, and what tenant `0`'s does.** A request is evaluated against its own tenant's `policies.json` and `pipes.json` (ingest, structured queries, pipes), its `query.*` keys, its `cors.allowed_origins`, and its `dedupe` block: whether its records are deduplicated, by which id, against the tenant's own store, which that folder's `dedupe.enabled` opens and closes on reload exactly as [the single-tenant one](/settings-directory#deduplication) does (every tenant's store is a share of the one Pebble instance at `/pebble`, each key led by its tenant), so `wavehouse_ingest_dedupe_disabled_total` ticks only across a tenant's own reload, whatever the other tenants' switches say. A tenant's seen ids are its own: the same event id is first seen under each tenant that sends it. Its `auth` block is its own too: each tenant's folder wires that tenant's token verifier (`jwks_url`, `role_claim`), built when the folder is adopted and rebuilt when its wiring changes, so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its provider's key set. Under another tenant's header a token is treated as invalid, and the request falls back to that tenant's `default_role` like any other unverifiable token, possibly after a rate-limited key refetch (see [Authentication](/settings-directory#authentication)). Keep `X-Tenant-ID` pinned at the proxy so a token is never presented under the wrong tenant. Tenants can still accept each other's tokens: those that leave `jwks_url` empty share the boot HMAC secret when `auth.jwt_secret` is set, so a token verifies under any of them (with no secret they validate no token at all), and those whose `jwks_url` names the same key set accept each other's tokens; isolate them by provider, or scope rows by a signed claim ([row-level security](/access-control#row-level-security)). A tenant whose `jwks_url` has not been fetched yet answers `503` with `Retry-After` to its token-bearing requests alone. A tenant that stops being served — its folder rejected or removed — loses its verifier and the JWKS refresh with it, and gets a fresh one when its folder is adopted again. The HMAC secret and the operator key stay boot config, shared by every tenant; the operator key is stamped with the request tenant's `admin_role`. A tenant's `clickhouse` and `schema` blocks are its own as well: each tenant reads and writes its own ClickHouse — one native pool per distinct address, database, user, password and `tls` tuple, shared by the tenants naming it, under the process-wide [connection ceiling](/settings-directory#clickhouse) — and discovers its own tables from its own database on its own `schema.refresh_interval`. The process still has one message queue, and its budget, `mq.max_bytes_gb`, follows tenant `0`'s folder. The queue is shared but addressed per tenant: an event is published on its tenant's subject (`ingest.{tenant}.{table}`), so a `GET /v1/stream` connection is authorized by its own tenant's `policies.json` and receives its own tenant's rows alone, the ingest worker inserts a row into its own tenant's ClickHouse, a failed row is parked under its own tenant's `dlq.enabled` and subject (`dlq.{tenant}.{table}`), and two tenants' tables of one name never share a batch. The query cache is one pool too, but its entries are keyed by tenant: identical `POST /v1/query` and pipe requests from two tenants are two entries and two queries to ClickHouse, and a tenant is never served another's cached rows. An insert invalidates the table's cached results under every tenant on the same ClickHouse address and database as the tenant it was ingested for, whatever their user or `tls` block, since they read the same tables; a tenant on no pool — its folder rejected or removed, or no pool could be opened for it, such as by the ceiling — is out of that fan-out while it is, and has its cached `POST /v1/query` results dropped the moment it is back on one, so a repaired or restored folder never serves query rows cached before the inserts it missed, and so does a tenant whose folder moves it to another address or database, whose cached rows came from other tables; a cached pipe result is left alone by all of this — no insert invalidates one, since it names no table — and stays until its TTL expires. Two settings weigh every tenant: the SSE keepalive, where the wheel runs at the shortest `stream.keepalive_interval` among the tenants being served, with that tenant's `stream.keepalive_buckets`; and the sweeper, which keeps the longest `stream.gap_window_minutes` among them, since every tenant's events share one message-queue stream and a purge is one bound over it. +**What a tenant's folder decides, and what tenant `0`'s does.** A request is evaluated against its own tenant's `policies.json` and `pipes.json` (ingest, structured queries, pipes), its `query.*` keys, its `cors.allowed_origins`, and its `dedupe` block: whether its records are deduplicated, by which id, against the tenant's own store, which that folder's `dedupe.enabled` opens and closes on reload exactly as [the single-tenant one](/settings-directory#deduplication) does (every tenant's store is a share of the one Pebble instance at `/pebble`, each key led by its tenant), so `wavehouse_ingest_dedupe_disabled_total` ticks only across a tenant's own reload, whatever the other tenants' switches say. A tenant's seen ids are its own: the same event id is first seen under each tenant that sends it. Its `auth` block is its own too: each tenant's folder wires that tenant's token verifier (`jwks_url`, `role_claim`), built when the folder is adopted and rebuilt when its wiring changes, so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its provider's key set. Under another tenant's header a token is treated as invalid, and the request falls back to that tenant's `default_role` like any other unverifiable token, possibly after a rate-limited key refetch (see [Authentication](/settings-directory#authentication)). Keep `X-Tenant-ID` pinned at the proxy so a token is never presented under the wrong tenant. Tenants can still accept each other's tokens: those that leave `jwks_url` empty share the boot HMAC secret when `auth.jwt_secret` is set, so a token verifies under any of them (with no secret they validate no token at all), and those whose `jwks_url` names the same key set accept each other's tokens; isolate them by provider, or scope rows by a signed claim ([row-level security](/access-control#row-level-security)). A tenant whose `jwks_url` has not been fetched yet answers `503` with `Retry-After` to its token-bearing requests alone. A tenant that stops being served — its folder rejected or removed — loses its verifier and the JWKS refresh with it, and gets a fresh one when its folder is adopted again. The HMAC secret and the operator key stay boot config, shared by every tenant; the operator key is stamped with the request tenant's `admin_role`. A tenant's `clickhouse` and `schema` blocks are its own as well: each tenant reads and writes its own ClickHouse — one native pool per distinct address, database, user, password and `tls` tuple, shared by the tenants naming it, under the process-wide [connection ceiling](/settings-directory#clickhouse) — and discovers its own tables from its own database on its own `schema.refresh_interval`. Its message queue is its own as well: its events are queued on a stream of their own, capped at its own `mq.max_bytes_gb` — at that budget its ingest answers `503` while every other tenant's keeps publishing — beside a dead-letter stream of its own at a tenth of it, and the history gap-fill replays from it is kept for its own `stream.gap_window_minutes`. Nothing checks what the tenants' budgets add up to against the disk, so size them together ([Message Queue](/settings-directory#message-queue)). An event is published on its tenant's subject (`ingest.{tenant}.{table}`), so a `GET /v1/stream` connection is authorized by its own tenant's `policies.json` and receives its own tenant's rows alone, the ingest worker inserts a row into its own tenant's ClickHouse, a failed row is parked under its own tenant's `dlq.enabled` and subject (`dlq.{tenant}.{table}`), and two tenants' tables of one name never share a batch. The query cache is one pool, but its entries are keyed by tenant: identical `POST /v1/query` and pipe requests from two tenants are two entries and two queries to ClickHouse, and a tenant is never served another's cached rows. An insert invalidates the table's cached results under every tenant on the same ClickHouse address and database as the tenant it was ingested for, whatever their user or `tls` block, since they read the same tables; a tenant on no pool — its folder rejected or removed, or no pool could be opened for it, such as by the ceiling — is out of that fan-out while it is, and has its cached `POST /v1/query` results dropped the moment it is back on one, so a repaired or restored folder never serves query rows cached before the inserts it missed, and so does a tenant whose folder moves it to another address or database, whose cached rows came from other tables; a cached pipe result is left alone by all of this — no insert invalidates one, since it names no table — and stays until its TTL expires. One setting weighs every tenant: the SSE keepalive, where the wheel runs at the shortest `stream.keepalive_interval` among the tenants being served, with that tenant's `stream.keepalive_buckets`. -**What a lost tenant `0` costs.** A `0` folder that a reload rejects or removes stops tenant `0` being served like any other, and what becomes of the shared settings depends on how they are read. `mq.max_bytes_gb` stays as tenant `0` last adopted it; tenant `0` leaves its ClickHouse pool (closed only once no served tenant names its tuple), and its schema registry and verifier are released with the folder, like any other tenant's; the `/v1/ops/*` routes, which resolve no tenant, verify against it, so a token there reads as invalid (`401`) rather than merely non-admin (`403`) until tenant `0` is served again — the operator key, which never consults a verifier, is unaffected. CORS does not stay either: the responses that read tenant `0`'s list — the tenant-exempt routes, the refusals, a preflight naming no tenant — carry no CORS headers until the folder is served again, while every other tenant's routes keep their own list. Tenant `0`'s own dedupe store closes, as any rejected or removed tenant's does, its seen ids kept for the folder that restores it. What is read per event follows the event's tenant, so tenant `0`'s events are the ones affected: with no ClickHouse to insert into, its rows fail and are parked on the DLQ whatever its switch said, and its open `GET /v1/stream` connections are ended, as any tenant's are when it stops being served — the other tenants' events are untouched. The sweeper keeps the longest gap window among the tenants still served, so tenant `0`'s history is purged at theirs, and with no tenant left being served the window is zero, which purges the acknowledged history gap-fill replays. A nested directory that has never served a tenant `0` — no `0` folder, or one rejected at boot — serves every other tenant from its own ClickHouse. Outside `/v1/ops/*`, a `/v1` request that sends no `X-Tenant-ID` resolves to tenant `0`, so with no `0` folder it answers `404 unknown tenant: 0` (`503` with a rejected one) — the SDK's `/v1/health` reachability ping included. +**What a lost tenant `0` costs.** A `0` folder that a reload rejects or removes stops tenant `0` being served like any other, and what becomes of the shared settings depends on how they are read. Tenant `0` leaves its ClickHouse pool (closed only once no served tenant names its tuple), and its schema registry and verifier are released with the folder, like any other tenant's; the `/v1/ops/*` routes, which resolve no tenant, verify against it, so a token there reads as invalid (`401`) rather than merely non-admin (`403`) until tenant `0` is served again — the operator key, which never consults a verifier, is unaffected. CORS does not stay either: the responses that read tenant `0`'s list — the tenant-exempt routes, the refusals, a preflight naming no tenant — carry no CORS headers until the folder is served again, while every other tenant's routes keep their own list. Tenant `0`'s own dedupe store closes, as any rejected or removed tenant's does, its seen ids kept for the folder that restores it. What is read per event follows the event's tenant, so tenant `0`'s events are the ones affected: with no ClickHouse to insert into, its rows fail and are parked on the DLQ whatever its switch said, and its open `GET /v1/stream` connections are ended, as any tenant's are when it stops being served — the other tenants' events are untouched. A nested directory that has never served a tenant `0` — no `0` folder, or one rejected at boot — serves every other tenant from its own ClickHouse. Outside `/v1/ops/*`, a `/v1` request that sends no `X-Tenant-ID` resolves to tenant `0`, so with no `0` folder it answers `404 unknown tenant: 0` (`503` with a rejected one) — the SDK's `/v1/health` reachability ping included. ### Upgrading behind a proxy that already sends `X-Tenant-ID` @@ -440,13 +440,9 @@ To drain before upgrading: If you skipped the drain, check `wavehouse_ingest_poison_total`, which counts both — `disposition="parked"` is recoverable from `dlq.{table}`, `disposition="dropped"` is gone — see [Dead Letter Queue](#dead-letter-queue-dlq) below. -## Upgrading across the tenant subject token - -Message-queue subjects now lead with the tenant: `ingest.{tenant}.{table}` and `dlq.{tenant}.{table}`, where a settings directory that holds the four files is tenant `0` (`ingest.0.clicks`). The subject change needs no drain of its own; the drain [the envelope upgrade above](#upgrading-across-the-v2-ingest-envelope) asks for still applies. The durable consumers filter `ingest.>`, which the previous subjects match, and a subject with no tenant token reads as tenant `0`'s, so the subject a message arrived on changes nothing about how it is inserted, streamed, or parked, and rows parked under the old `dlq.{table}` keep counting in `GET /v1/ops/dlq/stats`. The one gap is SSE gap-fill, which reads a tenant's own subject: a replay spanning the upgrade omits the events published before it, for one `stream.gap_window_minutes` (15 by default) — the same window as the envelope upgrade's. Clients that need them should backfill over REST. - ## Dead Letter Queue (DLQ) -A failed batch insert is retried row by row; while the tenant's `dlq.enabled` is `true` for the table (the seed default — a hot-reloadable [settings directory](/settings-directory#dead-letter-queue) key, overridable per table), the rows that fail again are published to the `WAVEHOUSE_DLQ` NATS stream under subjects `dlq.{tenant}.{table}` (`0` for a directory that holds the four files) instead of retrying forever. A batch whose tenant has no ClickHouse connection — one no longer served, or one no pool could be opened for, such as by the connection ceiling — skips the row-by-row retry, which no row of it could pass: its tenant's switch is read once for the whole batch, and a tenant no longer served has no switch to read, so its batch is always parked. Monitor DLQ depth via `GET /v1/ops/dlq/stats`. +A failed batch insert is retried row by row; while the tenant's `dlq.enabled` is `true` for the table (the seed default — a hot-reloadable [settings directory](/settings-directory#dead-letter-queue) key, overridable per table), the rows that fail again are published to the tenant's own dead-letter stream (`DLQ_{tenant}`) under subjects `dlq.{tenant}.{table}` (`0` for a directory that holds the four files) instead of retrying forever. A batch whose tenant has no ClickHouse connection — one no longer served, or one no pool could be opened for, such as by the connection ceiling — skips the row-by-row retry, which no row of it could pass: its tenant's switch is read once for the whole batch, and a tenant no longer served has no switch to read, so its batch is always parked. Monitor DLQ depth via `GET /v1/ops/dlq/stats`, per tenant (`?tenant=`; tenant `0` without it). ## Observability diff --git a/docs/src/content/docs/durability.md b/docs/src/content/docs/durability.md index 8548ab96c..98a8c9c27 100644 --- a/docs/src/content/docs/durability.md +++ b/docs/src/content/docs/durability.md @@ -33,8 +33,8 @@ WaveHouse does not currently expose a knob to relax this — `SyncAlways` is alw Because the publish blocks on `fsync`, **your typical ingest latency is your storage's typical `fsync` latency, and your worst-case publish is your storage's worst-case `fsync`.** When that tail is healthy (sub-millisecond to single-digit milliseconds) the guarantee is essentially free. When it is not, the same code path that handles every production message stalls: - Publishes block for the duration of the `fsync`, so a multi-second `fsync` tail is a multi-second ingest tail. -- The embedded server's stream/consumer setup and every publish run under the JetStream client's request timeout; a slow-enough substrate makes them exceed it. The boot-time symptom is `create stream: ... context deadline exceeded`. -- If the worker cannot drain to ClickHouse faster than producers publish, the stream fills toward [`mq.max_bytes_gb`](/settings-directory#message-queue) and the API returns `503` ([backpressure by construction](/ingest-pipeline#backpressure-and-durability-knobs)). +- The embedded server's stream/consumer setup and every publish run under the JetStream client's request timeout; a slow-enough substrate makes them exceed it. The symptom at a first boot, which opens every tenant's queue, is `open dlq stream: ... context deadline exceeded`; a later boot writes nothing, so the first publish is where it shows. +- If the worker cannot drain to ClickHouse faster than producers publish, a tenant's stream fills toward its [`mq.max_bytes_gb`](/settings-directory#message-queue) and the API returns `503` to that tenant ([backpressure by construction](/ingest-pipeline#backpressure-and-durability-knobs)). ## Where `SyncAlways` is cheap vs. expensive @@ -100,6 +100,6 @@ If you see any of these, benchmark the `/nats` volume as above: ## See also -- [Settings Directory → Message Queue](/settings-directory#message-queue) — `mq.max_bytes_gb`, the stream's disk budget (hot-reloadable); the SSE gap window inside it is [`stream.gap_window_minutes`](/settings-directory#streaming). +- [Settings Directory → Message Queue](/settings-directory#message-queue) — `mq.max_bytes_gb`, each tenant's queue's disk budget (hot-reloadable); the SSE gap window inside it is [`stream.gap_window_minutes`](/settings-directory#streaming). - [Deployment → Persistent Storage](/deployment#persistent-storage-required-for-containers) — `data_dir` must resolve to a host-backed volume. - [Ingest Pipeline → Backpressure and durability knobs](/ingest-pipeline#backpressure-and-durability-knobs) — the worker-side ack cost and the in-flight backpressure layers. diff --git a/docs/src/content/docs/ingest-pipeline.md b/docs/src/content/docs/ingest-pipeline.md index 870133cac..e602e0cfb 100644 --- a/docs/src/content/docs/ingest-pipeline.md +++ b/docs/src/content/docs/ingest-pipeline.md @@ -22,15 +22,15 @@ The pipeline is **insert-only**. (Upgrading across the v2 envelope? [Drain the q ## High-level shape -One process consumes a single durable JetStream consumer and fans events out to a goroutine per tenant table — the tenant is the subject's leading token. Each tenant's table batches independently and POSTs to ClickHouse over the HTTP interface (`JSONCompactEachRow`). On a bulk-insert failure the batch is re-inserted row by row, so a single poison row can't sink it: clean rows ack, and only the rows that fail again go to the dead-letter stream. A batch whose tenant has no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips that retry, which no row of it could pass, and meets the dead-letter switch once, whole; a tenant no longer served has no switch to read, so its batch is parked. An envelope the worker cannot *read* — malformed JSON, an unknown row `format` (what a pre-v2 message looks like), or columns and a row that don't pair — never reaches a table loop at all: `parseMsg` parks it on the same dead-letter stream, or, where the DLQ is off for the table, acks and drops it rather than redelivering a message that can never insert. A separate sweeper reclaims stream storage. +Each tenant's events are queued on a JetStream stream of its own. One process holds one durable consumer on each tenant's stream, delivered into one handler, and fans events out to a goroutine per tenant table — the tenant is the subject's leading token. Each tenant's table batches independently and POSTs to ClickHouse over the HTTP interface (`JSONCompactEachRow`). On a bulk-insert failure the batch is re-inserted row by row, so a single poison row can't sink it: clean rows ack, and only the rows that fail again go to the dead-letter stream. A batch whose tenant has no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips that retry, which no row of it could pass, and meets the dead-letter switch once, whole; a tenant no longer served has no switch to read, so its batch is parked. An envelope the worker cannot *read* — malformed JSON, an unknown row `format` (what a pre-v2 message looks like), or columns and a row that don't pair — never reaches a table loop at all: `parseMsg` parks it on the same dead-letter stream, or, where the DLQ is off for the table, acks and drops it rather than redelivering a message that can never insert. A separate sweeper reclaims stream storage. ```mermaid flowchart LR API["POST /v1/ingest"] -->|"publish ingest.TENANT.TABLE"| Stream subgraph NATS["Embedded NATS JetStream (in-process)"] - Stream["WAVEHOUSE stream
all ingest subjects
LimitsPolicy + DiscardNew"] - Cons["buffer-consumer
(durable, pull)"] + Stream["INGEST_TENANT stream, one per tenant
ingest.TENANT.>
LimitsPolicy + DiscardNew"] + Cons["buffer-consumer
(durable, pull, one per tenant stream)"] Stream --> Cons end @@ -46,7 +46,7 @@ flowchart LR TLa -->|"JSONCompactEachRow POST"| CH[("ClickHouse")] TLb --> CH TLc --> CH - TLa -.->|"poison rows"| DLQ["WAVEHOUSE_DLQ
dlq.TENANT.TABLE"] + TLa -.->|"poison rows"| DLQ["DLQ_TENANT stream
dlq.TENANT.TABLE"] D -.->|"unreadable envelope"| DLQ Sweep["Active Sweeper"] -.->|"reads AckFloor, purges"| Stream @@ -204,7 +204,7 @@ Messages still sitting in `msgChan` or the consumer's prefetch buffer at shutdow Delivery can end underneath a running worker: the durable consumer is deleted, or the MQ connection closes. The broker client reports that only through an asynchronous error callback and then stops delivering — no message ever arrives to say so, so a loop that only watches `msgChan` would wait forever while the API kept accepting events nothing writes. `mq.Consumer.Consume` therefore returns a `failed` channel next to `stop` (`mq.ErrDeliveryEnded`, wrapping the broker's reason), and `dispatchLoop` selects on it beside `ctx.Done()` and `msgChan`. On a failure it runs the same bottom-up drain as a shutdown — the rows already in hand are flushed and acked, not abandoned — and then reports the error on the worker's own `failed` channel. A consumer that cannot start at all takes the same path. -The worker does not try to revive the consumer. The app's ingest-worker component returns the error from `app.Run`, which stops every other component and exits non-zero, the same way any failed component does; the supervisor's restart recreates the durable consumer at boot, and everything unacked is redelivered (at-least-once). Passing conditions the client also reports through that callback (a missed heartbeat, a leadership change) are logged at `WARN` and do not end the worker. With the embedded broker (`DontListen`, no external client that could delete the durable) this path is hard to reach today; it matters once a remote broker or per-tenant consumers exist. +The worker does not try to revive the consumer. The app's ingest-worker component returns the error from `app.Run`, which stops every other component and exits non-zero, the same way any failed component does; the supervisor's restart recreates the durable consumer at boot, and everything unacked is redelivered (at-least-once). Passing conditions the client also reports through that callback (a missed heartbeat, a leadership change) are logged at `WARN` and do not end the worker. With the embedded broker (`DontListen`, no external client that could delete a durable) this path is hard to reach; the likeliest way in is a tenant's queue, opened at runtime, that the consumer cannot join. It matters more once a remote broker exists. ## Backpressure and durability knobs @@ -212,23 +212,23 @@ Several layers throttle the pipeline, inner to outer: 1. **`batch`** flushes at `maxBatch` rows or `maxWait`. 2. **`msgChan`** (cap `maxBatch*2`) — when full, the consume callback blocks and delivery pauses. -3. **`pullMaxMessages`** — nats.go's client-side prefetch buffer in front of `msgChan`. -4. **`maxAckPending`** — the server suspends delivery once this many messages are delivered-but-unacked. The outermost in-memory bound. -5. **`MaxBytes` + `DiscardNew`** on the stream (`mq.max_bytes_gb` in the [settings directory](/settings-directory#message-queue), resized in place on reload) — when disk fills (e.g. ClickHouse is down so nothing acks/purges), new publishes are rejected and the API returns 503. +3. **`pullMaxMessages`** — nats.go's client-side prefetch buffer in front of `msgChan`, shared by the tenants' streams (at least one message each). +4. **`maxAckPending`** — the server suspends a tenant's delivery once this many of its messages are delivered-but-unacked; no other tenant's delivery waits on it. The outermost in-memory bound. +5. **`MaxBytes` + `DiscardNew`** on each tenant's stream (its `mq.max_bytes_gb` in the [settings directory](/settings-directory#message-queue), resized in place on reload) — when it fills (e.g. ClickHouse is down so nothing acks/purges), that tenant's new publishes are rejected and the API returns 503. | Knob | Default | Meaning / invariant | | --- | --- | --- | | `maxBatch` | 500 | rows that trigger a flush (soft — coalescing can exceed it) | | `maxWait` | 5s | max time a row waits before its batch flushes | | `ackWait` | 60s | server redelivery timeout; **must exceed `maxWait` + flush time** or in-flight rows get redelivered → duplicate inserts | -| `pullMaxMessages` | 500 | client prefetch; keep `<= maxAckPending` | -| `maxAckPending` | 10,000 | server cap on unacked messages (backpressure) | +| `pullMaxMessages` | 500 | client prefetch, shared by the tenants' streams; keep `<= maxAckPending` | +| `maxAckPending` | 10,000 | server cap on a tenant's unacked messages (backpressure) | `DoubleAck` is used (not fire-and-forget `Ack`) because acking is what records "this data is durably in ClickHouse." With the embedded server's `SyncAlways`, every ack is an fsync and therefore *slow*, which is exactly why acks run in the background (`ackWg`) off the insert path. ## The Active Sweeper -The worker advances the consumer's `AckFloor` by acking; the sweep observes it to decide what is safe to purge. They never call each other — the consumer's `AckFloor` is their only contract. The sweeper (`internal/ingest`) owns the schedule and the window: each tick it calls `mq.Purger.PurgeAcked(buffer-consumer, now − gap window)`, where the window is the longest `stream.gap_window_minutes` among the tenants being served — every tenant's events share one stream and a purge is one bound over it, so purging less is the safe direction until each tenant has its own stream. The steps after the tick below are the embedded broker's implementation of that call. +The worker advances the consumer's `AckFloor` by acking; the sweep observes it to decide what is safe to purge. They never call each other — the consumer's `AckFloor` is their only contract. The sweeper (`internal/ingest`) owns the schedule and the window: each tick it calls `mq.Purger.PurgeAcked(buffer-consumer, cutoffs)` with each served tenant's cutoff at now − its own `stream.gap_window_minutes`; a tenant no longer served — its folder removed or rejected — is given none, and keeps none of the history it has acknowledged. The steps after the tick below are the embedded broker's implementation of that call, run on each tenant's stream at that tenant's cutoff. ```mermaid flowchart TD @@ -236,7 +236,7 @@ flowchart TD Read --> Gap["binary-search the gap-window sequence"] Gap --> Target["target = MIN(ackFloor + 1, gapSeq)"] Target --> Purge["stream.Purge below target"] - Purge -->|"deletes msgs that are BOTH
written to ClickHouse AND past the gap window"| Stream[("WAVEHOUSE stream")] + Purge -->|"deletes msgs that are BOTH
written to ClickHouse AND past the gap window"| Stream[("INGEST_TENANT stream")] ``` `MIN(ackFloor+1, gapSeq)` is the safety argument: never purge past what is in ClickHouse, and never past the SSE replay window. If ClickHouse is down the `AckFloor` stops advancing, purging freezes, and the stream fills toward `MaxBytes` — backpressure by construction. The sweeper is one of `app.Run`'s components (`Sweeper.Start` blocks until the run context is canceled), but an interrupted sweep is harmless and idempotent, so it returns on `ctx.Done()` with no drain of its own — unlike the worker's bounded `stopFunc`. @@ -248,7 +248,7 @@ Today this is a **single-process** design (embedded, in-process NATS — the "co ```mermaid flowchart TD subgraph Cluster["Clustered NATS (Replicas: 3)"] - S["WAVEHOUSE stream"] + S["one shared ingest stream"] end S --> P0["partition 0"] S --> P1["partition 1"] diff --git a/docs/src/content/docs/sdk/admin.md b/docs/src/content/docs/sdk/admin.md index ae1635f78..59db0e5a0 100644 --- a/docs/src/content/docs/sdk/admin.md +++ b/docs/src/content/docs/sdk/admin.md @@ -65,6 +65,13 @@ const { data } = await wh.dlq.list(); const { data } = await wh.dlq.table('clicks'); ``` +Each tenant has a dead-letter queue of its own, and the calls read tenant `0`'s without `tenant`. Over [a nested settings directory](/deployment#the-nested-settings-directory), pass `tenant` to read another's — a tenant whose folder was rejected or removed included, since its queue is kept — with the [operator key](/api#authentication), as for the schema reads above. A tenant with no dead-letter queue is a `404`: + +```ts +const { data } = await wh.dlq.list({ tenant: 'acme' }); +const { data: clicks } = await wh.dlq.table('clicks', { tenant: 'acme' }); +``` + `wh.dlq.stream()` exists in the API but is **not yet functional**: there is no server-side DLQ stream today (the SSE bridge only carries `ingest.>` subjects), so it connects and receives no events rather than failing. Live DLQ streaming is tracked in [#197](https://github.com/Wave-RF/WaveHouse/issues/197). --- diff --git a/docs/src/content/docs/sdk/reference.md b/docs/src/content/docs/sdk/reference.md index 56fc09d9e..af0cddef0 100644 --- a/docs/src/content/docs/sdk/reference.md +++ b/docs/src/content/docs/sdk/reference.md @@ -104,8 +104,8 @@ createClient(config) → WaveHouseClient ├── .settings (admin) │ └── .reload(opts?) → Promise> ├── .dlq (admin) -│ ├── .list() → Promise> -│ ├── .table(name) → Promise> +│ ├── .list(opts?) → Promise> +│ ├── .table(name, opts?) → Promise> │ └── .stream() → StreamController // not yet functional server-side — #197 └── .sys └── .health() → Promise> diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index 9555bf797..c0e7a19b7 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -31,7 +31,7 @@ A reload that fails validation is logged (and reported by the endpoint) and the "Previous good settings" is the in-memory snapshot of the running process, nothing more: there is no persisted copy of the files. A restart re-validates the directory from scratch and refuses to start on the same findings the reload rejected, so bad files never survive a restart silently — fix them (or run `wavehouse validate`) before bouncing the server. -A directory that holds one folder per tenant instead of the four files is [a nested settings directory](/deployment#the-nested-settings-directory): each folder is everything this page describes, but it is not watched, a rejected folder stops its tenant being served rather than keeping the previous settings, and the keys the whole process shares are not read from the tenant's own folder: they come from tenant `0`'s, bar the two that weigh every tenant being served: the SSE keepalive, which follows the shortest `stream.keepalive_interval` among them, and the sweeper's gap window, the longest `stream.gap_window_minutes` among them — that section lists which keys. +A directory that holds one folder per tenant instead of the four files is [a nested settings directory](/deployment#the-nested-settings-directory): each folder is everything this page describes, but it is not watched, a rejected folder stops its tenant being served rather than keeping the previous settings, and the keys the whole process shares are not read from the tenant's own folder: they come from tenant `0`'s, bar the one that weighs every tenant being served: the SSE keepalive, which follows the shortest `stream.keepalive_interval` among them — that section lists which keys. Every adoption — boot and every reload — goes through the same `Validate`, so the policy, the roles, and the pipes are checked with the current rules each time they are read; there is no stored copy that can skip validation. All four files are adopted as one snapshot: a request is evaluated against the policy, pipes, and tunables of a single adoption, never a mix. @@ -124,7 +124,7 @@ The tenant tunables. Every key is required (a missing one is a validation error) | `dedupe.id_field` | `event_id` | Dedup key field — see [Deduplication](#deduplication). | | `dedupe.require_id` | `false` | Reject rows missing the id field — see [Deduplication](#deduplication). | | `dedupe.tables.
.{id_field, require_id}` | `{}` | Optional per-table overrides; each entry overrides only the fields it names and inherits the rest. | -| `dlq.enabled` | `true` | Park poison rows — those that still fail after row-by-row isolation, and every row of a batch whose tenant has no ClickHouse connection — on the `WAVEHOUSE_DLQ` stream (`false`: leave them unacked for redelivery — except an envelope the worker cannot read, which is dropped and counted) — see [Dead Letter Queue](#dead-letter-queue). | +| `dlq.enabled` | `true` | Park poison rows — those that still fail after row-by-row isolation, and every row of a batch whose tenant has no ClickHouse connection — on the tenant's dead-letter stream (`DLQ_{tenant}`) (`false`: leave them unacked for redelivery — except an envelope the worker cannot read, which is dropped and counted) — see [Dead Letter Queue](#dead-letter-queue). | | `dlq.tables.
.enabled` | `{}` | Optional per-table override of the switch. | | `query.timestamp_bucket_seconds` | `60` | Bucket (seconds, `>= 0`) that a structured query's relative time range is truncated to, so near-identical queries share a cache entry; `0` disables bucketing. Read per query. | | `query.default_max_rows` | `10000` | Fallback result `LIMIT` (`>= 1`) applied to a structured query when the caller and policy specify none. A result-**shaping** default, not a resource limit — server-wide limits (memory, rows scanned, execution time) belong in ClickHouse, see [Server-side resource limits](/configuration#server-side-resource-limits). | @@ -132,7 +132,7 @@ The tenant tunables. Every key is required (a missing one is a validation error) | `stream.keepalive_interval` | `30` | Seconds (`>= 1`) a quiet `GET /v1/stream` connection may go without a write before the server sends a `:` keepalive comment — keep it under your proxy's idle timeout; see [Streaming](#streaming). | | `stream.keepalive_buckets` | `3` | Load-spreading (`>= 1`): connections are spread across N buckets so each tick nudges ~1/N of live streams. Most deployments leave it. | | `stream.gap_window_minutes` | `15` | Minutes (`>= 0`) of written-to-ClickHouse history the Active Sweeper keeps in NATS for `Last-Event-ID` gap-fill; applies from the next sweep. | -| `mq.max_bytes_gb` | `50` | Disk budget (GB, `>= 1`) for the embedded NATS `WAVEHOUSE` ingest stream; the `WAVEHOUSE_DLQ` stream gets a tenth of it. A reload updates the live streams in place. See [Message Queue](#message-queue). | +| `mq.max_bytes_gb` | `50` | Disk budget (GB, `>= 1`) for the tenant's embedded NATS ingest stream (`INGEST_{tenant}`); its dead-letter stream (`DLQ_{tenant}`) gets a tenth of it. A reload updates the live streams in place. See [Message Queue](#message-queue). | | `cors.allowed_origins` | `["*"]` | Allowed CORS origins, applied per request. `"*"` allows any browser origin. WaveHouse is a Bearer-token API — `Access-Control-Allow-Credentials` is intentionally never sent, so this allowlist controls *which origins can read responses*, not cookie scope. Tighten to your frontend's exact origin(s) in production (e.g. `["https://dashboard.example.com", "http://localhost:3000"]`). An empty list `[]` denies every browser origin (no `Access-Control-Allow-Origin` is ever sent); `"*"` is the only allow-all spelling. Over [a nested settings directory](/deployment#the-nested-settings-directory) each tenant's list decorates its own responses, the preflight included; which list answers a preflight, the tenant-exempt routes, and a refused request is [spelled out there](/deployment#multi-tenant-deployments). | ```json @@ -212,16 +212,16 @@ The `auth` block is the verifier wiring, minus the secrets. `jwks_url` (absolute A failed batch insert is retried row by row; a row that fails again on its own is a poison row. A batch whose tenant has no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the retry, which no row of it could pass, and every row of it is a poison row. `dlq.enabled` (seed default `true`) decides what happens to it, resolved per table (`dlq.tables.
.enabled` → global) at the moment of the failure, so a reload applies to the next poison row: -- `true` — the row is published to the `WAVEHOUSE_DLQ` NATS stream under `dlq.{tenant}.{table}` (`0` for a directory that holds the four files) with the failure in its headers, and its original is acked. Inspect it with `GET /v1/ops/dlq/stats` (admin-only). +- `true` — the row is published to the tenant's dead-letter stream (`DLQ_{tenant}`) under `dlq.{tenant}.{table}` (`0` for a directory that holds the four files) with the failure in its headers, and its original is acked. Inspect it with `GET /v1/ops/dlq/stats` (admin-only; `?tenant=` names the tenant). - `false` — the row is left unacked, so NATS redelivers it and it retries until it inserts or the switch is flipped back. For every row the worker **can read**, nothing is ever dropped either way — the choice is *park it* versus *keep retrying*. **One exception, new in this release:** an envelope the worker cannot read *at all* — malformed JSON, an unknown `format` (what a pre-v2 in-flight message looks like), or `columns` and `row` that do not pair — can never insert, so redelivering it forever would wedge the consumer. With the DLQ off for the table it is acked and **dropped**, logged at `ERROR` and counted by `wavehouse_ingest_poison_total` with `disposition="dropped"` (also labeled by `table` and `reason`; an envelope parked on the DLQ carries `disposition="parked"`). See [Ingest Pipeline](/ingest-pipeline) — and drain the ingest queue before upgrading. -For a tenant no longer served — its folder removed or rejected — there is no switch to read: its rows are always parked, so none of them sits unacked in the shared ingest queue, where it would stop the [Active Sweeper](/ingest-pipeline#the-active-sweeper) purging it. +For a tenant no longer served — its folder removed or rejected — there is no switch to read: its rows are always parked, so none of them sits unacked in its ingest queue, redelivered for as long as the tenant is away and stopping the [Active Sweeper](/ingest-pipeline#the-active-sweeper) purging that queue. -The `WAVEHOUSE_DLQ` stream always exists (an empty stream costs nothing) and the stats endpoint is always registered — the switch is purely behavioral, which is what makes it safe to reload. +A tenant's dead-letter stream exists from the moment the tenant is first served (an empty stream costs nothing) and the stats endpoint is always registered — the switch is purely behavioral, which is what makes it safe to reload. ## Message Queue -- `mq.max_bytes_gb` (seed default `50`) — disk budget for the embedded JetStream `WAVEHOUSE` stream that buffers ingested events until the worker writes them to ClickHouse; the `WAVEHOUSE_DLQ` stream gets a tenth of it. The stream runs `DiscardNew`, so when it's full new publishes are rejected and `POST /v1/ingest` returns `503` — [backpressure by construction](/ingest-pipeline#backpressure-and-durability-knobs). A reload updates both streams' limits in place without touching what's buffered: growing takes effect immediately; shrinking below what's currently on disk makes the stream refuse new publishes until the worker drains it back under the limit — nothing already accepted is dropped. If NATS rejects the update, the rest of the reload is still adopted, the failure is logged, and the next reload retries it. The two streams are resized as a pair: a failed DLQ resize undoes the ingest one so both stay on the previous budget, but if that undo fails too the ingest stream keeps the new limit and the DLQ the previous one until a later reload succeeds — the log line says which happened. Size it from [Durability & Storage](/durability). +- `mq.max_bytes_gb` (seed default `50`) — disk budget for the tenant's embedded JetStream ingest stream (`INGEST_{tenant}`), which buffers its ingested events until the worker writes them to ClickHouse; its dead-letter stream (`DLQ_{tenant}`) gets a tenth of it. Each tenant's pair of streams is its own, opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed. The ingest stream runs `DiscardNew`, so when it's full new publishes are rejected and `POST /v1/ingest` returns `503` for that tenant alone — [backpressure by construction](/ingest-pipeline#backpressure-and-durability-knobs). A reload updates both streams' limits in place without touching what's buffered: growing takes effect immediately; shrinking below what's currently on disk makes the ingest stream refuse new publishes until the worker drains it back under the limit — nothing already accepted is dropped — and a dead-letter stream holding more than a tenth of the new budget is kept at what it holds rather than shrunk, since shrinking it would delete its oldest parked rows; that is logged, and the stream then makes room for each new row by dropping its oldest, as a full one always does. If NATS rejects the update, the rest of the reload is still adopted, the failure is logged, and the next reload retries it. A queue NATS will not open at all refuses boot, like every other store; over [a nested directory](/deployment#the-nested-settings-directory) it costs that tenant alone, at boot or on reload — its ingest answers `503`, each publish and each reload trying the queue again — while every other tenant carries on. The two streams are resized as a pair: a failed DLQ resize undoes the ingest one so both stay on the previous budget, but if that undo fails too the ingest stream keeps the new limit and the DLQ the previous one until a later reload succeeds — the log line says which happened. Nothing checks the budget against the disk — neither one tenant's nor what the tenants' add up to — so a queue fills until its budget or the disk runs out, whichever comes first; size them together from [Durability & Storage](/durability). ## Streaming @@ -229,4 +229,4 @@ The `WAVEHOUSE_DLQ` stream always exists (an empty stream costs nothing) and the - `stream.keepalive_interval` (seed default `30`) — seconds a quiet connection may go without a write before the server sends a `:` keepalive comment. It exists to stay under whatever idle timeout sits between WaveHouse and the client; the default clears the common 55–60s proxy windows with margin, and a tighter edge (Azure Application Gateway 20s, CloudFront 30s) wants a lower value — see [Behind a reverse proxy → Idle timeouts](/reverse-proxy#idle-timeouts-by-provider). A reload rebuilds the keepalive wheel in place: live connections stay open and are redistributed across the new ring, each getting at most one full new period before its next keepalive. - `stream.keepalive_buckets` (seed default `3`) — spreads the keepalive writes across the interval (one bucket fires every `keepalive_interval ÷ keepalive_buckets`) so the server nudges ~1/N of connections per tick instead of all at once. It changes only how the writes are spread in time, never the period. -- `stream.gap_window_minutes` (seed default `15`) — minutes of already-written-to-ClickHouse history the Active Sweeper keeps in NATS so a reconnecting client's `Last-Event-ID` replay can bridge the gap; a drop longer than this resumes with a hole. Bounded by the stream's disk budget, [`mq.max_bytes_gb`](#message-queue). A reload applies from the next sweep (every minute). +- `stream.gap_window_minutes` (seed default `15`) — minutes of already-written-to-ClickHouse history the Active Sweeper keeps in NATS so a reconnecting client's `Last-Event-ID` replay can bridge the gap; a drop longer than this resumes with a hole. Bounded by the tenant's disk budget, [`mq.max_bytes_gb`](#message-queue). A reload applies from the next sweep (every minute). diff --git a/docs/src/content/docs/why-wavehouse.md b/docs/src/content/docs/why-wavehouse.md index a5b63e219..26ac9d70b 100644 --- a/docs/src/content/docs/why-wavehouse.md +++ b/docs/src/content/docs/why-wavehouse.md @@ -53,7 +53,7 @@ Even if you remember to batch client-side, a naive ingest path has no safe way t - **No backpressure channel.** If the merger falls behind, ClickHouse raises an error at the *next* insert. The client has already left. - **No DLQ.** Bad events that fail to insert are either lost or logged into ClickHouse's error log. Good luck replaying yesterday's dropped rows. -WaveHouse fixes all three at the gateway: validates every payload against the real `system.columns` schema before accepting, returns `503 Service Unavailable` with a `Retry-After` header when the NATS WAL fills, and routes failed batch inserts to a dedicated `WAVEHOUSE_DLQ` stream you can inspect via `GET /v1/ops/dlq/stats`. +WaveHouse fixes all three at the gateway: validates every payload against the real `system.columns` schema before accepting, returns `503 Service Unavailable` with a `Retry-After` header when the NATS WAL fills, and routes failed batch inserts to a dedicated dead-letter stream, one per tenant, you can inspect via `GET /v1/ops/dlq/stats`. ### No real-time push @@ -156,7 +156,7 @@ flowchart TB | Real-time push | WebSocket service + bridge from Kafka | Built in (`/v1/stream`) | | Schema validation | Custom code in ingest API | Built in (discovers `system.columns`) | | Row/column access control | Custom middleware or a dedicated service | Built in (Hasura-style, JWT-driven) | -| Dead letter queue | Custom retry + dead topic on Kafka | Built in (`WAVEHOUSE_DLQ`) | +| Dead letter queue | Custom retry + dead topic on Kafka | Built in (a dead-letter stream per tenant) | | Client SDK | Each team writes one | `@wavehouse/sdk` (TypeScript, one dependency, codegen) | The DIY path works — big teams run it — but the ops cost is not small. You're paying for a Kafka cluster (or Confluent bill), a second service you wrote from scratch, and all the debugging hours when the batching consumer stalls at 3 a.m. @@ -190,7 +190,7 @@ Tinybird wins on "zero ops to start." WaveHouse wins on "own your data plane and | Self-hosted | ✓ | ✓ | ✗ | ✓ | | Handles N-row inserts safely | ✗ merge blowup | ✓ via Kafka | ✓ | ✓ native | | Schema validation at the edge | ✗ | Custom | ✓ | ✓ (discovers schema) | -| Dead letter queue | ✗ | Custom | Partial | ✓ `WAVEHOUSE_DLQ` | +| Dead letter queue | ✗ | Custom | Partial | ✓ dead-letter stream per tenant | | Backpressure (503 + Retry-After) | ✗ | Custom | ✓ | ✓ | | Idempotent ingest (dedup by ID) | ✗ | Custom | ✓ | ✓ optional | | Real-time push (SSE) | ✗ | Custom service | ✗ | ✓ native, gap-fill | @@ -221,7 +221,7 @@ flowchart TB NATS --> BC["Buffer consumer
5-second batches"]:::wh BC --> CH[("ClickHouse")]:::store - BC -. "on failure" .-> DLQ["WAVEHOUSE_DLQ"]:::fail + BC -. "on failure" .-> DLQ["dead-letter stream"]:::fail ``` **Query path with tiered cache:** diff --git a/internal/api/dlq.go b/internal/api/dlq.go index 9de69ad56..267d2dc47 100644 --- a/internal/api/dlq.go +++ b/internal/api/dlq.go @@ -7,6 +7,7 @@ import ( "net/http" "github.com/Wave-RF/WaveHouse/internal/mq" + "github.com/Wave-RF/WaveHouse/internal/tenant" ) // DLQHandler exposes Dead Letter Queue statistics. @@ -18,18 +19,30 @@ func NewDLQHandler(stats mq.DeadLetterStats) *DLQHandler { return &DLQHandler{Counts: stats} } -// Stats returns per-table message counts on the dead-letter queue. -// Supports optional ?table= query parameter to filter by table name. +// Stats returns per-table message counts on one tenant's dead-letter queue: +// the tenant ?tenant= names (read strictly, as every ops read does — +// opsTenant), tenant.Default without it. The queue is the MQ's, not the +// settings', so it is looked up there: a tenant whose folder was rejected or +// removed is read like one being served, for as long as its queue is kept, +// and an id with no queue is a 404. Supports optional ?table= query parameter +// to filter by table name. func (h *DLQHandler) Stats(w http.ResponseWriter, r *http.Request) { - counts, err := h.Counts.DeadLetterCounts(r.Context(), r.URL.Query().Get("table")) + id, named, ok := opsTenant(w, r) + if !ok { + return + } + if !named { + id = tenant.Default + } + counts, err := h.Counts.DeadLetterCounts(r.Context(), id, r.URL.Query().Get("table")) if err != nil { - if !errors.Is(err, mq.ErrNoDeadLetterQueue) { - slog.ErrorContext(r.Context(), "dlq stats failed", "error", err) - writeJSONError(w, http.StatusInternalServerError, "stream info failed") + if errors.Is(err, mq.ErrNoDeadLetterQueue) { + writeJSONError(w, http.StatusNotFound, "no dead-letter queue for tenant: "+id.String()) return } - // No dead-letter queue: nothing can have been parked. - counts = mq.DeadLetterCounts{Tables: map[string]uint64{}} + slog.ErrorContext(r.Context(), "dlq stats failed", "tenant", id, "error", err) + writeJSONError(w, http.StatusInternalServerError, "stream info failed") + return } w.Header().Set("Content-Type", "application/json") diff --git a/internal/api/dlq_test.go b/internal/api/dlq_test.go index 4023e3541..d814cd8be 100644 --- a/internal/api/dlq_test.go +++ b/internal/api/dlq_test.go @@ -15,56 +15,53 @@ import ( "github.com/stretchr/testify/require" ) -// parkedMsg is a message as the ingest worker would hand it to the DLQ. -func parkedMsg(table string) *mq.Message { +// parkedMsg is a message as the ingest worker would hand it to the DLQ, +// parked under tenant id's table. +func parkedMsg(id tenant.ID, table string) *mq.Message { return (&testutil.MockMessage{ - MsgTopic: mq.Topic{Tenant: tenant.Default, Table: table}, + MsgTopic: mq.Topic{Tenant: id, Table: table}, MsgData: []byte(`{"table_name":"` + table + `"}`), }).Message() } -func TestDLQStats_EmptyWhenNoStream(t *testing.T) { - // The embedded MQ always has a dead-letter queue, so its absence comes - // from a mock. - handler := NewDLQHandler(&testutil.MockDeadLetterStats{Err: mq.ErrNoDeadLetterQueue}) - - req := httptest.NewRequestWithContext(context.Background(), http.MethodGet, "/v1/ops/dlq/stats", nil) +// dlqStats serves GET /v1/ops/dlq/stats with query through handler. +func dlqStats(t *testing.T, handler *DLQHandler, query string) *httptest.ResponseRecorder { + t.Helper() + req := httptest.NewRequestWithContext(t.Context(), http.MethodGet, "/v1/ops/dlq/stats"+query, nil) rec := httptest.NewRecorder() - handler.Stats(rec, req) + return rec +} - assert.Equal(t, http.StatusOK, rec.Code) - - var resp map[string]any - require.NoError(t, json.Unmarshal(rec.Body.Bytes(), &resp)) - - tables, ok := resp["tables"].(map[string]any) - require.True(t, ok) - assert.Empty(t, tables) - assert.Equal(t, float64(0), resp["total"]) +// A tenant with no dead-letter queue — never given one on this data +// directory, or an id nobody uses — is a 404 that names it, not an empty +// count that would read as "nothing parked" for a typo. +func TestDLQStats_NoQueueIs404(t *testing.T) { + // The embedded MQ opens a served tenant's queue at boot, so a queue's + // absence comes from a mock. + stats := &testutil.MockDeadLetterStats{Err: mq.ErrNoDeadLetterQueue} + + rec := dlqStats(t, NewDLQHandler(stats), "?tenant=acmee") + assert.Equal(t, http.StatusNotFound, rec.Code) + assert.Contains(t, rec.Body.String(), "no dead-letter queue for tenant: acmee") + testutil.AssertJSONErrorResponse(t, rec) + assert.Equal(t, tenant.ID("acmee"), stats.Tenant) } func TestDLQStats_ReturnsCorrectCounts(t *testing.T) { - dir := t.TempDir() - emb, err := mq.NewEmbedded(dir, 1024*1024) - require.NoError(t, err) - defer func() { _ = emb.Close() }() + emb := testutil.NewEmbeddedMQ(t, 1024*1024) ctx := context.Background() // Park messages on the dead-letter queue. for i := 0; i < 3; i++ { - require.NoError(t, emb.DeadLetter(ctx, parkedMsg("events"))) + require.NoError(t, emb.DeadLetter(ctx, parkedMsg(tenant.Default, "events"))) } for i := 0; i < 2; i++ { - require.NoError(t, emb.DeadLetter(ctx, parkedMsg("users"))) + require.NoError(t, emb.DeadLetter(ctx, parkedMsg(tenant.Default, "users"))) } - handler := NewDLQHandler(emb) - req := httptest.NewRequestWithContext(context.Background(), http.MethodGet, "/v1/ops/dlq/stats", nil) - rec := httptest.NewRecorder() - - handler.Stats(rec, req) + rec := dlqStats(t, NewDLQHandler(emb), "") assert.Equal(t, http.StatusOK, rec.Code) @@ -78,21 +75,20 @@ func TestDLQStats_ReturnsCorrectCounts(t *testing.T) { assert.Equal(t, float64(5), resp["total"]) } +func TestDLQStats_EmptyBeforeAnyFailure(t *testing.T) { + rec := dlqStats(t, NewDLQHandler(testutil.NewEmbeddedMQ(t, 1024*1024)), "") + assert.Equal(t, http.StatusOK, rec.Code) + assert.JSONEq(t, `{"tables":{},"total":0}`, rec.Body.String()) +} + func TestDLQStats_SingleTable(t *testing.T) { - dir := t.TempDir() - emb, err := mq.NewEmbedded(dir, 1024*1024) - require.NoError(t, err) - defer func() { _ = emb.Close() }() + emb := testutil.NewEmbeddedMQ(t, 1024*1024) ctx := context.Background() - require.NoError(t, emb.DeadLetter(ctx, parkedMsg("orders"))) + require.NoError(t, emb.DeadLetter(ctx, parkedMsg(tenant.Default, "orders"))) - handler := NewDLQHandler(emb) - req := httptest.NewRequestWithContext(context.Background(), http.MethodGet, "/v1/ops/dlq/stats", nil) - rec := httptest.NewRecorder() - - handler.Stats(rec, req) + rec := dlqStats(t, NewDLQHandler(emb), "") assert.Equal(t, http.StatusOK, rec.Code) @@ -105,33 +101,61 @@ func TestDLQStats_SingleTable(t *testing.T) { } func TestDLQStats_BrokerFailureIsAnError(t *testing.T) { - handler := NewDLQHandler(&testutil.MockDeadLetterStats{Err: errors.New("broker unavailable")}) - - req := httptest.NewRequestWithContext(context.Background(), http.MethodGet, "/v1/ops/dlq/stats", nil) - rec := httptest.NewRecorder() - - handler.Stats(rec, req) - + rec := dlqStats(t, NewDLQHandler(&testutil.MockDeadLetterStats{Err: errors.New("broker unavailable")}), "") assert.Equal(t, http.StatusInternalServerError, rec.Code, "a failed read is not an empty queue") } func TestDLQStats_PassesTheTableFilter(t *testing.T) { - emb, err := mq.NewEmbedded(t.TempDir(), 1024*1024) - require.NoError(t, err) - defer func() { _ = emb.Close() }() + emb := testutil.NewEmbeddedMQ(t, 1024*1024) ctx := context.Background() - require.NoError(t, emb.DeadLetter(ctx, parkedMsg("default.orders"))) - require.NoError(t, emb.DeadLetter(ctx, parkedMsg("users"))) + require.NoError(t, emb.DeadLetter(ctx, parkedMsg(tenant.Default, "default.orders"))) + require.NoError(t, emb.DeadLetter(ctx, parkedMsg(tenant.Default, "users"))) - handler := NewDLQHandler(emb) - req := httptest.NewRequestWithContext(ctx, http.MethodGet, "/v1/ops/dlq/stats?table=default.orders", nil) - rec := httptest.NewRecorder() - - handler.Stats(rec, req) + rec := dlqStats(t, NewDLQHandler(emb), "?table=default.orders") var resp map[string]any require.NoError(t, json.Unmarshal(rec.Body.Bytes(), &resp)) assert.Equal(t, map[string]any{"default.orders": float64(1)}, resp["tables"]) assert.Equal(t, float64(2), resp["total"]) } + +// ?tenant= reads that tenant's queue alone, and no parameter reads tenant 0's +// — the ops-read convention. The handler asks the MQ, not the settings, so a +// tenant the settings no longer serve (here, none at all) is read by name +// for as long as its queue is kept. +func TestDLQStats_ReadsTheNamedTenantsQueue(t *testing.T) { + emb := testutil.NewEmbeddedMQ(t, 1024*1024, tenant.Default, "acme") + ctx := context.Background() + require.NoError(t, emb.DeadLetter(ctx, parkedMsg(tenant.Default, "events"))) + for range 2 { + require.NoError(t, emb.DeadLetter(ctx, parkedMsg("acme", "events"))) + } + handler := NewDLQHandler(emb) + + rec := dlqStats(t, handler, "?tenant=acme") + assert.Equal(t, http.StatusOK, rec.Code) + assert.JSONEq(t, `{"tables":{"events":2},"total":2}`, rec.Body.String()) + + rec = dlqStats(t, handler, "?tenant=acme&table=users") + assert.Equal(t, http.StatusOK, rec.Code) + assert.JSONEq(t, `{"tables":{},"total":2}`, rec.Body.String()) + + rec = dlqStats(t, handler, "") + assert.Equal(t, http.StatusOK, rec.Code) + assert.JSONEq(t, `{"tables":{"events":1},"total":1}`, rec.Body.String(), "no parameter reads tenant 0") + + rec = dlqStats(t, handler, "?tenant=globex") + assert.Equal(t, http.StatusNotFound, rec.Code, "a tenant with no queue") +} + +// The parameter is read strictly, like every ops read's (opsTenant): a +// query that misparses must not fall back to tenant 0's counts. +func TestDLQStats_RefusesAMalformedTenant(t *testing.T) { + for _, query := range []string{"?tenant=a.b", "?tenant=", "?tenant=a&tenant=b", "?tenant=acme;x=1"} { + stats := &testutil.MockDeadLetterStats{} + rec := dlqStats(t, NewDLQHandler(stats), query) + assert.Equal(t, http.StatusBadRequest, rec.Code, query) + assert.Empty(t, stats.Tenant, "%s: nothing is read", query) + } +} diff --git a/internal/api/ingest.go b/internal/api/ingest.go index c429aaa41..029799b6c 100644 --- a/internal/api/ingest.go +++ b/internal/api/ingest.go @@ -707,7 +707,7 @@ func (h *IngestHandler) processRecord( slog.DebugContext(ctx, "publishing event to the ingest queue", "table", table, "scope", scope) if err := h.Publisher.Publish(ctx, mq.Topic{Tenant: store.Tenant(), Table: table, Scope: scope}, payload); err != nil { if errors.Is(err, mq.ErrQueueFull) { - slog.WarnContext(ctx, "ingest queue is full", "table", table, "scope", scope) + slog.WarnContext(ctx, "ingest queue is full", "error", err, "table", table, "scope", scope) return false, nil, &requestAbort{Status: http.StatusServiceUnavailable, Message: "service unavailable", RetryAfter: "30"} } slog.ErrorContext(ctx, "failed to publish to the ingest queue", "error", err, "table", table, "scope", scope) diff --git a/internal/api/router_test.go b/internal/api/router_test.go index e0240ba6f..03a39c769 100644 --- a/internal/api/router_test.go +++ b/internal/api/router_test.go @@ -15,7 +15,6 @@ import ( "github.com/Wave-RF/WaveHouse/internal/auth" "github.com/Wave-RF/WaveHouse/internal/discovery" - "github.com/Wave-RF/WaveHouse/internal/mq" "github.com/Wave-RF/WaveHouse/internal/pipes" "github.com/Wave-RF/WaveHouse/internal/policy" "github.com/Wave-RF/WaveHouse/internal/settings" @@ -332,9 +331,7 @@ func TestNewRouter_RoutesRegistered(t *testing.T) { pub := &testutil.MockPublisher{} hub := stream.NewHub(nil, nil, nil) - emb, err := mq.NewEmbedded(t.TempDir(), 1024*1024) - require.NoError(t, err) - t.Cleanup(func() { _ = emb.Close() }) + emb := testutil.NewEmbeddedMQ(t, 1024*1024) deps := Dependencies{ Tenants: testTenants(), diff --git a/internal/app/app.go b/internal/app/app.go index f853a76b3..dc15bf5d7 100644 --- a/internal/app/app.go +++ b/internal/app/app.go @@ -18,8 +18,7 @@ // handed whole to each component's wiring function, which derives the // per-call getters the internal packages take: keyed by the request's store // for the handlers, by tenant id for the async paths (perTenant), and fixed -// to the default tenant for the process-wide resources #583 has not yet made -// per tenant (defaultSetting). +// to the default tenant for the ops gate of a flat directory (defaultSetting). package app import ( @@ -90,9 +89,9 @@ type App struct { listener net.Listener // tenants is the registry every tenant-aware path resolves through, and - // the owner of every reload. The one process-wide resource left, the MQ, - // still follows its default tenant, through defaultStore: tenant 0's - // store as of its last adoption (defaultSetting). + // the owner of every reload. defaultStore is tenant 0's store as of its + // last adoption, which the ops gate of a flat directory reads its admin + // role from (defaultSetting). tenants *settings.Registry defaultStore atomic.Pointer[settings.Store] // policies is the default tenant's policy, for the ops gate of a flat diff --git a/internal/app/app_test.go b/internal/app/app_test.go index fc5d27553..2ecd77d98 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -237,7 +237,7 @@ func TestReload_DrivesTheRegisteredHooks(t *testing.T) { a := newApp(t, cfg, Options{}) dedup := a.dedup.For(tenant.Default) require.False(t, dedup.Open()) - require.Equal(t, int64(1<<30), a.mq.MaxBytes()) + require.Equal(t, int64(1<<30), a.mq.MaxBytes(tenant.Default)) rewriteSettings(t, dir, map[string]any{ "dedupe": map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "tables": map[string]any{}}, @@ -246,8 +246,8 @@ func TestReload_DrivesTheRegisteredHooks(t *testing.T) { _, adopted := a.tenants.Reload("test") require.True(t, adopted) assert.True(t, dedup.Open(), "dedupe hook opened the store") - // How the budget is split across the MQ's queues is internal/mq's to test. - assert.Equal(t, int64(2<<30), a.mq.MaxBytes(), "mq hook applied the new byte budget") + // How the budget is split across the tenant's queues is internal/mq's to test. + assert.Equal(t, int64(2<<30), a.mq.MaxBytes(tenant.Default), "mq hook applied the new byte budget") rewriteSettings(t, dir, map[string]any{"mq": map[string]any{"max_bytes_gb": 2}}) _, adopted = a.tenants.Reload("test") @@ -397,12 +397,13 @@ func TestNew_NestedWithoutAnOperatorKeyWarnsTheOpsTreeIsClosed(t *testing.T) { }) } -// The process-wide resources follow tenant 0 alone: another tenant's reload -// never moves them, and a rejected 0 folder leaves them as they were rather -// than reconfiguring them from nothing. The dedupe stores are per tenant -// (story 7), so each follows its own folder instead — the contrast the -// same reloads show. -func TestReload_NestedHooksFollowTheDefaultTenant(t *testing.T) { +// A tenant's queue budget and dedupe store follow its own folder alone: +// another tenant's reload moves neither. A rejected or removed folder keeps +// its tenant's queue at the budget it last had — removing never touches +// data — while its dedupe store closes, its seen ids kept. CORS is read per +// request, so a lost 0 folder is felt at once on the routes that read tenant +// 0's list. +func TestReload_NestedHooksFollowEachTenant(t *testing.T) { dedupeOn := map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "tables": map[string]any{}} grown := map[string]any{"dedupe": dedupeOn, "mq": map[string]any{"max_bytes_gb": 2}} root := writeNestedSettings(t, map[string]map[string]any{ @@ -413,7 +414,8 @@ func TestReload_NestedHooksFollowTheDefaultTenant(t *testing.T) { dedup0, dedupAcme := a.dedup.For(tenant.Default), a.dedup.For("acme") require.False(t, dedup0.Open()) require.False(t, dedupAcme.Open()) - require.Equal(t, int64(1<<30), a.mq.MaxBytes()) + require.Equal(t, int64(1<<30), a.mq.MaxBytes(tenant.Default)) + require.Equal(t, int64(1<<30), a.mq.MaxBytes("acme"), "each served tenant's queue opens at boot at its own budget") // CORS is per tenant, not a hook's: a tenant route reads its own tenant's // list and the exempt routes tenant 0's (the seed's ["*"] in every folder // here), both through the registry, so a lost 0 folder is felt at once. @@ -434,26 +436,27 @@ func TestReload_NestedHooksFollowTheDefaultTenant(t *testing.T) { _, adopted := a.tenants.Reload("test") require.True(t, adopted) assert.True(t, dedupAcme.Open(), "acme's dedupe switch opens acme's own store") + assert.Equal(t, int64(2<<30), a.mq.MaxBytes("acme"), "acme's budget resizes acme's own queue") assert.False(t, dedup0.Open(), "and moves nothing of tenant 0's") - assert.Equal(t, int64(1<<30), a.mq.MaxBytes()) + assert.Equal(t, int64(1<<30), a.mq.MaxBytes(tenant.Default)) rewriteSettings(t, filepath.Join(root, "0"), grown) _, adopted = a.tenants.Reload("test") require.True(t, adopted) assert.True(t, dedup0.Open()) - assert.Equal(t, int64(2<<30), a.mq.MaxBytes()) + assert.Equal(t, int64(2<<30), a.mq.MaxBytes(tenant.Default)) rewriteSettings(t, filepath.Join(root, "0"), invalidQuery) _, adopted = a.tenants.Reload("test") require.False(t, adopted) assert.False(t, dedup0.Open(), "a rejected 0 folder closes tenant 0's own store, which answers no request now") assert.True(t, dedupAcme.Open(), "and costs acme nothing") - assert.Equal(t, int64(2<<30), a.mq.MaxBytes(), "the process-wide budget stays as tenant 0 last adopted it") + assert.Equal(t, int64(2<<30), a.mq.MaxBytes(tenant.Default), "tenant 0's queue is kept at the budget it last had") assert.Empty(t, allowOrigin("/version"), "the exempt routes read tenant 0 through the registry, which is no longer serving it") assert.Equal(t, "*", allowOrigin("/v1/health", "acme"), "acme's own routes keep acme's list") - // A removed 0 folder is the same: the registry forgets the tenant, the - // process keeps the wiring it last adopted. + // A removed 0 folder is the same: the registry forgets the tenant, and + // its queue stays at the budget it last had. require.NoError(t, os.RemoveAll(filepath.Join(root, "0"))) _, adopted = a.tenants.Reload("test") require.True(t, adopted) @@ -461,7 +464,8 @@ func TestReload_NestedHooksFollowTheDefaultTenant(t *testing.T) { require.False(t, known) assert.False(t, dedup0.Open()) assert.True(t, dedupAcme.Open()) - assert.Equal(t, int64(2<<30), a.mq.MaxBytes()) + assert.Equal(t, int64(2<<30), a.mq.MaxBytes(tenant.Default)) + assert.Equal(t, int64(2<<30), a.mq.MaxBytes("acme")) assert.Empty(t, allowOrigin("/version")) assert.Equal(t, "*", allowOrigin("/v1/health", "acme")) } @@ -682,12 +686,11 @@ func gapWindow(minutes int) map[string]any { return map[string]any{"stream": map[string]any{"keepalive_interval": 30, "keepalive_buckets": 3, "gap_window_minutes": minutes}} } -// One ingest stream holds every tenant's events and the sweeper purges below -// one sequence, so it keeps the longest gap window among the tenants being -// served — every tenant's gap-fill history is inside it (a stream per tenant -// will honor each tenant's own, #583 story 5b). A flat directory's single +// Each tenant being served keeps its own stream.gap_window_minutes, since +// each has a queue of its own; a rejected tenant is not served, so it is not +// named and keeps no history (mq.Purger.PurgeAcked). A flat directory's single // tenant gets exactly its own window. -func TestLongestGapWindow(t *testing.T) { +func TestGapWindows(t *testing.T) { open := func(t *testing.T, dir string) *settings.Registry { t.Helper() guardGlobals(t) @@ -697,22 +700,21 @@ func TestLongestGapWindow(t *testing.T) { } t.Run("flat directory", func(t *testing.T) { - assert.Equal(t, 45*time.Minute, longestGapWindow(open(t, writeSettings(t, gapWindow(45))))) + assert.Equal(t, map[tenant.ID]time.Duration{tenant.Default: 45 * time.Minute}, gapWindows(open(t, writeSettings(t, gapWindow(45))))) }) t.Run("nested directory", func(t *testing.T) { root := writeNestedSettings(t, map[string]map[string]any{"acme": gapWindow(15), "globex": gapWindow(60), "initech": gapWindow(30)}) tenants := open(t, root) - assert.Equal(t, 60*time.Minute, longestGapWindow(tenants)) + assert.Equal(t, map[tenant.ID]time.Duration{"acme": 15 * time.Minute, "globex": 60 * time.Minute, "initech": 30 * time.Minute}, gapWindows(tenants)) - // A rejected tenant is not being served, so its window is not weighed. rewriteSettings(t, filepath.Join(root, "globex"), invalidQuery) tenants.Reload("test") - assert.Equal(t, 30*time.Minute, longestGapWindow(tenants)) + assert.Equal(t, map[tenant.ID]time.Duration{"acme": 15 * time.Minute, "initech": 30 * time.Minute}, gapWindows(tenants)) }) - t.Run("no tenant served keeps nothing", func(t *testing.T) { - assert.Zero(t, longestGapWindow(open(t, writeNestedSettings(t, map[string]map[string]any{"acme": invalidQuery})))) + t.Run("no tenant served names none", func(t *testing.T) { + assert.Empty(t, gapWindows(open(t, writeNestedSettings(t, map[string]map[string]any{"acme": invalidQuery})))) }) } diff --git a/internal/app/wire.go b/internal/app/wire.go index 009f9920e..60cbdcd82 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -65,18 +65,13 @@ func (a *App) wireSettings() error { return fmt.Errorf("settings directory %s invalid, refusing to start — findings above; `wavehouse validate` reproduces them, `wavehouse bootstrap` writes a starter directory", a.cfg.Settings.Dir) } a.tenants = tenants - // Registered first: hooks run in registration order, and every other one - // reads tenant 0 through the store this one tracks. + // Registered first: hooks run in registration order, so every reload + // updates the tracked store before any other hook runs. a.trackDefaultStore() a.onDefaultAdopt(a.trackDefaultStore) a.policies = func() *policy.Policy { return defaultSetting(a, (*settings.Store).Policy) } - switch _, served := tenants.For(tenant.Default); { - case !tenants.Nested(): - if a.policies() == nil { - slog.Warn("no policy adopted — every token-based request is denied until policies.json defines one (fail closed)") - } - case !served: - slog.Warn("nested settings directory with no tenant 0 being served: the MQ byte budget is still configured from tenant 0's config.json, so it runs unconfigured until a 0 folder is adopted") + if !tenants.Nested() && a.policies() == nil { + slog.Warn("no policy adopted — every token-based request is denied until policies.json defines one (fail closed)") } return nil } @@ -84,20 +79,18 @@ func (a *App) wireSettings() error { // trackDefaultStore remembers tenant 0's store as of its last adoption. The // registry stops handing out a rejected tenant's store and forgets a removed // one, but the store keeps its last adopted document either way — and that is -// what the process-wide resources go on following (defaultSetting). +// what defaultSetting goes on reading. func (a *App) trackDefaultStore() { if store, ok := a.tenants.For(tenant.Default); ok { a.defaultStore.Store(store) } } -// defaultSetting reads one setting of the default tenant, which the one -// process-wide resource left (the MQ) follows until #583 gives each tenant -// its own. It reads tenant 0's last adopted document, so a -// 0 folder a reload rejected or removed leaves every reader as it was — -// the MQ's byte budget a hook reconciles and the one read per request (the ops -// gate's admin role) alike. A nested directory that has never served a tenant -// 0 reads T's zero value, which wireSettings warned about at boot. +// defaultSetting reads one setting of the default tenant: the admin role the +// ops gate of a flat directory reads per request. It reads tenant 0's last +// adopted document, so a 0 folder a reload rejected or removed leaves its +// reader as it was. A nested directory that has never served a tenant 0 reads +// T's zero value, and its ops gate reads no policy at all. func defaultSetting[T any](a *App, get func(*settings.Store) T) T { store := a.defaultStore.Load() if store == nil { @@ -108,8 +101,8 @@ func defaultSetting[T any](a *App, get func(*settings.Store) T) T { } // onDefaultAdopt registers fn to run after each reload that adopts the -// default tenant, so a nested directory's other tenants never move the -// process-wide resources, and a rejected 0 folder leaves them as they were. +// default tenant, so a nested directory's other tenants never move what +// follows it, and a rejected 0 folder leaves that as it was. func (a *App) onDefaultAdopt(fn func()) { a.tenants.AfterAdopt(func(adopted []tenant.ID) { if slices.Contains(adopted, tenant.Default) { @@ -137,20 +130,16 @@ func shortestKeepalive(tenants *settings.Registry) (period time.Duration, bucket return period, buckets } -// longestGapWindow is the shape of the one purge bound every tenant's events -// share: the ingest queue is one stream and the sweeper purges below one -// sequence, so the history kept is the longest stream.gap_window_minutes -// among the tenants being served — purging less, never more, so every -// tenant's gap-fill history survives — at the cost of one tenant holding the -// others' history for longer, which a stream per tenant will end (#583 story -// 5b). A flat directory's one tenant gets exactly its own window; -// with no tenant served the zero window purges everything acknowledged. -func longestGapWindow(tenants *settings.Registry) time.Duration { - var window time.Duration - for _, store := range tenants.All() { - window = max(window, store.GapWindow()) +// gapWindows is the history the sweeper keeps for each tenant being served: +// its own stream.gap_window_minutes, since each tenant's events have a queue +// of their own. A tenant it does not name — removed or rejected — keeps no +// history (mq.Purger.PurgeAcked). +func gapWindows(tenants *settings.Registry) map[tenant.ID]time.Duration { + windows := map[tenant.ID]time.Duration{} + for id, store := range tenants.All() { + windows[id] = store.GapWindow() } - return window + return windows } // served reports whether the registry is serving tenant id: what the @@ -185,10 +174,11 @@ func perTenant[T any](tenants *settings.Registry, get func(*settings.Store) T) f // miss reads as DLQ on, not as the zero value perTenant would give: off lets // the worker drop a message it cannot read, and not knowing the tenant is no // reason to destroy its row. Parked, it survives until the tenant resolves. -// So a removed or rejected tenant's queued rows are parked under its own -// subject rather than left unacked for its return: an unacked row holds the -// ack floor, the sweeper stops purging, and the one shared stream fills -// toward mq.max_bytes_gb until every tenant's ingest answers 503. +// So a removed or rejected tenant's queued rows are parked in its own +// dead-letter queue rather than left unacked for its return: unacked, each +// would be redelivered every ack wait for as long as the tenant is away, and +// would hold the tenant's ack floor, so the sweeper could purge none of its +// queue past it. func dlqFor(tenants *settings.Registry) func(tenant.ID, string) bool { return func(id tenant.ID, table string) bool { store, ok := tenants.For(id) @@ -541,15 +531,23 @@ func (a *App) wireDedupe() error { } // wireMQ starts the MQ — the embedded NATS under data_dir/nats, the one -// place the implementation is chosen; everything after it sees mq.Broker. -// mq.max_bytes_gb is hot-reloadable: after each adoption the new budget is -// handed to the MQ, which owns how it is split across its queues and keeps -// them consistent (see mq.Broker.SetMaxBytes). +// place the implementation is chosen; everything after it sees mq.Broker — +// and hands it each served tenant's mq.max_bytes_gb, which opens that +// tenant's queue the first time. The budget is hot-reloadable: after every +// reload the registry applies, each served tenant's is handed over again, +// and the MQ owns how it is split across the tenant's queues and keeps them +// consistent (see mq.Broker.SetMaxBytes). A tenant no longer served keeps +// its queue at the budget it last had. A queue that cannot be opened or +// resized follows the registry's rule for the shape: a flat directory +// refuses boot, like every other store, and on a reload logs it, keeping the +// previous budget; a nested directory logs it at boot too, so it never costs +// the process — the tenant's ingest answers 503 until a reload opens its +// queue. The hook is registered before the boot apply, as the dedupe one is. func (a *App) wireMQ() error { dir := filepath.Join(a.cfg.DataDir, "nats") config.WarnIfFreshDataDir("nats", dir) var broker mq.Broker - broker, err := mq.NewEmbedded(dir, defaultSetting(a, (*settings.Store).MQMaxBytes)) + broker, err := mq.NewEmbedded(dir) if err != nil { config.LogStorageInitError("mq", dir, err) return fmt.Errorf("mq open: %w", err) @@ -570,17 +568,26 @@ func (a *App) wireMQ() error { // Rooted in the App's stop context, so a reload caught mid-hook by // SIGTERM gives up rather than holding the drain past // server.shutdown_timeout. - a.onDefaultAdopt(func() { - mb := defaultSetting(a, (*settings.Store).MQMaxBytes) - if mb == broker.MaxBytes() { - return - } - if err := broker.SetMaxBytes(a.stopCtx, mb); err != nil { - slog.Error("mq stream resize failed; the next reload retries", "error", err) - return + reconcile := func() error { + var errs []error + for id, store := range a.tenants.All() { + mb := store.MQMaxBytes() + if mb == broker.MaxBytes(id) { + continue + } + if err := broker.SetMaxBytes(a.stopCtx, id, mb); err != nil { + slog.Error("mq queue not reconciled with settings; the next reload retries", "tenant", id, "error", err) + errs = append(errs, fmt.Errorf("tenant %s: %w", id, err)) + continue + } + slog.Info("mq queue reconciled with settings", "tenant", id, "max_bytes_gb", mb>>30) } - slog.Info("mq stream limits reconciled with settings", "max_bytes_gb", mb>>30) - }) + return errors.Join(errs...) + } + a.tenants.AfterAdopt(func([]tenant.ID) { _ = reconcile() }) + if err := reconcile(); err != nil && !a.tenants.Nested() { + return fmt.Errorf("mq open: %w", err) + } return nil } @@ -597,11 +604,11 @@ func (a *App) wireCache() error { } // wireSweeper adds the active sweeper — purges messages that are both -// written to ClickHouse and older than the SSE gap window (the longest -// stream.gap_window_minutes among the tenants served, re-read every sweep — -// see longestGapWindow). Runs every minute. +// written to ClickHouse and older than their tenant's SSE gap window (its own +// stream.gap_window_minutes, re-read every sweep — see gapWindows). Runs +// every minute. func (a *App) wireSweeper() { - sweeper := ingest.NewSweeper(a.mq, func() time.Duration { return longestGapWindow(a.tenants) }) + sweeper := ingest.NewSweeper(a.mq, func() map[tenant.ID]time.Duration { return gapWindows(a.tenants) }) a.add(component{name: "sweeper", run: func(ctx context.Context) error { sweeper.Start(ctx) return nil diff --git a/internal/ingest/sweeper.go b/internal/ingest/sweeper.go index 85383ead5..367e22af9 100644 --- a/internal/ingest/sweeper.go +++ b/internal/ingest/sweeper.go @@ -7,35 +7,34 @@ import ( "time" "github.com/Wave-RF/WaveHouse/internal/mq" + "github.com/Wave-RF/WaveHouse/internal/tenant" ) // Sweeper implements the Active Sweeper pattern. It runs every minute and // asks the MQ to purge the ingest events that satisfy BOTH conditions: // - ACKed by the buffer consumer (written to ClickHouse) -// - Older than the gap window (no longer needed for SSE replay) +// - Older than their tenant's gap window (no longer needed for SSE replay) // -// This guarantees: healthy state keeps exactly gap_window of rolling data; -// ClickHouse down freezes purging; a catastrophic outage fills the queue to -// its byte budget and triggers backpressure (mq.ErrQueueFull). How the MQ -// finds the purge point is its own business (see mq.Purger). +// This guarantees: healthy state keeps exactly each tenant's gap_window of +// rolling data; ClickHouse down freezes purging; a catastrophic outage fills +// a tenant's queue to its byte budget and triggers backpressure +// (mq.ErrQueueFull). How the MQ finds the purge point is its own business +// (see mq.Purger). type Sweeper struct { purger mq.Purger - // gapWindow is the history to keep, read on every sweep so a reload of - // stream.gap_window_minutes applies from the next sweep without a - // restart. The ingest queue is one stream for every tenant and a purge - // is one bound over it, so in production this is the longest window - // among the tenants being served (internal/app's longestGapWindow); a - // tenant's own window follows once the streams are per tenant (#583 - // story 5b). - gapWindow func() time.Duration + // gapWindows is the history to keep for each tenant being served, read on + // every sweep so a reload of stream.gap_window_minutes applies from the + // next sweep without a restart. A tenant it does not name — one removed + // or rejected — keeps no history (mq.Purger.PurgeAcked). + gapWindows func() map[tenant.ID]time.Duration } -// NewSweeper creates the Active Sweeper. gapWindow is resolved per sweep. +// NewSweeper creates the Active Sweeper. gapWindows is resolved per sweep. // TODO: (future) need leader election or shared lock to only run one instance of the sweeper in clustered mode -func NewSweeper(purger mq.Purger, gapWindow func() time.Duration) *Sweeper { +func NewSweeper(purger mq.Purger, gapWindows func() map[tenant.ID]time.Duration) *Sweeper { return &Sweeper{ - purger: purger, - gapWindow: gapWindow, + purger: purger, + gapWindows: gapWindows, } } @@ -54,7 +53,13 @@ func (s *Sweeper) Start(ctx context.Context) { } func (s *Sweeper) sweep(ctx context.Context) { - _, err := s.purger.PurgeAcked(ctx, BufferConsumerName, time.Now().Add(-s.gapWindow())) + now := time.Now() + windows := s.gapWindows() + cutoffs := make(map[tenant.ID]time.Time, len(windows)) + for id, window := range windows { + cutoffs[id] = now.Add(-window) + } + _, err := s.purger.PurgeAcked(ctx, BufferConsumerName, cutoffs) if err != nil { if errors.Is(err, mq.ErrConsumerNotFound) { // Consumer may not exist yet if no messages have been ingested. diff --git a/internal/ingest/sweeper_test.go b/internal/ingest/sweeper_test.go index dfd4e02f3..9b1eabab2 100644 --- a/internal/ingest/sweeper_test.go +++ b/internal/ingest/sweeper_test.go @@ -7,6 +7,7 @@ import ( "time" "github.com/Wave-RF/WaveHouse/internal/mq" + "github.com/Wave-RF/WaveHouse/internal/tenant" "github.com/Wave-RF/WaveHouse/internal/testutil" "github.com/stretchr/testify/assert" "github.com/stretchr/testify/require" @@ -15,11 +16,11 @@ import ( // The purge-point arithmetic is the MQ's (internal/mq/purge_test.go); the // sweeper owns only when to ask and what window to ask for. -func TestSweep_AsksForTheBufferConsumerAndTheGapWindow(t *testing.T) { +func TestSweep_AsksForTheBufferConsumerAndEachTenantsGapWindow(t *testing.T) { t.Parallel() - gapWindow := 5 * time.Minute + windows := map[tenant.ID]time.Duration{"acme": 5 * time.Minute, "globex": time.Hour} purger := &testutil.MockPurger{Purged: true} - s := NewSweeper(purger, func() time.Duration { return gapWindow }) + s := NewSweeper(purger, func() map[tenant.ID]time.Duration { return windows }) before := time.Now() s.sweep(context.Background()) @@ -28,29 +29,35 @@ func TestSweep_AsksForTheBufferConsumerAndTheGapWindow(t *testing.T) { require.Len(t, purger.Calls, 1) call := purger.Calls[0] assert.Equal(t, BufferConsumerName, call.Consumer) - assert.False(t, call.OlderThan.Before(before.Add(-gapWindow)), "cutoff is now - gap window") - assert.False(t, call.OlderThan.After(after.Add(-gapWindow)), "cutoff is now - gap window") + require.Len(t, call.OlderThan, 2, "one cutoff per tenant served") + for id, window := range windows { + cutoff := call.OlderThan[id] + assert.False(t, cutoff.Before(before.Add(-window)), "%s: cutoff is now - its own gap window", id) + assert.False(t, cutoff.After(after.Add(-window)), "%s: cutoff is now - its own gap window", id) + } } -func TestSweep_RereadsTheGapWindowEverySweep(t *testing.T) { +func TestSweep_RereadsTheGapWindowsEverySweep(t *testing.T) { t.Parallel() - gapWindow := time.Minute + windows := map[tenant.ID]time.Duration{"acme": time.Minute} purger := &testutil.MockPurger{} - s := NewSweeper(purger, func() time.Duration { return gapWindow }) + s := NewSweeper(purger, func() map[tenant.ID]time.Duration { return windows }) s.sweep(context.Background()) - gapWindow = time.Hour // a settings reload + windows = map[tenant.ID]time.Duration{"acme": time.Hour, "globex": time.Minute} // a settings reload s.sweep(context.Background()) require.Len(t, purger.Calls, 2) - assert.Greater(t, purger.Calls[0].OlderThan.Sub(purger.Calls[1].OlderThan), 50*time.Minute) + assert.Greater(t, purger.Calls[0].OlderThan["acme"].Sub(purger.Calls[1].OlderThan["acme"]), 50*time.Minute) + assert.NotContains(t, purger.Calls[0].OlderThan, tenant.ID("globex")) + assert.Contains(t, purger.Calls[1].OlderThan, tenant.ID("globex"), "a tenant adopted since is named from the next sweep") } func TestSweep_ErrorsDoNotPanic(t *testing.T) { t.Parallel() for _, err := range []error{mq.ErrConsumerNotFound, errors.New("broker unavailable")} { purger := &testutil.MockPurger{Err: err} - s := NewSweeper(purger, func() time.Duration { return time.Minute }) + s := NewSweeper(purger, func() map[tenant.ID]time.Duration { return map[tenant.ID]time.Duration{"acme": time.Minute} }) s.sweep(context.Background()) assert.Len(t, purger.Calls, 1) } @@ -62,7 +69,7 @@ func TestSweep_ErrorsDoNotPanic(t *testing.T) { func TestStart_ContextCancellation(t *testing.T) { t.Parallel() - s := NewSweeper(&testutil.MockPurger{}, func() time.Duration { return 5 * time.Minute }) + s := NewSweeper(&testutil.MockPurger{}, func() map[tenant.ID]time.Duration { return nil }) ctx, cancel := context.WithCancel(context.Background()) cancel() // Cancel immediately. diff --git a/internal/ingest/worker.go b/internal/ingest/worker.go index 618b2b357..9af5b1ef3 100644 --- a/internal/ingest/worker.go +++ b/internal/ingest/worker.go @@ -116,10 +116,12 @@ const ( // maxAckPending, and ackWait > defaultMaxWait + CH flush (else in-flight // messages are redelivered mid-processing → duplicate inserts). const ( - // Server-side cap on unacked messages; suspends delivery when hit (backpressure). + // Server-side cap on a tenant's unacked messages; suspends that tenant's + // delivery when hit (backpressure), and no other tenant's. maxAckPending = 10_000 // TODO: raise if NATS delivery becomes the bottleneck - // Client prefetch buffer in front of msgChan (was the implicit jetstream default). + // Client prefetch buffer in front of msgChan (was the implicit jetstream + // default), shared by the tenants' queues (mq.Consumer.Consume). pullMaxMessages = 500 // Redelivery timeout. 60s ≈ 5s batch + ~30s HTTP timeout + margin. @@ -219,8 +221,9 @@ func waitOrDeadline(ctx context.Context, wg *sync.WaitGroup) error { } } -// dispatchLoop owns the single JetStream consumer and fans every message out to -// a tableLoop per tenant table (lazily spawned on first sight of one). It does +// dispatchLoop owns the one consumer — held on every tenant's queue — and fans +// every message out to a tableLoop per tenant table (lazily spawned on first +// sight of one). It does // no batching itself — it parses just enough to route — so a low-volume table // can never strand another table's rows behind a shared timer. It is the ONLY // goroutine that watches ctx; tableLoops stop via channel-close, which gives a @@ -230,11 +233,13 @@ func (w *IngestWorker) dispatchLoop(ctx context.Context, cons mq.Consumer) { msgChan := make(chan *mq.Message, w.maxBatch*2) - // Pull consumer with a push-like callback (the client prefetches pullMaxMessages). - // Hand off to msgChan only, so the consume goroutine never blocks on flush work. + // Pull consumer with a push-like callback (the client prefetches pullMaxMessages, + // shared by the tenants' queues). It runs on one delivery goroutine per tenant, + // so the handoff is a channel send, safe from all of them at once. Hand off to + // msgChan only, so a consume goroutine never blocks on flush work. // The handoff also watches ctx: stop (deferred below) does not wait for a // delivery already in the handler, so once this loop has stopped draining - // msgChan a full channel would otherwise pin the client's delivery goroutine + // msgChan a full channel would otherwise pin a delivery goroutine // forever. A message dropped here is unacked and simply redelivered. stop, deliveryEnded, err := cons.Consume(func(msg *mq.Message) { select { diff --git a/internal/ingest/worker_test.go b/internal/ingest/worker_test.go index c3a0988ed..7fe9130e2 100644 --- a/internal/ingest/worker_test.go +++ b/internal/ingest/worker_test.go @@ -122,9 +122,7 @@ func TestStartIngestWorker_Validation(t *testing.T) { { name: "nil cache", setup: func(t *testing.T) (Queue, cache.Cache) { - emb, err := mq.NewEmbedded(t.TempDir(), 1024*1024) - require.NoError(t, err) - t.Cleanup(func() { _ = emb.Close() }) + emb := testutil.NewEmbeddedMQ(t, 1024*1024) return emb, nil }, wantErrSub: "cache is nil", @@ -154,9 +152,7 @@ func TestStartIngestWorker_EndToEnd(t *testing.T) { t.Parallel() // ── Embedded MQ ── - emb, err := mq.NewEmbedded(t.TempDir(), 4*1024*1024) - require.NoError(t, err) - t.Cleanup(func() { _ = emb.Close() }) + emb := testutil.NewEmbeddedMQ(t, 4*1024*1024) // ── ClickHouse stub: capture each request body, return 200 ── var ( @@ -247,9 +243,7 @@ func TestStartIngestWorker_EndToEnd(t *testing.T) { func TestStartIngestWorker_StopFunc_RespectsShutdownDeadline(t *testing.T) { t.Parallel() - emb, err := mq.NewEmbedded(t.TempDir(), 1024*1024) - require.NoError(t, err) - t.Cleanup(func() { _ = emb.Close() }) + emb := testutil.NewEmbeddedMQ(t, 1024*1024) // ClickHouse stub that blocks until we say go — keeps the worker's // flush goroutine alive past the stop call. @@ -298,9 +292,7 @@ func TestStartIngestWorker_StopFunc_RespectsShutdownDeadline(t *testing.T) { func TestStartIngestWorker_StopFunc_CleanShutdown(t *testing.T) { t.Parallel() - emb, err := mq.NewEmbedded(t.TempDir(), 1024*1024) - require.NoError(t, err) - t.Cleanup(func() { _ = emb.Close() }) + emb := testutil.NewEmbeddedMQ(t, 1024*1024) // chURL is never dialed: with no messages there is no flush, so a dummy // host/port is fine. @@ -1094,9 +1086,7 @@ func TestDispatchLoop_PerTableBatching_NoCrossTableContamination(t *testing.T) { batchB = maxBatch ) - emb, err := mq.NewEmbedded(t.TempDir(), 8*1024*1024) - require.NoError(t, err) - t.Cleanup(func() { _ = emb.Close() }) + emb := testutil.NewEmbeddedMQ(t, 8*1024*1024) // CH stub: count rows (newlines in the JSONCompactEachRow body) per target table. var ( @@ -1187,9 +1177,7 @@ func TestDispatchLoop_PartialBatchWaitsForOwnTrigger(t *testing.T) { total = 4 // 3 → one full batch on the size trigger; 1 leftover ) - emb, err := mq.NewEmbedded(t.TempDir(), 8*1024*1024) - require.NoError(t, err) - t.Cleanup(func() { _ = emb.Close() }) + emb := testutil.NewEmbeddedMQ(t, 8*1024*1024) // CH stub counts rows and sleeps briefly, so the 4th row is reliably buffered // before the first (3-row) flush completes — that's when the old code would @@ -1791,9 +1779,7 @@ func TestDispatchLoop_BatchesPerTenantTable(t *testing.T) { t.Parallel() const maxBatch = 2 - emb, err := mq.NewEmbedded(t.TempDir(), 8*1024*1024) - require.NoError(t, err) - t.Cleanup(func() { _ = emb.Close() }) + emb := testutil.NewEmbeddedMQ(t, 8*1024*1024, "acme", "globex") // CH stub: record each INSERT's body under the database it named. var ( diff --git a/internal/mq/embedded.go b/internal/mq/embedded.go index 4219c34d5..dbe35fa5f 100644 --- a/internal/mq/embedded.go +++ b/internal/mq/embedded.go @@ -5,12 +5,16 @@ import ( "errors" "fmt" "log/slog" + "maps" + "math" + "slices" "strings" "sync" "sync/atomic" "time" "github.com/Wave-RF/WaveHouse/internal/observability" + "github.com/Wave-RF/WaveHouse/internal/tenant" natsserver "github.com/nats-io/nats-server/v2/server" "github.com/nats-io/nats.go" "github.com/nats-io/nats.go/jetstream" @@ -44,45 +48,73 @@ func (slogNATSLogger) Tracef(format string, v ...any) { slog.Debug(fmt.Sprintf(format, v...), "component", "nats") } -// EmbeddedNATS runs an in-process NATS server with JetStream. +// EmbeddedNATS runs an in-process NATS server with JetStream, and gives each +// tenant a queue of its own: an ingest stream and a dead-letter stream +// (subject.go names them), each with its own byte cap, and the durable +// consumers on the ingest one. Nothing outside this package sees that layout. type EmbeddedNATS struct { server *natsserver.Server conn *nats.Conn js jetstream.JetStream - limitMu sync.Mutex - maxBytes int64 // the ingest stream cap both streams were last reconciled to + // mu guards queues and consumers, and serializes opening or resizing a + // tenant's queue with registering a consumer, so a queue opened while a + // consumer registers is never missed by it. It is held across the + // JetStream calls that open or resize a queue. + mu sync.Mutex + queues map[tenant.ID]*tenantQueue + // consumers are the durable consumers held on every tenant's queue, each + // joined to a queue as it opens. + consumers []*fanIn +} + +// tenantQueue is what the broker knows of one tenant's queue. +type tenantQueue struct { + // ingest and dlq report whether each of the tenant's streams exists. + ingest, dlq bool + // maxBytes is the budget last applied in full (MaxBytes); asked is the + // budget last asked for, which a publish or park that finds a stream + // missing opens it at. Both are read back from the ingest stream at boot, + // so a tenant no longer served keeps the budget it last had. + maxBytes, asked int64 } // EmbeddedNATS is the one implementation of every mq interface. var _ Broker = (*EmbeddedNATS)(nil) const ( - // dlqShare is the DLQ stream's slice of the byte budget: a tenth of the - // ingest stream's cap. + // dlqShare is a tenant's dead-letter stream's slice of its byte budget: a + // tenth of the ingest stream's cap. dlqShare = 10 - // resizeTimeout bounds the JetStream calls SetMaxBytes makes to apply a - // new cap — both streams share it. A settings reload holds the store's - // lock while its hooks run, so an in-process JetStream call that never - // returns would otherwise block every later reload. + // resizeTimeout bounds the JetStream calls SetMaxBytes makes to open a + // tenant's queue or apply a new cap to it — both streams share it. A + // settings reload holds the store's lock while its hooks run, so an + // in-process JetStream call that never returns would otherwise block every + // later reload. resizeTimeout = 10 * time.Second - // rollbackTimeout is the undo's own budget when the DLQ resize fails: - // in-process JetStream fails by stalling rather than erroring, so the - // likely cause is that resizeTimeout has just run out, and an undo on + // rollbackTimeout is the undo's own budget when the dead-letter resize + // fails: in-process JetStream fails by stalling rather than erroring, so + // the likely cause is that resizeTimeout has just run out, and an undo on // that context would fail without touching the stream. SetMaxBytes runs // for at most the sum of the two. rollbackTimeout = 5 * time.Second ) -// NewEmbedded starts an embedded NATS server with JetStream enabled and -// both streams in place: the ingest stream capped at maxBytes and the DLQ -// stream at a tenth of it. The DLQ stream is always present — an empty -// limits-policy stream costs nothing, and whether a poison row lands on it is -// the ingest worker's decision at the moment of the failure. -// The server logs through slog's default logger. The stream names are fixed -// (see subject.go) — the embedded server is private to this process, so -// there's no namespacing to do. -func NewEmbedded(storeDir string, maxBytes int64) (*EmbeddedNATS, error) { +// errNoQueue is why a publish or park finds no queue it can open: no budget +// has been asked for the tenant yet (see SetMaxBytes). Publish reports it as +// ErrQueueFull. +var errNoQueue = errors.New("no queue is open for it yet") + +// NewEmbedded starts an embedded NATS server with JetStream over storeDir and +// takes stock of the tenants' queues already there: a consumer created later +// is held on every one of them, those of tenants no longer served included, +// whose queued rows still have to reach the ingest worker. The pair of streams +// an earlier build kept for every tenant together is deleted, since its +// subjects overlap every tenant's; the events it held are not carried over. A +// tenant's queue is opened by SetMaxBytes, the first time its budget is +// applied, or by a publish or park that finds it missing, at the budget last +// asked for it. The server logs through slog's default logger. +func NewEmbedded(storeDir string) (*EmbeddedNATS, error) { opts := &natsserver.Options{ DontListen: true, JetStream: true, @@ -93,6 +125,18 @@ func NewEmbedded(storeDir string, maxBytes int64) (*EmbeddedNATS, error) { // channel" panic) and os.Exit(0)s past its cleanup. WaveHouse owns // the lifecycle; Close() shuts the server down. See #287. NoSigs: true, + // JetStream counts every stream's byte cap as reserved disk and + // refuses a stream once the caps together pass this limit — by + // default 75% of the free disk at boot. A tenant's mq.max_bytes_gb + // caps that tenant's queue and nothing else; what the tenants' caps + // add up to against the disk is #138's to decide, not a limit the + // server enforces on the side, so its own is set out of reach. Half + // the int64 range, not all of it: the server subtracts its count of + // reserved bytes from this limit, and a stream whose store fails to + // open releases a reservation it never made (nats-server 2.14.6), so + // the count can fall below zero — at the top of the range that + // subtraction overflows, and every stream after it is refused. + JetStreamMaxStore: math.MaxInt64 / 2, } ns, err := natsserver.NewServer(opts) @@ -119,114 +163,285 @@ func NewEmbedded(storeDir string, maxBytes int64) (*EmbeddedNATS, error) { return nil, fmt.Errorf("jetstream new: %w", err) } - if _, err := js.CreateOrUpdateStream(context.Background(), ingestStreamConfig(maxBytes)); err != nil { - nc.Close() - ns.Shutdown() - return nil, fmt.Errorf("create stream: %w", err) + e := &EmbeddedNATS{server: ns, conn: nc, js: js, queues: map[tenant.ID]*tenantQueue{}} + if err := e.takeStock(context.Background()); err != nil { + _ = e.Close() + return nil, err } - if _, err := js.CreateOrUpdateStream(context.Background(), dlqStreamConfig(maxBytes/dlqShare)); err != nil { - nc.Close() - ns.Shutdown() - return nil, fmt.Errorf("create dlq stream: %w", err) + return e, nil +} + +// takeStock deletes the pair of streams an earlier build kept for every +// tenant together, then records every tenant stream on disk, with the budget +// its ingest stream last had. +func (e *EmbeddedNATS) takeStock(ctx context.Context) error { + for _, name := range []string{legacyIngestStream, legacyDLQStream} { + if err := e.deleteLegacy(ctx, name); err != nil { + return err + } + } + streams := e.js.ListStreams(ctx) + for info := range streams.Info() { + name := info.Config.Name + if id, ok := streamTenant(ingestStreamPrefix, name); ok { + q := e.queue(id) + q.ingest = true + q.maxBytes, q.asked = info.Config.MaxBytes, info.Config.MaxBytes + } else if id, ok := streamTenant(dlqStreamPrefix, name); ok { + e.queue(id).dlq = true + } } + if err := streams.Err(); err != nil { + return fmt.Errorf("list streams: %w", err) + } + return nil +} - return &EmbeddedNATS{server: ns, conn: nc, js: js, maxBytes: maxBytes}, nil +// deleteLegacy deletes one stream of the pair an earlier build kept for every +// tenant together, logging what it held; one that is not there is nothing to +// do. +func (e *EmbeddedNATS) deleteLegacy(ctx context.Context, name string) error { + s, err := e.js.Stream(ctx, name) + if errors.Is(err, jetstream.ErrStreamNotFound) { + return nil + } + if err != nil { + return fmt.Errorf("look up stream %s: %w", name, err) + } + held := s.CachedInfo().State.Msgs + if err := e.js.DeleteStream(ctx, name); err != nil { + return fmt.Errorf("delete stream %s: %w", name, err) + } + slog.Warn("mq: deleted the stream an earlier build kept for every tenant together; its messages are not carried over", + "component", "nats", "stream", name, "messages", held) + return nil +} + +// queue returns what the broker knows of tenant id's queue, recording the +// tenant first if it knows nothing. Under e.mu (or before e is shared). +func (e *EmbeddedNATS) queue(id tenant.ID) *tenantQueue { + q := e.queues[id] + if q == nil { + q = &tenantQueue{} + e.queues[id] = q + } + return q } -// ingestStreamConfig is the WAVEHOUSE stream. LimitsPolicy: standard +// ingestTenants lists the tenants whose ingest stream exists, in id order. +// Under e.mu. +func (e *EmbeddedNATS) ingestTenants() []tenant.ID { + var ids []tenant.ID + for id, q := range e.queues { + if q.ingest { + ids = append(ids, id) + } + } + slices.Sort(ids) + return ids +} + +// ingestStreamConfig is tenant id's ingest stream. LimitsPolicy: standard // append-only log; the Active Sweeper handles message purging. MaxBytes caps -// disk usage to protect the shared ClickHouse/NATS disk. DiscardNew rejects -// new messages when full, propagating backpressure to the upstream API. -func ingestStreamConfig(maxBytes int64) jetstream.StreamConfig { +// the tenant's share of the disk. DiscardNew rejects new messages when full, +// propagating backpressure to the upstream API — for this tenant alone. +func ingestStreamConfig(id tenant.ID, maxBytes int64) jetstream.StreamConfig { return jetstream.StreamConfig{ - Name: ingestStream, - Subjects: []string{ingestAll}, + Name: ingestStreamName(id), + Subjects: []string{tenantSubjects(ingestPrefix, id)}, Retention: jetstream.LimitsPolicy, MaxBytes: maxBytes, Discard: jetstream.DiscardNew, } } -// dlqStreamConfig is the WAVEHOUSE_DLQ stream. DiscardOld: a full DLQ drops -// its oldest parked rows rather than refusing new ones — backpressure belongs -// to the ingest stream, not the dead-letter one. -func dlqStreamConfig(maxBytes int64) jetstream.StreamConfig { +// dlqStreamConfig is tenant id's dead-letter stream. DiscardOld: a full one +// drops its oldest parked rows rather than refusing new ones — backpressure +// belongs to the ingest stream, not the dead-letter one. +func dlqStreamConfig(id tenant.ID, maxBytes int64) jetstream.StreamConfig { return jetstream.StreamConfig{ - Name: dlqStream, - Subjects: []string{dlqAll}, + Name: dlqStreamName(id), + Subjects: []string{tenantSubjects(dlqPrefix, id)}, Retention: jetstream.LimitsPolicy, MaxBytes: maxBytes, Discard: jetstream.DiscardOld, } } -// MaxBytes reports the ingest stream cap both streams were last reconciled to -// (by NewEmbedded, then by each successful SetMaxBytes). -func (e *EmbeddedNATS) MaxBytes() int64 { - e.limitMu.Lock() - defer e.limitMu.Unlock() - return e.maxBytes +// MaxBytes reports the budget tenant id's queue was last given in full (by +// SetMaxBytes, or read back from disk at boot), 0 when it has none. +func (e *EmbeddedNATS) MaxBytes(id tenant.ID) int64 { + e.mu.Lock() + defer e.mu.Unlock() + if q := e.queues[id]; q != nil { + return q.maxBytes + } + return 0 } -// SetMaxBytes applies a new byte budget to both streams in place (the -// hot-reloadable mq.max_bytes_gb): the ingest stream takes maxBytes and the -// DLQ stream a tenth of it. JetStream applies a limit change to a live stream -// without touching its messages: growing takes effect immediately; shrinking -// the ingest stream below its current size makes DiscardNew refuse new -// publishes until the worker drains it — nothing buffered is dropped. +// SetMaxBytes applies tenant id's byte budget (its hot-reloadable +// mq.max_bytes_gb) to its queue: the ingest stream takes maxBytes and the +// dead-letter stream a tenth of it. A tenant with no queue yet has one opened, +// its dead-letter stream first, so no row is queued that could not be parked, +// and every registered consumer joins it. No other tenant's queue is touched. +// +// JetStream applies a limit change to a live stream without touching its +// messages: growing takes effect immediately; shrinking the ingest stream +// below its current size makes DiscardNew refuse new publishes until the +// worker drains it — nothing buffered is dropped. The dead-letter stream is +// DiscardOld, which would delete its oldest parked rows to fit a smaller cap, +// so it is never capped below the bytes it holds (#532): it keeps what it +// has, and that is logged. // -// The pair moves together where it can. If the DLQ update fails after the -// ingest one succeeded, the ingest resize is undone so the 10:1 pair stays at +// The pair moves together where it can. If the dead-letter update fails after +// the ingest one succeeded, the ingest resize is undone so the pair stays at // the previous budget, and the next call retries both. Safe in that direction // — the ingest stream is DiscardNew, so shrinking it back drops nothing // stored. The undo is best effort: if it fails too, the ingest stream stays at -// the new limit and the DLQ at the previous, and the error says so. On any -// error MaxBytes keeps reporting the previous budget, so a later call with the -// new budget reapplies both. +// the new limit and the dead-letter one at the previous, and the error says +// so. On any error MaxBytes keeps reporting the previous budget, so a later +// call with the new budget reapplies both. // // The JetStream calls are bounded by resizeTimeout, plus rollbackTimeout for // the undo, both rooted in ctx. That is deliberate: ctx is the process's stop // context, so a reload caught mid-hook by a stop gives up — undo included — // rather than holding the drain past server.shutdown_timeout. A cancellation // between the two updates is therefore the one way to leave the pair split, -// and only for the rest of a process that is exiting: the next boot -// reconciles both streams from the adopted settings. -func (e *EmbeddedNATS) SetMaxBytes(ctx context.Context, maxBytes int64) error { - e.limitMu.Lock() - defer e.limitMu.Unlock() - if maxBytes == e.maxBytes { +// and only for the rest of a process that is exiting: the next boot applies +// the adopted settings to it again. +func (e *EmbeddedNATS) SetMaxBytes(ctx context.Context, id tenant.ID, maxBytes int64) error { + if _, err := tenant.Parse(string(id)); err != nil { + return fmt.Errorf("tenant: %w", err) + } + e.mu.Lock() + defer e.mu.Unlock() + q := e.queue(id) + q.asked = maxBytes + if q.ingest && q.dlq && maxBytes == q.maxBytes { return nil } + return e.apply(ctx, id, q, maxBytes) +} +// apply brings tenant id's queue to maxBytes: opening it when its ingest +// stream is missing, resizing it otherwise (see SetMaxBytes). Under e.mu. +func (e *EmbeddedNATS) apply(ctx context.Context, id tenant.ID, q *tenantQueue, maxBytes int64) error { resizeCtx, cancel := context.WithTimeout(ctx, resizeTimeout) defer cancel() - if _, err := e.js.UpdateStream(resizeCtx, ingestStreamConfig(maxBytes)); err != nil { + if !q.ingest { + if err := e.applyDLQ(resizeCtx, id, q, maxBytes); err != nil { + return err + } + if _, err := e.js.CreateOrUpdateStream(resizeCtx, ingestStreamConfig(id, maxBytes)); err != nil { + return fmt.Errorf("open ingest stream: %w", err) + } + q.ingest, q.maxBytes = true, maxBytes + for _, f := range e.consumers { + if err := f.join(resizeCtx, id); err != nil { + f.fail(fmt.Errorf("tenant %s: %w: join its queue: %w", id, ErrDeliveryEnded, err)) + } + } + return nil + } + if _, err := e.js.UpdateStream(resizeCtx, ingestStreamConfig(id, maxBytes)); err != nil { return fmt.Errorf("resize ingest stream: %w", err) } - if _, err := e.js.CreateOrUpdateStream(resizeCtx, dlqStreamConfig(maxBytes/dlqShare)); err != nil { - // The undo runs on its own budget, not the one the DLQ call has - // likely just exhausted. + if err := e.applyDLQ(resizeCtx, id, q, maxBytes); err != nil { + // The undo runs on its own budget, not the one the dead-letter call + // has likely just exhausted. rollbackCtx, cancelRollback := context.WithTimeout(ctx, rollbackTimeout) defer cancelRollback() - if _, rollbackErr := e.js.UpdateStream(rollbackCtx, ingestStreamConfig(e.maxBytes)); rollbackErr != nil { - return fmt.Errorf("resize dlq stream: %w (ingest stream rollback failed, so it stays at the new limit and the dlq at the previous: %w)", err, rollbackErr) + if _, rollbackErr := e.js.UpdateStream(rollbackCtx, ingestStreamConfig(id, q.maxBytes)); rollbackErr != nil { + return fmt.Errorf("%w (ingest stream rollback failed, so it stays at the new limit and the dlq at the previous: %w)", err, rollbackErr) } - return fmt.Errorf("resize dlq stream: %w (ingest stream restored to the previous limit)", err) + return fmt.Errorf("%w (ingest stream restored to the previous limit)", err) } - e.maxBytes = maxBytes + q.maxBytes = maxBytes return nil } -// Publish stores data on topic's ingest subject. A topic without a valid -// tenant is refused before anything is sent (see subject). A stream at its -// byte budget (DiscardNew) refuses the publish; that is reported as -// ErrQueueFull. +// applyDLQ gives tenant id's dead-letter stream a tenth of maxBytes, creating +// it when it is missing, but never caps it below the bytes it holds: those +// stay, the cap is what they take, and the stream then drops its oldest row +// to make room for each new one, as any full dead-letter stream does. Under +// e.mu. +func (e *EmbeddedNATS) applyDLQ(ctx context.Context, id tenant.ID, q *tenantQueue, maxBytes int64) error { + limit := maxBytes / dlqShare + verb := "resize" + s, err := e.js.Stream(ctx, dlqStreamName(id)) + switch { + case errors.Is(err, jetstream.ErrStreamNotFound): + verb = "open" + case err != nil: + return fmt.Errorf("dlq stream info: %w", err) + default: + // A stream's size fits an int64 as its cap does; the bound is + // checked rather than assumed. + if held := s.CachedInfo().State.Bytes; held <= math.MaxInt64 && int64(held) > limit { + slog.Warn("mq: dead-letter queue kept at what it holds rather than shrunk to its budget, so no parked row is deleted", + "component", "nats", "tenant", id, "held_bytes", held, "budget_bytes", limit) + limit = int64(held) + } + } + if _, err := e.js.CreateOrUpdateStream(ctx, dlqStreamConfig(id, limit)); err != nil { + return fmt.Errorf("%s dlq stream: %w", verb, err) + } + q.dlq = true + return nil +} + +// reopen opens tenant id's queue at the budget last asked for it, for a +// publish or park that found one of its streams missing. errNoQueue when no +// budget has been asked for the tenant yet: a reload can make a tenant +// resolvable an instant before its budget arrives. +func (e *EmbeddedNATS) reopen(ctx context.Context, id tenant.ID) error { + e.mu.Lock() + defer e.mu.Unlock() + q := e.queues[id] + if q == nil || q.asked == 0 { + return fmt.Errorf("tenant %s: %w", id, errNoQueue) + } + // What is missing is asked of JetStream rather than read off the flags, + // which may still say the stream the publish just missed exists — or it + // may be back already, opened by a caller that held mu first. + for _, name := range []string{ingestStreamName(id), dlqStreamName(id)} { + _, err := e.js.Stream(ctx, name) + switch { + case errors.Is(err, jetstream.ErrStreamNotFound): + if name == ingestStreamName(id) { + q.ingest = false + } else { + q.dlq = false + } + case err != nil: + return fmt.Errorf("stream info: %w", err) + } + } + if q.ingest && q.dlq { + return nil + } + return e.apply(ctx, id, q, q.asked) +} + +// Publish stores data on topic's ingest subject, in its tenant's queue. A +// topic without a valid tenant is refused before anything is sent (see +// subject). A tenant with no queue has one opened at the budget last asked +// for it (see SetMaxBytes). A queue that cannot be opened — none asked for +// yet, or JetStream refused it — and a queue at its byte budget (DiscardNew) +// are reported as ErrQueueFull: either way the tenant's queue takes nothing +// now, and a retry is the caller's answer. func (e *EmbeddedNATS) Publish(ctx context.Context, topic Topic, data []byte, opts ...PublishOpt) error { subj, err := subject(ingestPrefix, topic) if err != nil { return err } err = e.publish(ctx, subj, data, opts) + if errors.Is(err, jetstream.ErrNoStreamResponse) { + if openErr := e.reopen(ctx, topic.Tenant); openErr != nil { + return fmt.Errorf("%w: %w", ErrQueueFull, openErr) + } + err = e.publish(ctx, subj, data, opts) + } if err != nil && strings.Contains(err.Error(), "maximum bytes exceeded") { // The server reports a full store as a generic store failure whose // text is the only thing that names the cause. @@ -235,12 +450,23 @@ func (e *EmbeddedNATS) Publish(ctx context.Context, topic Topic, data []byte, op return err } -// DeadLetter stores msg's data on its topic's DLQ subject — the subject it -// arrived on with the ingest prefix swapped for the DLQ one, nothing decoded -// or re-encoded. The DLQ stream is DiscardOld, so a full DLQ drops its oldest -// parked rows rather than refusing. +// DeadLetter stores msg's data on its topic's dead-letter subject, in its +// tenant's queue — the subject it arrived on with the ingest prefix swapped +// for the dead-letter one, nothing decoded or re-encoded. The dead-letter +// stream is DiscardOld, so a full one drops its oldest parked rows rather than +// refusing. A dead-letter stream found missing is opened again with its +// tenant's queue, as Publish does. func (e *EmbeddedNATS) DeadLetter(ctx context.Context, msg *Message, opts ...PublishOpt) error { - return e.publish(ctx, dlqPrefix+msg.topicKey, msg.Data, opts) + subj := dlqPrefix + msg.topicKey + err := e.publish(ctx, subj, msg.Data, opts) + if errors.Is(err, jetstream.ErrNoStreamResponse) { + if id, ok := keyTenant(msg.topicKey); ok { + if err = e.reopen(ctx, id); err == nil { + err = e.publish(ctx, subj, msg.Data, opts) + } + } + } + return err } func (e *EmbeddedNATS) publish(ctx context.Context, subj string, data []byte, opts []PublishOpt) error { @@ -255,7 +481,10 @@ func (e *EmbeddedNATS) publish(ctx context.Context, subj string, data []byte, op observability.InjectHeaders(ctx, headers) msg.Header = nats.Header(headers) - _, err := e.js.PublishMsg(ctx, msg) + // No retry on "no responders": in-process, that only ever means no + // stream holds the subject — a tenant with no queue, which the callers + // open rather than wait out. + _, err := e.js.PublishMsg(ctx, msg, jetstream.WithRetryAttempts(0)) return err } @@ -275,58 +504,170 @@ func wrapMsg(ctx context.Context, m jetstream.Msg) *Message { ) } +// Subscribe holds a durable explicit-ack consumer named consumerName on every +// tenant's queue, those opened later included, and delivers each message to +// handler with the trace context its headers carry, until ctx is done. A +// tenant's queue that cannot be joined when it opens is logged: its events +// reach handler from the next boot. func (e *EmbeddedNATS) Subscribe(ctx context.Context, consumerName string, handler func(msg *Message) error) error { - cons, err := e.js.CreateOrUpdateConsumer(ctx, ingestStream, jetstream.ConsumerConfig{ - Durable: consumerName, - FilterSubject: ingestAll, - AckPolicy: jetstream.AckExplicitPolicy, - }) - if err != nil { + f := e.newFanIn(ctx, jetstream.ConsumerConfig{Durable: consumerName, AckPolicy: jetstream.AckExplicitPolicy}) + f.fail = func(err error) { + slog.Error("mq: a tenant's events do not reach this consumer until the next boot", "component", "nats", "consumer", consumerName, "error", err) + } + if err := e.register(ctx, f); err != nil { return fmt.Errorf("create consumer: %w", err) } - - cctx, err := cons.Consume(func(m jetstream.Msg) { + stop, err := f.start(func(m jetstream.Msg) { msg := wrapMsg(observability.ExtractHeaders(ctx, m.Headers()), m) if err := handler(msg); err != nil { _ = msg.Nak() } - }) + }, 0, false) if err != nil { return fmt.Errorf("consume: %w", err) } go func() { <-ctx.Done() - cctx.Stop() + stop() }() return nil } // CreateConsumer creates or updates a durable explicit-ack pull consumer on -// the ingest stream. ctx becomes every delivered Message.Ctx (see -// ConsumerManager); it does not stop delivery — Consumer.Consume's stop does. +// every tenant's queue, and joins each queue opened later. ctx becomes every +// delivered Message.Ctx (see ConsumerManager); it does not stop delivery — +// Consumer.Consume's stop does. func (e *EmbeddedNATS) CreateConsumer(ctx context.Context, cfg ConsumerConfig) (Consumer, error) { - cons, err := e.js.CreateOrUpdateConsumer(ctx, ingestStream, jetstream.ConsumerConfig{ - Durable: cfg.Durable, - FilterSubject: ingestAll, - AckPolicy: jetstream.AckExplicitPolicy, - AckWait: cfg.AckWait, - MaxAckPending: cfg.MaxAckPending, - }) - if err != nil { + c := &workerConsumer{ + fanIn: e.newFanIn(ctx, jetstream.ConsumerConfig{ + Durable: cfg.Durable, + AckPolicy: jetstream.AckExplicitPolicy, + AckWait: cfg.AckWait, + MaxAckPending: cfg.MaxAckPending, + }), + failed: make(chan error, 1), + } + c.fail = func(err error) { + // Exactly one error, and nothing once stop has been called. + if c.stopped.Load() { + return + } + select { + case c.failed <- err: + default: + } + } + if err := e.register(ctx, c.fanIn); err != nil { return nil, fmt.Errorf("create consumer: %w", err) } - return &jsConsumer{cons: cons, ctx: ctx}, nil + return c, nil +} + +// newFanIn is a fanIn over cfg, not yet holding any durable; the caller sets +// its fail and registers it. +func (e *EmbeddedNATS) newFanIn(ctx context.Context, cfg jetstream.ConsumerConfig) *fanIn { + return &fanIn{ + e: e, + ctx: ctx, + cfg: cfg, + handles: map[tenant.ID]jetstream.Consumer{}, + running: map[tenant.ID]jetstream.ConsumeContext{}, + } } -// jsConsumer is the Consumer over a JetStream pull consumer. -type jsConsumer struct { - cons jetstream.Consumer - ctx context.Context // each delivered Message.Ctx (see ConsumerManager) +// register holds f's durable on every tenant's queue there is and registers +// f, so every queue opened from here on is joined too. +func (e *EmbeddedNATS) register(ctx context.Context, f *fanIn) error { + e.mu.Lock() + defer e.mu.Unlock() + for _, id := range e.ingestTenants() { + if err := f.join(ctx, id); err != nil { + return fmt.Errorf("tenant %s: %w", id, err) + } + } + e.consumers = append(e.consumers, f) + return nil +} + +// unregister stops joining f to the queues that open from here on. Under +// e.mu. +func (e *EmbeddedNATS) unregister(f *fanIn) { + e.consumers = slices.DeleteFunc(e.consumers, func(c *fanIn) bool { return c == f }) } -func (c *jsConsumer) Consume(handler func(msg *Message), prefetch int) (func(), <-chan error, error) { +// fanIn is one durable consumer held on every tenant's ingest stream — the +// ingest worker's (CreateConsumer) or the hub bridge's (Subscribe) — +// delivering them all into one handler: each tenant's messages on a +// goroutine of their own, so a tenant's arrive in order and different +// tenants' concurrently, and a handler blocked on one tenant holds back that +// tenant alone. Its fields are guarded by e.mu, bar stopped. +type fanIn struct { + e *EmbeddedNATS + ctx context.Context // each delivered Message.Ctx (CreateConsumer), or where Subscribe extracts trace context into + cfg jetstream.ConsumerConfig + + // fail reports a tenant's delivery that ended on its own, or a queue that + // could not be joined when it opened. + fail func(error) + + // handles is the durable on each tenant's ingest stream; running, the + // delivery started on each once deliver is set. + handles map[tenant.ID]jetstream.Consumer + running map[tenant.ID]jetstream.ConsumeContext + deliver func(jetstream.Msg) + // prefetch is the fetch-ahead asked for across the tenants together; 0 + // leaves each tenant the client default. + prefetch int + // watch reports a delivery that ends on its own through fail. + watch bool + stopped atomic.Bool +} + +// join holds f's durable on tenant id's ingest stream — looked up first, and +// created or updated only when missing or configured otherwise, so a boot +// over thousands of queues writes nothing it need not — and starts delivery +// on it when f is delivering. Under e.mu. +func (f *fanIn) join(ctx context.Context, id tenant.ID) error { + stream := ingestStreamName(id) + c, err := f.e.js.Consumer(ctx, stream, f.cfg.Durable) + if err != nil || !sameConsumer(c.CachedInfo().Config, f.cfg) { + if c, err = f.e.js.CreateOrUpdateConsumer(ctx, stream, f.cfg); err != nil { + return err + } + } + f.handles[id] = c + if f.deliver == nil || f.stopped.Load() { + return nil + } + return f.run(id) +} + +// sameConsumer reports whether a durable holds the fields this package sets; +// a zero field in want is the server's default, whatever that resolved to. +func sameConsumer(have, want jetstream.ConsumerConfig) bool { + return have.AckPolicy == want.AckPolicy && + have.FilterSubject == want.FilterSubject && + (want.AckWait == 0 || have.AckWait == want.AckWait) && + (want.MaxAckPending == 0 || have.MaxAckPending == want.MaxAckPending) +} + +// share is one tenant's part of the fetch-ahead: the total spread over the +// tenants' queues, at least one each. 0 leaves the client default. Under +// e.mu. +func (f *fanIn) share() int { + if f.prefetch <= 0 { + return 0 + } + return max(1, f.prefetch/max(1, len(f.handles))) +} + +// run starts delivery from tenant id's durable, once. Under e.mu. +func (f *fanIn) run(id tenant.ID) error { + if _, ok := f.running[id]; ok { + return nil + } // The client reports what goes wrong after Consume returns only through // this handler, never through Consume's own error. It calls it for // passing conditions too (a missed heartbeat, a leadership change) and @@ -339,37 +680,80 @@ func (c *jsConsumer) Consume(handler func(msg *Message), prefetch int) (func(), opts := []jetstream.PullConsumeOpt{ jetstream.ConsumeErrHandler(func(_ jetstream.ConsumeContext, err error) { lastErr.Store(&err) - slog.Warn("mq: consumer reported an error", "component", "nats", "error", err) + slog.Warn("mq: consumer reported an error", "component", "nats", "tenant", id, "error", err) }), } - if prefetch > 0 { - opts = append(opts, jetstream.PullMaxMessages(prefetch)) + if n := f.share(); n > 0 { + opts = append(opts, jetstream.PullMaxMessages(n)) } - cctx, err := c.cons.Consume(func(m jetstream.Msg) { - handler(wrapMsg(c.ctx, m)) - }, opts...) + cctx, err := f.handles[id].Consume(f.deliver, opts...) if err != nil { - return nil, nil, fmt.Errorf("consume: %w", err) + return err + } + f.running[id] = cctx + if !f.watch { + return nil } - - var stopped atomic.Bool - failed := make(chan error, 1) go func() { <-cctx.Closed() - if stopped.Load() { + if f.stopped.Load() { return } - if reason := lastErr.Load(); reason != nil { - failed <- fmt.Errorf("%w: %w", ErrDeliveryEnded, *reason) - return + reason := ErrDeliveryEnded + if r := lastErr.Load(); r != nil { + reason = fmt.Errorf("%w: %w", ErrDeliveryEnded, *r) } - failed <- ErrDeliveryEnded + f.fail(fmt.Errorf("tenant %s: %w", id, reason)) }() - stop := func() { - stopped.Store(true) - cctx.Stop() + return nil +} + +// start begins delivery to deliver from every tenant's durable, and from +// each queue joined later, fetching about prefetch messages ahead across the +// tenants together (see share); watch reports a delivery that ends on its own +// through fail. The returned stop ends every delivery and stops joining new +// queues, without waiting. +func (f *fanIn) start(deliver func(jetstream.Msg), prefetch int, watch bool) (stop func(), err error) { + f.e.mu.Lock() + defer f.e.mu.Unlock() + f.deliver, f.prefetch, f.watch = deliver, prefetch, watch + stop = func() { + f.e.mu.Lock() + defer f.e.mu.Unlock() + f.stopped.Store(true) + for _, cctx := range f.running { + cctx.Stop() + } + f.e.unregister(f) + } + for _, id := range slices.Sorted(maps.Keys(f.handles)) { + if err := f.run(id); err != nil { + f.stopped.Store(true) + for _, cctx := range f.running { + cctx.Stop() + } + f.e.unregister(f) + return nil, fmt.Errorf("tenant %s: %w", id, err) + } + } + return stop, nil +} + +// workerConsumer is the Consumer CreateConsumer returns: a fanIn with the +// failed channel its contract promises. +type workerConsumer struct { + *fanIn + failed chan error +} + +func (c *workerConsumer) Consume(handler func(msg *Message), prefetch int) (func(), <-chan error, error) { + stop, err := c.start(func(m jetstream.Msg) { + handler(wrapMsg(c.ctx, m)) + }, prefetch, true) + if err != nil { + return nil, nil, fmt.Errorf("consume: %w", err) } - return stop, failed, nil + return stop, c.failed, nil } // stream resolves a stream handle by name. @@ -433,40 +817,71 @@ func (s *jsStream) consumerAckFloor(ctx context.Context, consumer string) (uint6 return info.AckFloor.Stream, nil } -// PurgeAcked purges the ingest stream below MIN(consumer's ack floor + 1, -// first sequence stored at or after olderThan) — see purgeAcked. -func (e *EmbeddedNATS) PurgeAcked(ctx context.Context, consumer string, olderThan time.Time) (bool, error) { - s, err := e.stream(ctx, ingestStream) - if err != nil { - return false, fmt.Errorf("get stream: %w", err) +// PurgeAcked purges each tenant's ingest stream below MIN(consumer's ack +// floor + 1, first sequence stored at or after the tenant's cutoff) — see +// purgeAcked. A tenant olderThan does not name is purged up to its ack floor. +// A failure on one tenant's stream is joined into the error and the sweep +// goes on to the next; a done ctx ends it. +func (e *EmbeddedNATS) PurgeAcked(ctx context.Context, consumer string, olderThan map[tenant.ID]time.Time) (bool, error) { + e.mu.Lock() + ids := e.ingestTenants() + e.mu.Unlock() + + now := time.Now() + var ( + errs []error + tenants int + ) + for _, id := range ids { + if err := ctx.Err(); err != nil { + errs = append(errs, err) + break + } + cutoff, ok := olderThan[id] + if !ok { + cutoff = now + } + s, err := e.stream(ctx, ingestStreamName(id)) + if err != nil { + errs = append(errs, fmt.Errorf("tenant %s: get stream: %w", id, err)) + continue + } + report, err := purgeAcked(ctx, s, consumer, cutoff) + if err != nil { + errs = append(errs, fmt.Errorf("tenant %s: %w", id, err)) + continue + } + // The sweep's own log lines: their detail is in sequences, which only + // this package speaks. Per tenant at Debug, since a sweep reaches + // every tenant each minute; the summary below is the Info line. + switch { + case report.purged: + tenants++ + slog.DebugContext(ctx, "sweeper: purged", + "tenant", id, + "purged_below_seq", report.target, + "ack_floor", report.ackFloor, + "gap_seq", report.gapSeq, + ) + case report.gapSeq == 0: + slog.DebugContext(ctx, "sweeper: all messages within gap window, skipping purge", "tenant", id) + } } - report, err := purgeAcked(ctx, s, consumer, olderThan) - if err != nil { - return false, err + if tenants > 0 { + slog.InfoContext(ctx, "sweeper: purged", "tenants", tenants) } - // The sweep's own log lines: their detail is in sequences, which only - // this package speaks. - switch { - case report.purged: - slog.InfoContext(ctx, "sweeper: purged", - "purged_below_seq", report.target, - "ack_floor", report.ackFloor, - "gap_seq", report.gapSeq, - ) - case report.gapSeq == 0: - slog.DebugContext(ctx, "sweeper: all messages within gap window, skipping purge") - } - return report.purged, nil -} - -// DeadLetterCounts reads the DLQ stream's per-subject counts and keys them by -// table across every tenant (see DeadLetterCounts.Tables). The table filter -// matches that table's unscoped subject under any tenant, so it is applied -// to the parsed topic rather than as a subject filter; a scoped topic counts -// under "table.scope". A subject written before the tenant led it counts -// under its table like any other (parseTopicKey). -func (e *EmbeddedNATS) DeadLetterCounts(ctx context.Context, table string) (DeadLetterCounts, error) { - s, err := e.stream(ctx, dlqStream) + return tenants > 0, errors.Join(errs...) +} + +// DeadLetterCounts reads tenant id's dead-letter stream's per-subject counts +// and keys them by table. The table filter matches that table's unscoped +// subject, so it is applied to the parsed topic rather than as a subject +// filter; a scoped topic counts under "table.scope". +func (e *EmbeddedNATS) DeadLetterCounts(ctx context.Context, id tenant.ID, table string) (DeadLetterCounts, error) { + if _, err := tenant.Parse(string(id)); err != nil { + return DeadLetterCounts{}, fmt.Errorf("tenant: %w", err) + } + s, err := e.stream(ctx, dlqStreamName(id)) if err != nil { if errors.Is(err, jetstream.ErrStreamNotFound) { return DeadLetterCounts{}, fmt.Errorf("%w: %w", ErrNoDeadLetterQueue, err) @@ -474,7 +889,7 @@ func (e *EmbeddedNATS) DeadLetterCounts(ctx context.Context, table string) (Dead return DeadLetterCounts{}, fmt.Errorf("get dlq stream: %w", err) } - state, err := s.state(ctx, dlqAll) + state, err := s.state(ctx, tenantSubjects(dlqPrefix, id)) if err != nil { return DeadLetterCounts{}, fmt.Errorf("dlq stream info: %w", err) } @@ -495,21 +910,20 @@ func (e *EmbeddedNATS) DeadLetterCounts(ctx context.Context, table string) (Dead return counts, nil } -// ReplaySince creates an ephemeral consumer on topic's ingest subject starting at -// since (DeliverByStartTime) and drains it to send until caught up. The -// consumer is ack-less and expires on its own once idle. Caught up is the -// client's no-messages or request-timeout answer to a pull; any other pull -// failure (a closed connection, a deleted consumer) is returned so the caller -// knows the replay ended short rather than empty. A done ctx ends the drain -// between pulls and returns ctx's error. A topic without a valid tenant is -// refused like a publish (see subject): the subject it names is exact, so -// events published before the tenant led the subject are not replayed. +// ReplaySince creates an ephemeral consumer on topic's ingest subject, in its +// tenant's queue, starting at since (DeliverByStartTime) and drains it to send +// until caught up. The consumer is ack-less and expires on its own once idle. +// Caught up is the client's no-messages or request-timeout answer to a pull; +// any other pull failure (a closed connection, a deleted consumer) is returned +// so the caller knows the replay ended short rather than empty. A done ctx +// ends the drain between pulls and returns ctx's error. A topic without a +// valid tenant is refused like a publish (see subject). func (e *EmbeddedNATS) ReplaySince(ctx context.Context, topic Topic, since time.Time, send func(data []byte) bool) error { subj, err := subject(ingestPrefix, topic) if err != nil { return err } - cons, err := e.js.CreateOrUpdateConsumer(ctx, ingestStream, jetstream.ConsumerConfig{ + cons, err := e.js.CreateOrUpdateConsumer(ctx, ingestStreamName(topic.Tenant), jetstream.ConsumerConfig{ FilterSubject: subj, DeliverPolicy: jetstream.DeliverByStartTimePolicy, OptStartTime: &since, diff --git a/internal/mq/embedded_test.go b/internal/mq/embedded_test.go index ab355df7e..fe4fd45ee 100644 --- a/internal/mq/embedded_test.go +++ b/internal/mq/embedded_test.go @@ -2,6 +2,9 @@ package mq import ( "context" + "fmt" + "os" + "path/filepath" "sync" "testing" "time" @@ -13,16 +16,62 @@ import ( "github.com/stretchr/testify/require" ) -// newTestEmbedded spins up an EmbeddedNATS with a temporary store directory -// that is cleaned up by the test framework. -func newTestEmbedded(t *testing.T) *EmbeddedNATS { +// testBudget is the byte budget newTestEmbedded opens each queue at. +const testBudget = 64 << 20 + +// openEmbedded starts an EmbeddedNATS over dir, closed by the test framework. +func openEmbedded(t *testing.T, dir string) *EmbeddedNATS { t.Helper() - e, err := NewEmbedded(t.TempDir(), 64<<20) + e, err := NewEmbedded(dir) require.NoError(t, err) t.Cleanup(func() { _ = e.Close() }) return e } +// newTestEmbedded spins up an EmbeddedNATS over a temporary store directory +// with a queue open for each of tenants — tenant.Default when none is named — +// at testBudget. +func newTestEmbedded(t *testing.T, tenants ...tenant.ID) *EmbeddedNATS { + t.Helper() + e := openEmbedded(t, t.TempDir()) + if len(tenants) == 0 { + tenants = []tenant.ID{tenant.Default} + } + for _, id := range tenants { + require.NoError(t, e.SetMaxBytes(t.Context(), id, testBudget)) + } + return e +} + +// streamConfig is the stored config of the named stream. +func streamConfig(t *testing.T, e *EmbeddedNATS, name string) jetstream.StreamConfig { + t.Helper() + s, err := e.js.Stream(t.Context(), name) + require.NoError(t, err) + return s.CachedInfo().Config +} + +// ackAll consumes every message delivered to consumer on the ingest queue, +// acknowledging each, until n have been acked. +func ackAll(t *testing.T, e *EmbeddedNATS, consumer string, n int) { + t.Helper() + ctx := t.Context() + cons, err := e.CreateConsumer(ctx, ConsumerConfig{Durable: consumer, MaxAckPending: 100}) + require.NoError(t, err) + acked := make(chan error, n) + stop, _, err := cons.Consume(func(msg *Message) { acked <- msg.DoubleAck(ctx) }, 10) + require.NoError(t, err) + t.Cleanup(stop) + for range n { + select { + case err := <-acked: + require.NoError(t, err) + case <-time.After(5 * time.Second): + t.Fatal("timed out waiting for acks") + } + } +} + func TestEmbeddedNATS_PublishSubscribe(t *testing.T) { // No t.Parallel(): each embedded server uses DontListen+InProcessServer, // but starting several in parallel still slows tests unnecessarily. @@ -83,7 +132,7 @@ func TestEmbeddedNATS_PublishHeaders(t *testing.T) { // Read the stored message back raw: the option headers are on the wire // exactly as set, exact-key, with Add appending rather than replacing. - s, err := e.js.Stream(ctx, ingestStream) + s, err := e.js.Stream(ctx, "INGEST_0") require.NoError(t, err) raw, err := s.GetLastMsgForSubject(ctx, "ingest.0.hdr") require.NoError(t, err) @@ -92,24 +141,30 @@ func TestEmbeddedNATS_PublishHeaders(t *testing.T) { assert.Equal(t, []byte("x"), raw.Data) } -func TestNewEmbedded_CreatesBothStreams(t *testing.T) { - e := newTestEmbedded(t) - ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) - defer cancel() - - assert.Equal(t, int64(64<<20), e.MaxBytes()) - - ingest, err := e.js.Stream(ctx, ingestStream) - require.NoError(t, err) - assert.Equal(t, int64(64<<20), ingest.CachedInfo().Config.MaxBytes) - - // The DLQ stream is always present, at a tenth of the budget. - dlq, err := e.js.Stream(ctx, dlqStream) - require.NoError(t, err) - cfg := dlq.CachedInfo().Config - assert.Equal(t, []string{"dlq.>"}, cfg.Subjects) - assert.Equal(t, int64(64<<20)/10, cfg.MaxBytes) - assert.Equal(t, jetstream.DiscardOld, cfg.Discard) +// A tenant's first budget opens its queue: an ingest stream holding its +// subjects alone at the budget, refusing when full, and a dead-letter stream +// at a tenth of it, dropping its oldest when full. No other tenant gets one. +func TestEmbeddedNATS_SetMaxBytes_OpensTheTenantsQueue(t *testing.T) { + e := openEmbedded(t, t.TempDir()) + assert.Zero(t, e.MaxBytes("acme"), "no budget applied yet") + + require.NoError(t, e.SetMaxBytes(t.Context(), "acme", testBudget)) + assert.Equal(t, int64(testBudget), e.MaxBytes("acme")) + + ingest := streamConfig(t, e, "INGEST_acme") + assert.Equal(t, []string{"ingest.acme.>"}, ingest.Subjects) + assert.Equal(t, int64(testBudget), ingest.MaxBytes) + assert.Equal(t, jetstream.DiscardNew, ingest.Discard) + dlq := streamConfig(t, e, "DLQ_acme") + assert.Equal(t, []string{"dlq.acme.>"}, dlq.Subjects) + assert.Equal(t, int64(testBudget)/10, dlq.MaxBytes) + assert.Equal(t, jetstream.DiscardOld, dlq.Discard) + + assert.Zero(t, e.MaxBytes("globex")) + _, err := e.js.Stream(t.Context(), "INGEST_globex") + require.ErrorIs(t, err, jetstream.ErrStreamNotFound, "another tenant's queue opens with its own budget") + + require.Error(t, e.SetMaxBytes(t.Context(), "a.b", testBudget), "a tenant outside the grammar has no queue") } func TestEmbeddedNATS_StreamHandle(t *testing.T) { @@ -120,7 +175,7 @@ func TestEmbeddedNATS_StreamHandle(t *testing.T) { _, err := e.stream(ctx, "NO_SUCH_STREAM") require.Error(t, err, "an unknown stream is an error, not a nil handle") - s, err := e.stream(ctx, ingestStream) + s, err := e.stream(ctx, "INGEST_0") require.NoError(t, err) empty, err := s.state(ctx, "") @@ -196,11 +251,11 @@ func TestEmbeddedNATS_StreamHandle(t *testing.T) { } // TestEmbeddedNATS_CreateConsumer_Config pins the ConsumerConfig → broker -// mapping: AckWait (redelivery timing) and MaxAckPending (ingest backpressure) -// are checkable nowhere else, and a dropped field would compile and pass -// every delivery test. +// mapping on every tenant's queue: AckWait (redelivery timing) and +// MaxAckPending (ingest backpressure, per tenant) are checkable nowhere else, +// and a dropped field would compile and pass every delivery test. func TestEmbeddedNATS_CreateConsumer_Config(t *testing.T) { - e := newTestEmbedded(t) + e := newTestEmbedded(t, "acme", "globex") ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() @@ -211,17 +266,16 @@ func TestEmbeddedNATS_CreateConsumer_Config(t *testing.T) { }) require.NoError(t, err) - s, err := e.js.Stream(ctx, ingestStream) - require.NoError(t, err) - cons, err := s.Consumer(ctx, "cfg") - require.NoError(t, err) - info, err := cons.Info(ctx) - require.NoError(t, err) - assert.Equal(t, "cfg", info.Config.Durable) - assert.Equal(t, "ingest.>", info.Config.FilterSubject, "the consumer sees every topic") - assert.Equal(t, jetstream.AckExplicitPolicy, info.Config.AckPolicy) - assert.Equal(t, 42*time.Second, info.Config.AckWait) - assert.Equal(t, 123, info.Config.MaxAckPending) + for _, stream := range []string{"INGEST_acme", "INGEST_globex"} { + cons, err := e.js.Consumer(ctx, stream, "cfg") + require.NoError(t, err, stream) + cfg := cons.CachedInfo().Config + assert.Equal(t, "cfg", cfg.Durable) + assert.Empty(t, cfg.FilterSubject, "%s: the durable sees the whole of its tenant's stream", stream) + assert.Equal(t, jetstream.AckExplicitPolicy, cfg.AckPolicy) + assert.Equal(t, 42*time.Second, cfg.AckWait) + assert.Equal(t, 123, cfg.MaxAckPending) + } } func TestEmbeddedNATS_ReplaySince(t *testing.T) { @@ -262,7 +316,7 @@ func TestEmbeddedNATS_ReplaySince(t *testing.T) { func TestEmbeddedNATS_DefaultLogger(t *testing.T) { // NewEmbedded without a logger should not panic — it falls back to the // default slog logger. - e, err := NewEmbedded(t.TempDir(), 64<<20) + e, err := NewEmbedded(t.TempDir()) require.NoError(t, err) t.Cleanup(func() { _ = e.Close() }) } @@ -301,29 +355,32 @@ func TestSlogNATSLogger_Levels(t *testing.T) { } func TestEmbeddedNATS_SetMaxBytes(t *testing.T) { - e := newTestEmbedded(t) + e := newTestEmbedded(t, "acme", "globex") ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() - require.NoError(t, e.SetMaxBytes(ctx, 128<<20)) - assert.Equal(t, int64(128<<20), e.MaxBytes()) + require.NoError(t, e.SetMaxBytes(ctx, "acme", 128<<20)) + assert.Equal(t, int64(128<<20), e.MaxBytes("acme")) - ingest, err := e.js.Stream(ctx, ingestStream) - require.NoError(t, err) - assert.Equal(t, int64(128<<20), ingest.CachedInfo().Config.MaxBytes) + ingest := streamConfig(t, e, "INGEST_acme") + assert.Equal(t, int64(128<<20), ingest.MaxBytes) // Everything but the limit is preserved. - assert.Equal(t, []string{"ingest.>"}, ingest.CachedInfo().Config.Subjects) - assert.Equal(t, jetstream.DiscardNew, ingest.CachedInfo().Config.Discard) + assert.Equal(t, []string{"ingest.acme.>"}, ingest.Subjects) + assert.Equal(t, jetstream.DiscardNew, ingest.Discard) - // The DLQ stream follows at a tenth of the budget. - dlq, err := e.js.Stream(ctx, dlqStream) - require.NoError(t, err) - assert.Equal(t, int64(128<<20)/10, dlq.CachedInfo().Config.MaxBytes) - assert.Equal(t, jetstream.DiscardOld, dlq.CachedInfo().Config.Discard) + // The dead-letter stream follows at a tenth of the budget. + dlq := streamConfig(t, e, "DLQ_acme") + assert.Equal(t, int64(128<<20)/10, dlq.MaxBytes) + assert.Equal(t, jetstream.DiscardOld, dlq.Discard) + + // No other tenant's queue moves. + assert.Equal(t, int64(testBudget), e.MaxBytes("globex")) + assert.Equal(t, int64(testBudget), streamConfig(t, e, "INGEST_globex").MaxBytes) + assert.Equal(t, int64(testBudget)/10, streamConfig(t, e, "DLQ_globex").MaxBytes) // The budget already in effect is a no-op, not an error. - require.NoError(t, e.SetMaxBytes(ctx, 128<<20)) - assert.Equal(t, int64(128<<20), e.MaxBytes()) + require.NoError(t, e.SetMaxBytes(ctx, "acme", 128<<20)) + assert.Equal(t, int64(128<<20), e.MaxBytes("acme")) } func TestEmbeddedNATS_SetMaxBytes_DLQFailureRollsBackIngest(t *testing.T) { @@ -331,27 +388,23 @@ func TestEmbeddedNATS_SetMaxBytes_DLQFailureRollsBackIngest(t *testing.T) { ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() - // Put the DLQ stream where the update can't follow: JetStream refuses to - // change a live stream's retention policy, so recreating it as a work - // queue makes the DLQ resize fail after the ingest resize has already - // succeeded. - require.NoError(t, e.js.DeleteStream(ctx, dlqStream)) + // Put the dead-letter stream where the update can't follow: JetStream + // refuses to change a live stream's retention policy, so recreating it as + // a work queue makes the dead-letter resize fail after the ingest resize + // has already succeeded. + require.NoError(t, e.js.DeleteStream(ctx, "DLQ_0")) _, err := e.js.CreateStream(ctx, jetstream.StreamConfig{ - Name: dlqStream, Subjects: []string{dlqAll}, Retention: jetstream.WorkQueuePolicy, MaxBytes: (64 << 20) / 10, + Name: "DLQ_0", Subjects: []string{"dlq.0.>"}, Retention: jetstream.WorkQueuePolicy, MaxBytes: testBudget / 10, }) require.NoError(t, err) - err = e.SetMaxBytes(ctx, 128<<20) + err = e.SetMaxBytes(ctx, tenant.Default, 128<<20) require.Error(t, err) assert.Contains(t, err.Error(), "ingest stream restored to the previous limit") - assert.Equal(t, int64(64<<20), e.MaxBytes(), "the budget in effect is unchanged, so the next call retries both") + assert.Equal(t, int64(testBudget), e.MaxBytes(tenant.Default), "the budget in effect is unchanged, so the next call retries both") - ingest, err := e.js.Stream(ctx, ingestStream) - require.NoError(t, err) - assert.Equal(t, int64(64<<20), ingest.CachedInfo().Config.MaxBytes, "the ingest resize is undone so the pair stays at the previous limit") - dlq, err := e.js.Stream(ctx, dlqStream) - require.NoError(t, err) - assert.Equal(t, int64(64<<20)/10, dlq.CachedInfo().Config.MaxBytes) + assert.Equal(t, int64(testBudget), streamConfig(t, e, "INGEST_0").MaxBytes, "the ingest resize is undone so the pair stays at the previous limit") + assert.Equal(t, int64(testBudget)/10, streamConfig(t, e, "DLQ_0").MaxBytes) } func TestEmbeddedNATS_SetMaxBytes_IngestFailureChangesNothing(t *testing.T) { @@ -359,13 +412,86 @@ func TestEmbeddedNATS_SetMaxBytes_IngestFailureChangesNothing(t *testing.T) { ctx, cancel := context.WithCancel(t.Context()) cancel() // a stop caught mid-reload: the first JetStream call gives up - err := e.SetMaxBytes(ctx, 128<<20) + err := e.SetMaxBytes(ctx, tenant.Default, 128<<20) require.ErrorIs(t, err, context.Canceled) - assert.Equal(t, int64(64<<20), e.MaxBytes()) + assert.Equal(t, int64(testBudget), e.MaxBytes(tenant.Default)) - dlq, err := e.js.Stream(t.Context(), dlqStream) - require.NoError(t, err) - assert.Equal(t, int64(64<<20)/10, dlq.CachedInfo().Config.MaxBytes, "the dlq is not touched when the ingest resize fails") + assert.Equal(t, int64(testBudget)/10, streamConfig(t, e, "DLQ_0").MaxBytes, "the dead-letter stream is not touched when the ingest resize fails") +} + +// A tenant whose queue JetStream will not open — here, a file where its +// dead-letter stream's store would go — is refused on its own: SetMaxBytes +// errors and applies no budget, and a publish is refused as a full queue, +// while every other tenant's queue opens after it (which a store limit at +// the very top of the int64 range would refuse: see NewEmbedded). Once the +// cause is gone, a publish opens the queue at the budget last asked for it. +func TestEmbeddedNATS_SetMaxBytes_AQueueThatCannotOpen(t *testing.T) { + dir := t.TempDir() + // The dead-letter stream is the first of the pair to open. A failed open + // removes what was in the way, so the obstacle is put back before each + // attempt meant to fail. + block := filepath.Join(dir, "jetstream", "$G", "streams", dlqStreamName("acme")) + obstruct := func() { + t.Helper() + require.NoError(t, os.MkdirAll(filepath.Dir(block), 0o750)) + require.NoError(t, os.WriteFile(block, nil, 0o600)) + } + obstruct() + e := openEmbedded(t, dir) + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + + require.Error(t, e.SetMaxBytes(ctx, "acme", testBudget)) + assert.Zero(t, e.MaxBytes("acme"), "no budget applied") + + require.NoError(t, e.SetMaxBytes(ctx, "globex", testBudget), "one tenant's failed open costs the next nothing") + require.NoError(t, e.Publish(ctx, Topic{Tenant: "globex", Table: "t"}, []byte("x"))) + + obstruct() + err := e.Publish(ctx, Topic{Tenant: "acme", Table: "t"}, []byte("x")) + require.ErrorIs(t, err, ErrQueueFull, "the tenant's queue takes nothing; a retry is the answer") + + if err := os.Remove(block); err != nil { + require.ErrorIs(t, err, os.ErrNotExist) + } + require.NoError(t, e.Publish(ctx, Topic{Tenant: "acme", Table: "t"}, []byte("x"))) + assert.Equal(t, int64(testBudget), e.MaxBytes("acme")) +} + +// A budget that shrinks a tenant's dead-letter stream below what it holds +// would have DiscardOld delete the oldest parked rows to fit (#532), so the +// stream keeps what it holds, capped at that, and every row survives. +func TestEmbeddedNATS_SetMaxBytes_NeverShrinksTheDeadLetterQueueBelowWhatItHolds(t *testing.T) { + e := openEmbedded(t, t.TempDir()) + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + require.NoError(t, e.SetMaxBytes(ctx, "acme", 10<<20)) + + payload := make([]byte, 1<<10) + for range 200 { + msg := NewMessage(ctx, Topic{Tenant: "acme", Table: "t"}, payload, time.Now(), nil, nil, nil) + require.NoError(t, e.DeadLetter(ctx, msg)) + } + dlqState := func() jetstream.StreamState { + t.Helper() + s, err := e.js.Stream(ctx, "DLQ_acme") + require.NoError(t, err) + return s.CachedInfo().State + } + held := dlqState().Bytes + require.Greater(t, held, uint64(100<<10), "the rows take more than a tenth of the budget below") + + // Shrunk to a 1 MB budget: a tenth of it is less than the stream holds. + require.NoError(t, e.SetMaxBytes(ctx, "acme", 1<<20)) + assert.Equal(t, int64(1<<20), e.MaxBytes("acme"), "the budget applies") + assert.Equal(t, int64(1<<20), streamConfig(t, e, "INGEST_acme").MaxBytes) + assert.Equal(t, held, uint64(streamConfig(t, e, "DLQ_acme").MaxBytes), "capped at what it holds, not at a tenth") //nolint:gosec // G115: a stream cap is never negative + assert.Equal(t, uint64(200), dlqState().Msgs, "no parked row is deleted") + + // A budget whose tenth covers what it holds applies as usual. + require.NoError(t, e.SetMaxBytes(ctx, "acme", 4<<20)) + assert.Equal(t, int64(4<<20)/10, streamConfig(t, e, "DLQ_acme").MaxBytes) + assert.Equal(t, uint64(200), dlqState().Msgs) } func TestEmbeddedNATS_ReplaySince_PullFailureIsAnError(t *testing.T) { @@ -410,11 +536,11 @@ func TestEmbeddedNATS_ReplaySince_StopsWhenContextIsDone(t *testing.T) { } func TestEmbeddedNATS_DeadLetter(t *testing.T) { - e := newTestEmbedded(t) + e := newTestEmbedded(t, tenant.Default, "acme") ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() - empty, err := e.DeadLetterCounts(ctx, "") + empty, err := e.DeadLetterCounts(ctx, tenant.Default, "") require.NoError(t, err) assert.Equal(t, DeadLetterCounts{Tables: map[string]uint64{}}, empty) @@ -426,52 +552,58 @@ func TestEmbeddedNATS_DeadLetter(t *testing.T) { park(Topic{Tenant: tenant.Default, Table: "default.orders"}, "o1") park(Topic{Tenant: tenant.Default, Table: "default.orders"}, "o2") park(Topic{Tenant: tenant.Default, Table: "users"}, "u1") - // Another tenant's table of the same name counts with it: one queue, one - // count, until the queue is per tenant. So does a subject parked before - // the tenant led it — the queue is never drained, so those stay. + // Another tenant's table of the same name is its own queue and its own + // count. park(Topic{Tenant: "acme", Table: "users"}, "acme-u1") - _, err = e.js.Publish(ctx, "dlq.users", []byte("pre-tenant")) - require.NoError(t, err) - // Parked under the same topic on the DLQ stream, headers intact, and - // nothing lands on the ingest stream. - dlq, err := e.js.Stream(ctx, dlqStream) + // Parked under the same topic on the tenant's dead-letter stream, headers + // intact, and nothing lands on the ingest stream. + dlq, err := e.js.Stream(ctx, "DLQ_0") require.NoError(t, err) raw, err := dlq.GetLastMsgForSubject(ctx, "dlq.0.default%2Eorders") require.NoError(t, err) assert.Equal(t, []byte("o2"), raw.Data) assert.Equal(t, "boom", raw.Header.Get("X-DLQ-Error")) - ingest, err := e.stream(ctx, ingestStream) + ingest, err := e.stream(ctx, "INGEST_0") require.NoError(t, err) st, err := ingest.state(ctx, "") require.NoError(t, err) assert.Zero(t, st.Msgs) - all, err := e.DeadLetterCounts(ctx, "") + all, err := e.DeadLetterCounts(ctx, tenant.Default, "") require.NoError(t, err) - assert.Equal(t, DeadLetterCounts{Tables: map[string]uint64{"default.orders": 2, "users": 3}, Total: 5}, all, "table names come back decoded, summed across tenants") + assert.Equal(t, DeadLetterCounts{Tables: map[string]uint64{"default.orders": 2, "users": 1}, Total: 3}, all, "table names come back decoded, the tenant's own alone") - one, err := e.DeadLetterCounts(ctx, "default.orders") + one, err := e.DeadLetterCounts(ctx, tenant.Default, "default.orders") require.NoError(t, err) - assert.Equal(t, DeadLetterCounts{Tables: map[string]uint64{"default.orders": 2}, Total: 5}, one, "Total is every parked message, filter or not") + assert.Equal(t, DeadLetterCounts{Tables: map[string]uint64{"default.orders": 2}, Total: 3}, one, "Total is every parked message of the tenant, filter or not") - users, err := e.DeadLetterCounts(ctx, "users") + acme, err := e.DeadLetterCounts(ctx, "acme", "") require.NoError(t, err) - assert.Equal(t, map[string]uint64{"users": 3}, users.Tables, "the filter is by table under any tenant, the pre-tenant subject included") + assert.Equal(t, DeadLetterCounts{Tables: map[string]uint64{"users": 1}, Total: 1}, acme) - none, err := e.DeadLetterCounts(ctx, "never_failed") + none, err := e.DeadLetterCounts(ctx, tenant.Default, "never_failed") require.NoError(t, err) assert.Empty(t, none.Tables) } +// A tenant with no queue — one never given a budget on this data directory — +// has nothing parked, which is not the same as a failed read. func TestEmbeddedNATS_DeadLetterCounts_NoQueue(t *testing.T) { e := newTestEmbedded(t) ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() - require.NoError(t, e.js.DeleteStream(ctx, dlqStream)) - _, err := e.DeadLetterCounts(ctx, "") + _, err := e.DeadLetterCounts(ctx, "globex", "") + require.ErrorIs(t, err, ErrNoDeadLetterQueue) + + require.NoError(t, e.js.DeleteStream(ctx, "DLQ_0")) + _, err = e.DeadLetterCounts(ctx, tenant.Default, "") require.ErrorIs(t, err, ErrNoDeadLetterQueue) + + _, err = e.DeadLetterCounts(ctx, "a.b", "") + require.Error(t, err, "an id outside the grammar names no stream") + assert.NotErrorIs(t, err, ErrNoDeadLetterQueue) } func TestEmbeddedNATS_DeadLetterCounts_BrokerFailureIsNotAnEmptyQueue(t *testing.T) { @@ -482,19 +614,19 @@ func TestEmbeddedNATS_DeadLetterCounts_BrokerFailureIsNotAnEmptyQueue(t *testing // A lookup that fails for any reason other than "no such stream" must not // read as an empty queue. e.conn.Close() - _, err := e.DeadLetterCounts(ctx, "") + _, err := e.DeadLetterCounts(ctx, tenant.Default, "") require.Error(t, err) assert.NotErrorIs(t, err, ErrNoDeadLetterQueue) } func TestEmbeddedNATS_DeadLetter_IsAPrefixSwap(t *testing.T) { - e := newTestEmbedded(t) + e := newTestEmbedded(t, "a") ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() // A subject this package would never write (four tokens) still parks - // under the very same tail: nothing on the dead-letter path decodes or - // re-encodes it. + // under the very same tail, in the queue of the tenant its first token + // names: nothing on the dead-letter path decodes or re-encodes it. _, err := e.js.Publish(ctx, "ingest.a.b.c.d", []byte("foreign")) require.NoError(t, err) @@ -513,29 +645,72 @@ func TestEmbeddedNATS_DeadLetter_IsAPrefixSwap(t *testing.T) { assert.Equal(t, Topic{Table: "a.b.c.d"}, msg.Topic(), "a foreign tail is the table of no tenant") require.NoError(t, e.DeadLetter(ctx, msg)) - dlq, err := e.js.Stream(ctx, dlqStream) + dlq, err := e.js.Stream(ctx, "DLQ_a") require.NoError(t, err) raw, err := dlq.GetLastMsgForSubject(ctx, "dlq.a.b.c.d") require.NoError(t, err) assert.Equal(t, []byte("foreign"), raw.Data) } -func TestEmbeddedNATS_Publish_QueueFull(t *testing.T) { - e, err := NewEmbedded(t.TempDir(), 4<<10) +// A dead-letter stream that has gone missing is opened again with its +// tenant's queue, at a tenth of the budget last asked for it, rather than +// leaving the row to be redelivered. +func TestEmbeddedNATS_DeadLetter_ReopensAMissingQueue(t *testing.T) { + e := newTestEmbedded(t, "acme") + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + + require.NoError(t, e.js.DeleteStream(ctx, "DLQ_acme")) + require.NoError(t, e.DeadLetter(ctx, NewMessage(ctx, Topic{Tenant: "acme", Table: "t"}, []byte("x"), time.Now(), nil, nil, nil))) + assert.Equal(t, int64(testBudget)/10, streamConfig(t, e, "DLQ_acme").MaxBytes) + counts, err := e.DeadLetterCounts(ctx, "acme", "") require.NoError(t, err) - t.Cleanup(func() { _ = e.Close() }) + assert.Equal(t, uint64(1), counts.Total) +} + +func TestEmbeddedNATS_Publish_QueueFull(t *testing.T) { + e := openEmbedded(t, t.TempDir()) ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() + require.NoError(t, e.SetMaxBytes(ctx, "acme", 4<<10)) + require.NoError(t, e.SetMaxBytes(ctx, "globex", 4<<10)) // DiscardNew refuses the publish that would pass the byte budget; that is // the backpressure signal, named so callers need not read broker errors. payload := make([]byte, 1<<10) + var err error for range 8 { - if err = e.Publish(ctx, Topic{Tenant: tenant.Default, Table: "full"}, payload); err != nil { + if err = e.Publish(ctx, Topic{Tenant: "acme", Table: "full"}, payload); err != nil { break } } require.ErrorIs(t, err, ErrQueueFull) + + // Only the tenant at its budget is refused: the next one has a budget of + // its own. + require.NoError(t, e.Publish(ctx, Topic{Tenant: "globex", Table: "full"}, payload)) +} + +// A tenant's queue opens at the budget last asked for it when a publish finds +// it missing, and a tenant never given a budget has no queue to publish to: +// that is refused as a full queue, and nothing is opened for it. +func TestEmbeddedNATS_Publish_OpensTheQueueAtTheLastBudget(t *testing.T) { + e := newTestEmbedded(t, "acme") + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + + for _, name := range []string{"INGEST_acme", "DLQ_acme"} { + require.NoError(t, e.js.DeleteStream(ctx, name)) + } + require.NoError(t, e.Publish(ctx, Topic{Tenant: "acme", Table: "t"}, []byte("x"))) + assert.Equal(t, int64(testBudget), streamConfig(t, e, "INGEST_acme").MaxBytes) + assert.Equal(t, int64(testBudget)/10, streamConfig(t, e, "DLQ_acme").MaxBytes) + + err := e.Publish(ctx, Topic{Tenant: "globex", Table: "t"}, []byte("x")) + require.ErrorIs(t, err, ErrQueueFull) + assert.Contains(t, err.Error(), "globex") + _, err = e.js.Stream(ctx, "INGEST_globex") + require.ErrorIs(t, err, jetstream.ErrStreamNotFound) } func TestEmbeddedNATS_PurgeAcked(t *testing.T) { @@ -544,7 +719,7 @@ func TestEmbeddedNATS_PurgeAcked(t *testing.T) { defer cancel() // No consumer yet: the sentinel the sweeper keys its "not yet" warning on. - _, err := e.PurgeAcked(ctx, "buffer", time.Now()) + _, err := e.PurgeAcked(ctx, "buffer", map[tenant.ID]time.Time{tenant.Default: time.Now()}) require.ErrorIs(t, err, ErrConsumerNotFound) for i := range 4 { @@ -570,7 +745,7 @@ func TestEmbeddedNATS_PurgeAcked(t *testing.T) { t.Fatal("timed out waiting for acks") } } - s, err := e.stream(ctx, ingestStream) + s, err := e.stream(ctx, "INGEST_0") require.NoError(t, err) require.Eventually(t, func() bool { floor, err := s.consumerAckFloor(ctx, "buffer") @@ -578,12 +753,12 @@ func TestEmbeddedNATS_PurgeAcked(t *testing.T) { }, 5*time.Second, 20*time.Millisecond) // Everything is acked-or-not but nothing is old enough: keep it all. - purged, err := e.PurgeAcked(ctx, "buffer", time.Now().Add(-time.Hour)) + purged, err := e.PurgeAcked(ctx, "buffer", map[tenant.ID]time.Time{tenant.Default: time.Now().Add(-time.Hour)}) require.NoError(t, err) assert.False(t, purged) // Everything is old enough: only the acked two go. - purged, err = e.PurgeAcked(ctx, "buffer", time.Now().Add(time.Hour)) + purged, err = e.PurgeAcked(ctx, "buffer", map[tenant.ID]time.Time{tenant.Default: time.Now().Add(time.Hour)}) require.NoError(t, err) assert.True(t, purged) st, err := s.state(ctx, "") @@ -592,8 +767,147 @@ func TestEmbeddedNATS_PurgeAcked(t *testing.T) { assert.Equal(t, uint64(2), st.Msgs) } +// Each tenant's queue is purged at its own cutoff and below its own ack +// floor: a tenant keeping an hour of history keeps it while the next one's +// goes, and a tenant the cutoffs do not name — one no longer served — keeps +// no history at all. +func TestEmbeddedNATS_PurgeAcked_EachTenantAtItsOwnCutoff(t *testing.T) { + e := newTestEmbedded(t, "acme", "globex", "initech") + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + + for _, id := range []tenant.ID{"acme", "globex", "initech"} { + for i := range 2 { + require.NoError(t, e.Publish(ctx, Topic{Tenant: id, Table: "p"}, []byte{byte(i)})) + } + } + ackAll(t, e, "buffer", 6) + for _, id := range []tenant.ID{"acme", "globex", "initech"} { + s, err := e.stream(ctx, ingestStreamName(id)) + require.NoError(t, err) + require.Eventually(t, func() bool { + floor, err := s.consumerAckFloor(ctx, "buffer") + return err == nil && floor == 2 + }, 5*time.Second, 20*time.Millisecond, id) + } + + purged, err := e.PurgeAcked(ctx, "buffer", map[tenant.ID]time.Time{ + "acme": time.Now().Add(-time.Hour), // an hour of history: all of it inside the window + "globex": time.Now().Add(time.Hour), // everything older than the cutoff + }) + require.NoError(t, err) + assert.True(t, purged) + msgs := func(id tenant.ID) uint64 { + s, err := e.stream(ctx, ingestStreamName(id)) + require.NoError(t, err) + st, err := s.state(ctx, "") + require.NoError(t, err) + return st.Msgs + } + assert.Equal(t, uint64(2), msgs("acme"), "kept for its own window") + assert.Zero(t, msgs("globex"), "past its own window") + assert.Zero(t, msgs("initech"), "a tenant the cutoffs do not name keeps nothing it has acknowledged") +} + +// The isolation per-tenant queues buy: a tenant at MaxAckPending, or one +// whose handler is stuck, holds back its own delivery and no other tenant's — +// each tenant's messages arrive on a delivery of their own, in order. +func TestEmbeddedNATS_Consume_OneTenantsBacklogDoesNotHoldAnother(t *testing.T) { + e := newTestEmbedded(t, "acme", "globex", "initech") + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + + cons, err := e.CreateConsumer(ctx, ConsumerConfig{Durable: "buffer", MaxAckPending: 2}) + require.NoError(t, err) + release := make(chan struct{}) + var mu sync.Mutex + delivered := map[tenant.ID][]byte{} + stop, _, err := cons.Consume(func(msg *Message) { + id := msg.Topic().Tenant + mu.Lock() + delivered[id] = append(delivered[id], msg.Data[0]) + mu.Unlock() + if id == "initech" { + <-release // never returns until the test ends + } + if id == "globex" { + _ = msg.Ack() // acme never acks: its delivery stops at MaxAckPending + } + }, 12) + require.NoError(t, err) + t.Cleanup(func() { + close(release) + stop() + }) + + for i := range 5 { + for _, id := range []tenant.ID{"acme", "globex", "initech"} { + require.NoError(t, e.Publish(ctx, Topic{Tenant: id, Table: "t"}, []byte{byte(i)})) + } + } + counts := func() (acme, globex, initech int) { + mu.Lock() + defer mu.Unlock() + return len(delivered["acme"]), len(delivered["globex"]), len(delivered["initech"]) + } + require.Eventually(t, func() bool { + acme, globex, initech := counts() + return acme == 2 && globex == 5 && initech == 1 + }, 5*time.Second, 20*time.Millisecond, "globex is delivered in full while acme waits on its acks and initech on its handler") + time.Sleep(200 * time.Millisecond) + acme, globex, initech := counts() + assert.Equal(t, 2, acme, "no more than MaxAckPending unacked, for acme alone") + assert.Equal(t, 5, globex) + assert.Equal(t, 1, initech, "a stuck handler holds back its own tenant alone") + mu.Lock() + defer mu.Unlock() + assert.Equal(t, []byte{0, 1, 2, 3, 4}, delivered["globex"], "in the order published") +} + +// A tenant's queue opened after the consumer started is joined to it: both +// consumer paths deliver its events as they do the queues that were there +// first, whether those were opened in this process or found on disk. +func TestEmbeddedNATS_ConsumersJoinQueuesOpenedLater(t *testing.T) { + e := newTestEmbedded(t, "acme") + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + + worker := make(chan Topic, 4) + cons, err := e.CreateConsumer(ctx, ConsumerConfig{Durable: "buffer", MaxAckPending: 10}) + require.NoError(t, err) + stop, _, err := cons.Consume(func(msg *Message) { + _ = msg.Ack() + worker <- msg.Topic() + }, 4) + require.NoError(t, err) + t.Cleanup(stop) + hub := make(chan Topic, 4) + require.NoError(t, e.Subscribe(ctx, "hub-bridge", func(msg *Message) error { + _ = msg.Ack() + hub <- msg.Topic() + return nil + })) + + require.NoError(t, e.SetMaxBytes(ctx, "globex", testBudget)) + for _, id := range []tenant.ID{"acme", "globex"} { + require.NoError(t, e.Publish(ctx, Topic{Tenant: id, Table: "t"}, []byte("x"))) + } + for name, got := range map[string]chan Topic{"worker": worker, "hub": hub} { + var topics []Topic + for range 2 { + select { + case topic := <-got: + topics = append(topics, topic) + case <-time.After(5 * time.Second): + t.Fatalf("%s: timed out; delivered %v", name, topics) + } + } + assert.ElementsMatch(t, []Topic{{Tenant: "acme", Table: "t"}, {Tenant: "globex", Table: "t"}}, topics, name) + } +} + func TestEmbeddedNATS_Consume_ReportsDeliveryEndingOnItsOwn(t *testing.T) { - e := newTestEmbedded(t) + e := newTestEmbedded(t, "acme", "globex") ctx, cancel := context.WithTimeout(t.Context(), 30*time.Second) defer cancel() @@ -609,22 +923,23 @@ func TestEmbeddedNATS_Consume_ReportsDeliveryEndingOnItsOwn(t *testing.T) { case <-time.After(200 * time.Millisecond): } - // Deleting the durable underneath a running Consume is terminal: the - // client stops the subscription on its own, and no message will ever say - // so. It must reach the caller. - require.NoError(t, e.js.DeleteConsumer(ctx, ingestStream, "doomed")) + // Deleting one tenant's durable underneath a running Consume is terminal + // for that tenant: the client stops the subscription on its own, and no + // message will ever say so. It must reach the caller. + require.NoError(t, e.js.DeleteConsumer(ctx, "INGEST_globex", "doomed")) select { case err := <-failed: require.ErrorIs(t, err, ErrDeliveryEnded) require.ErrorIs(t, err, jetstream.ErrConsumerDeleted, "the broker's reason is kept") + assert.Contains(t, err.Error(), "globex", "the tenant is named") case <-ctx.Done(): t.Fatal("delivery ended underneath the consumer and nothing was reported") } } func TestEmbeddedNATS_Consume_StopIsNotAFailure(t *testing.T) { - e := newTestEmbedded(t) + e := newTestEmbedded(t, "acme", "globex") ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() @@ -639,6 +954,29 @@ func TestEmbeddedNATS_Consume_StopIsNotAFailure(t *testing.T) { t.Fatalf("our own stop was reported as a failure: %v", err) case <-time.After(time.Second): } + // Nor is a queue opened after the stop joined to it. + require.NoError(t, e.SetMaxBytes(ctx, "initech", testBudget)) + _, err = e.js.Consumer(ctx, "INGEST_initech", "stopped") + require.ErrorIs(t, err, jetstream.ErrConsumerNotFound) +} + +// The fetch-ahead asked for is shared by the tenants' queues, at least one +// each, so the rows held client-side stay about what the caller asked for +// however many tenants there are. +func TestFanIn_SharesThePrefetch(t *testing.T) { + t.Parallel() + handles := func(n int) map[tenant.ID]jetstream.Consumer { + m := map[tenant.ID]jetstream.Consumer{} + for i := range n { + m[tenant.ID(fmt.Sprint(i))] = nil + } + return m + } + assert.Equal(t, 500, (&fanIn{prefetch: 500, handles: handles(1)}).share()) + assert.Equal(t, 250, (&fanIn{prefetch: 500, handles: handles(2)}).share()) + assert.Equal(t, 1, (&fanIn{prefetch: 500, handles: handles(1000)}).share(), "at least one per tenant") + assert.Equal(t, 500, (&fanIn{prefetch: 500}).share(), "no tenant yet") + assert.Zero(t, (&fanIn{handles: handles(3)}).share(), "0 leaves the client default") } // Nothing lands on the default tenant by omission (#583): the tenant is a @@ -652,7 +990,7 @@ func TestEmbeddedNATS_Publish_RefusesATopicWithoutATenant(t *testing.T) { require.Error(t, e.Publish(ctx, topic, []byte("x")), "%+v", topic) require.Error(t, e.ReplaySince(ctx, topic, time.Time{}, func([]byte) bool { return true }), "%+v", topic) } - s, err := e.stream(ctx, ingestStream) + s, err := e.stream(ctx, "INGEST_0") require.NoError(t, err) st, err := s.state(ctx, "") require.NoError(t, err) @@ -661,7 +999,7 @@ func TestEmbeddedNATS_Publish_RefusesATopicWithoutATenant(t *testing.T) { // Two tenants, one table name: a replay of one never carries the other's rows. func TestEmbeddedNATS_ReplaySince_IsPerTenant(t *testing.T) { - e := newTestEmbedded(t) + e := newTestEmbedded(t, "acme", "globex") ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() @@ -677,34 +1015,69 @@ func TestEmbeddedNATS_ReplaySince_IsPerTenant(t *testing.T) { assert.Equal(t, []string{"acme1", "acme2"}, got) } -// A message published before the tenant led the subject (#583 story 5) is -// still delivered after the upgrade — the durable consumers filter ingest.> -// — and reads as the default tenant's, so it inserts, streams and parks as -// it did. -func TestEmbeddedNATS_PreTenantSubjectsStillDeliver(t *testing.T) { - e := newTestEmbedded(t) +// A boot over a directory an earlier build wrote deletes the pair of streams +// it kept for every tenant together: their subjects overlap every tenant's, +// so no tenant's queue could open beside them. +func TestNewEmbedded_DeletesTheStreamsAnEarlierBuildShared(t *testing.T) { + dir := t.TempDir() + old, err := NewEmbedded(dir) + require.NoError(t, err) ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) defer cancel() - - _, err := e.js.Publish(ctx, "ingest.events", []byte("old")) + for name, subj := range map[string]string{legacyIngestStream: "ingest.>", legacyDLQStream: "dlq.>"} { + _, err := old.js.CreateStream(ctx, jetstream.StreamConfig{Name: name, Subjects: []string{subj}}) + require.NoError(t, err) + } + _, err = old.js.Publish(ctx, "ingest.events", []byte("pre-tenant")) require.NoError(t, err) + require.NoError(t, old.Close()) - got := make(chan *Message, 1) - require.NoError(t, e.Subscribe(ctx, "upgrade", func(msg *Message) error { - got <- msg - return nil - })) - var msg *Message - select { - case msg = <-got: - case <-time.After(5 * time.Second): - t.Fatal("timed out waiting for delivery") + e := openEmbedded(t, dir) + for _, name := range []string{legacyIngestStream, legacyDLQStream} { + _, err := e.js.Stream(ctx, name) + require.ErrorIs(t, err, jetstream.ErrStreamNotFound, name) } - assert.Equal(t, Topic{Tenant: tenant.Default, Table: "events"}, msg.Topic()) - require.NoError(t, e.DeadLetter(ctx, msg)) - dlq, err := e.js.Stream(ctx, dlqStream) + require.NoError(t, e.SetMaxBytes(ctx, tenant.Default, testBudget)) + require.NoError(t, e.Publish(ctx, Topic{Tenant: tenant.Default, Table: "events"}, []byte("x"))) +} + +// A boot takes stock of the queues on disk: each keeps the budget it last +// had, and a consumer created afterwards is held on every one of them — a +// tenant no longer served, which is never given a budget again, included — +// so what such a tenant had queued still reaches the worker. +func TestNewEmbedded_TakesStockOfTheQueuesOnDisk(t *testing.T) { + dir := t.TempDir() + ctx, cancel := context.WithTimeout(t.Context(), 10*time.Second) + defer cancel() + first, err := NewEmbedded(dir) + require.NoError(t, err) + require.NoError(t, first.SetMaxBytes(ctx, "acme", 8<<20)) + for i := range 2 { + require.NoError(t, first.Publish(ctx, Topic{Tenant: "acme", Table: "t"}, []byte{byte(i)})) + } + require.NoError(t, first.Close()) + + e := openEmbedded(t, dir) + assert.Equal(t, int64(8<<20), e.MaxBytes("acme"), "the budget is read back") + + got := make(chan byte, 2) + cons, err := e.CreateConsumer(ctx, ConsumerConfig{Durable: "buffer", MaxAckPending: 10}) require.NoError(t, err) - raw, err := dlq.GetLastMsgForSubject(ctx, "dlq.events") + stop, _, err := cons.Consume(func(msg *Message) { + _ = msg.Ack() + got <- msg.Data[0] + }, 4) require.NoError(t, err) - assert.Equal(t, []byte("old"), raw.Data, "parked under the tail it arrived on") + t.Cleanup(stop) + for i := range 2 { + select { + case b := <-got: + assert.Equal(t, byte(i), b) + case <-time.After(5 * time.Second): + t.Fatal("the queued rows of a tenant given no budget this boot were not delivered") + } + } + // And a publish to it opens nothing new: the queue is there at its budget. + require.NoError(t, e.Publish(ctx, Topic{Tenant: "acme", Table: "t"}, []byte("x"))) + assert.Equal(t, int64(8<<20), streamConfig(t, e, "INGEST_acme").MaxBytes) } diff --git a/internal/mq/mq.go b/internal/mq/mq.go index 2c5de5662..34620f859 100644 --- a/internal/mq/mq.go +++ b/internal/mq/mq.go @@ -141,24 +141,30 @@ func WithHeader(key, value string) PublishOpt { } } -// ErrQueueFull is returned by Publisher.Publish when the ingest queue is at -// its byte budget and refuses new events — the backpressure signal the API -// turns into a 503 with Retry-After. +// ErrQueueFull is returned by Publisher.Publish when the topic's tenant's +// ingest queue refuses new events — it is at its byte budget, or the tenant +// has no queue open yet — the backpressure signal the API turns into a 503 +// with Retry-After. var ErrQueueFull = errors.New("ingest queue is full") // Publisher appends events to the ingest queue. type Publisher interface { - // Publish stores data as one event on topic. ErrQueueFull when the queue - // is at its byte budget. + // Publish stores data as one event on topic, in the ingest queue of the + // topic's tenant. ErrQueueFull when that queue is at its byte budget, or + // the tenant has no queue open yet (see Broker.SetMaxBytes). Publish(ctx context.Context, topic Topic, data []byte, opts ...PublishOpt) error Close() error } // Subscriber delivers every event on the ingest queue, across all tenants -// and topics. +// and topics: each tenant's in the order it was published, and different +// tenants' concurrently. type Subscriber interface { // Subscribe registers a handler for incoming events under a durable - // consumer named consumerName. + // consumer named consumerName, held on every tenant's queue — those + // opened after Subscribe included. The handler runs on one delivery + // goroutine per tenant, one message at a time, so it must be safe to + // call concurrently for different tenants. // // CONTRACT: If the handler intends to return an error to trigger automatic // redelivery, it MUST NOT manually call msg.Ack() or msg.Nak() beforehand. @@ -181,27 +187,32 @@ type ConsumerConfig struct { // AckWait is the redelivery timeout: a message not acked within it is // delivered again. AckWait time.Duration - // MaxAckPending caps unacked messages broker-side; delivery pauses when - // hit (backpressure). + // MaxAckPending caps unacked messages broker-side, per tenant: delivery + // of a tenant's events pauses when that tenant's unacked ones hit it + // (backpressure), and no other tenant's does. MaxAckPending int } // Consumer is a live durable consumer created by ConsumerManager. type Consumer interface { - // Consume delivers each message to handler on the client's delivery - // goroutine, so a handler that blocks holds delivery back — that is the - // backpressure the ingest worker relies on. Up to prefetch messages are - // fetched ahead (0 = the client default). The returned stop asks delivery - // to end and returns without waiting: a handler invocation already in - // flight, or one for a message already queued client-side, may still run - // after stop returns, so a handler must not write to anything the caller - // tears down right after stopping. + // Consume delivers each message to handler on a delivery goroutine of + // its tenant's: one per tenant, so a tenant's messages arrive in order, + // one at a time, while different tenants' arrive concurrently — handler + // must be safe for that. A handler that blocks holds back its tenant's + // delivery — that is the backpressure the ingest worker relies on. About + // prefetch messages are fetched ahead across the tenants together, at + // least one per tenant (0 = the client default, per tenant). The returned + // stop asks delivery to end and returns without waiting: a handler + // invocation already in flight, or one for a message already queued + // client-side, may still run after stop returns, so a handler must not + // write to anything the caller tears down right after stopping. // // Delivery can also end on its own after Consume has returned: the broker // or the client gives up on the consumer (it was deleted, the connection - // closed). That is reported on failed — exactly one error, and nothing - // once stop has been called — because no message will ever arrive to say - // so. A caller that ignores failed waits forever on a dead consumer. + // closed), or a tenant's queue opened later could not be joined. That is + // reported on failed — exactly one error, and nothing once stop has been + // called — because no message will ever arrive to say so. A caller that + // ignores failed waits forever on a dead consumer. Consume(handler func(msg *Message), prefetch int) (stop func(), failed <-chan error, err error) } @@ -209,7 +220,8 @@ type Consumer interface { // broker's reason when it gave one. var ErrDeliveryEnded = errors.New("consumer delivery ended") -// ConsumerManager creates durable consumers on the ingest queue. A delivered +// ConsumerManager creates durable consumers on the ingest queue, held on +// every tenant's queue — those opened later included. A delivered // Message.Ctx is the ctx given to CreateConsumer: unlike Subscriber, the // consumer path does not extract the trace context carried in the message // headers, because its one consumer (the ingest worker) batches across @@ -220,51 +232,53 @@ type ConsumerManager interface { // DeadLetterer parks messages on the dead-letter queue. type DeadLetterer interface { - // DeadLetter stores msg's data on the dead-letter queue under msg's topic, - // with the headers the options set. It does not ack msg: the caller acks - // once the parking is confirmed, so a failure here leaves the original to - // be redelivered. + // DeadLetter stores msg's data on the dead-letter queue of msg's tenant, + // under msg's topic, with the headers the options set. It does not ack + // msg: the caller acks once the parking is confirmed, so a failure here + // leaves the original to be redelivered. DeadLetter(ctx context.Context, msg *Message, opts ...PublishOpt) error } -// DeadLetterCounts is what is parked on the dead-letter queue. +// DeadLetterCounts is what is parked on one tenant's dead-letter queue. type DeadLetterCounts struct { - // Tables maps table name → parked messages, for the tables asked about, - // summed across tenants: one queue serves every tenant until each has its - // own (#583 story 5b), so one count covers them all. Scope is - // not broken out yet (it is inert until #235): a message parked under a - // scoped topic counts under "table.scope", not under its table. + // Tables maps table name → parked messages, for the tables asked about. + // Scope is not broken out yet (it is inert until #235): a message parked + // under a scoped topic counts under "table.scope", not under its table. Tables map[string]uint64 - // Total is every parked message, whatever the filter. + // Total is every parked message of the tenant, whatever the filter. Total uint64 } // ErrNoDeadLetterQueue is returned by DeadLetterStats.DeadLetterCounts when -// the dead-letter queue does not exist (nothing can have been parked). Any -// other failure to read it is a plain error. +// the tenant has no dead-letter queue (nothing can have been parked for it). +// Any other failure to read it is a plain error. var ErrNoDeadLetterQueue = errors.New("dead-letter queue not found") -// DeadLetterStats reports on the dead-letter queue. +// DeadLetterStats reports on the dead-letter queues. type DeadLetterStats interface { - // DeadLetterCounts counts parked messages per table; a non-empty table - // narrows Tables to that one (its unscoped messages, under any tenant — - // see DeadLetterCounts.Tables). - DeadLetterCounts(ctx context.Context, table string) (DeadLetterCounts, error) + // DeadLetterCounts counts tenant id's parked messages per table — a + // tenant served, rejected, or removed alike, for as long as its queue is + // kept. A non-empty table narrows Tables to that one (its unscoped + // messages). + DeadLetterCounts(ctx context.Context, id tenant.ID, table string) (DeadLetterCounts, error) } // ErrConsumerNotFound is returned by Purger.PurgeAcked when the named -// consumer does not exist (yet). +// consumer does not exist (yet) on a tenant's queue. var ErrConsumerNotFound = errors.New("consumer not found") // Purger reclaims ingest-queue storage. type Purger interface { - // PurgeAcked removes the ingest events that are BOTH acknowledged by the - // named durable consumer (everything before its first unacked event) AND - // stored before olderThan. Either bound alone keeps the event: unacked - // events are not yet written, and recent ones are still needed for replay. - // Reports whether anything was removed. ErrConsumerNotFound when the - // consumer has not been created. - PurgeAcked(ctx context.Context, consumer string, olderThan time.Time) (purged bool, err error) + // PurgeAcked removes, from each tenant's ingest queue, the events that + // are BOTH acknowledged by the named durable consumer (everything before + // its first unacked event) AND stored before that tenant's cutoff in + // olderThan. Either bound alone keeps the event: unacked events are not + // yet written, and recent ones are still needed for replay. A tenant + // olderThan does not name — one no longer served — keeps no history: + // everything it has acknowledged goes. Reports whether anything was + // removed. ErrConsumerNotFound when the consumer has not been created on + // some tenant's queue; the other tenants' are purged all the same. + PurgeAcked(ctx context.Context, consumer string, olderThan map[tenant.ID]time.Time) (purged bool, err error) } // Replayer re-delivers stored events for SSE gap-fill. @@ -278,7 +292,7 @@ type Replayer interface { } // Broker is everything the process wiring needs from the MQ: every interface -// above plus the lifecycle and the byte budget. EmbeddedNATS is the one +// above plus the lifecycle and the byte budgets. EmbeddedNATS is the one // implementation; internal/app depends on this, not on it. type Broker interface { Publisher @@ -288,15 +302,17 @@ type Broker interface { DeadLetterStats Purger Replayer - // SetMaxBytes applies a new byte budget (the hot-reloadable - // mq.max_bytes_gb) to the queues as a whole — how it is split between - // them is the implementation's. On an error the implementation restores - // the previous budget where it can (best effort: the error says when it - // could not, and a canceled ctx abandons the restore too), and MaxBytes - // keeps reporting the previous budget so the next call retries. - // MaxBytes reports the budget last applied in full. - SetMaxBytes(ctx context.Context, maxBytes int64) error - MaxBytes() int64 + // SetMaxBytes applies tenant id's byte budget (its hot-reloadable + // mq.max_bytes_gb) to that tenant's queues — how it is split between them + // is the implementation's — opening them if the tenant has none yet. No + // other tenant's queues are touched. On an error the implementation + // restores the previous budget where it can (best effort: the error says + // when it could not, and a canceled ctx abandons the restore too), and + // MaxBytes keeps reporting the previous budget so the next call retries. + // MaxBytes reports the budget last applied in full for id, 0 when none + // has been. + SetMaxBytes(ctx context.Context, id tenant.ID, maxBytes int64) error + MaxBytes(id tenant.ID) int64 // Stats reports the broker counters the system gauges observe. Stats() (observability.MQStats, error) } diff --git a/internal/mq/subject.go b/internal/mq/subject.go index 8d67ace48..489ad2414 100644 --- a/internal/mq/subject.go +++ b/internal/mq/subject.go @@ -12,25 +12,53 @@ import ( // The embedded broker's naming. Private to this package: everything else // addresses events by Topic. const ( - // ingestStream / dlqStream are the JetStream stream names. Hardcoded — the - // embedded NATS server is private to the WaveHouse process, so there is - // nothing to namespace against. - ingestStream = "WAVEHOUSE" - dlqStream = "WAVEHOUSE_DLQ" + // Each tenant's queue is a pair of JetStream streams named after it: + // INGEST_ and DLQ_. The prefixes differ in their first + // letter, so no tenant id makes one kind's name the other's, and the + // tenant grammar (tenant.Parse: letters, digits, '_' and '-', at most + // tenant.MaxLen bytes) keeps every name inside JetStream's. No namespacing + // beyond that: the embedded server is private to the WaveHouse process. + ingestStreamPrefix = "INGEST_" + dlqStreamPrefix = "DLQ_" + + // legacyIngestStream / legacyDLQStream are the one pair an earlier build + // kept for every tenant. Their subjects (ingest.> and dlq.>) overlap every + // tenant's, and JetStream refuses a stream whose subjects overlap + // another's, so NewEmbedded deletes them. + legacyIngestStream = "WAVEHOUSE" + legacyDLQStream = "WAVEHOUSE_DLQ" // A topic's subject is .
[.]: the tenant id // verbatim — its grammar (tenant.Parse) admits only letters, digits, '_' // and '-', so it is one token as it is — then the table and scope each // as one encoded token. Tenant first so one wildcard selects a tenant's - // traffic (ingest.acme.>). The same topic has the same tail on both - // streams, so parking a message on the DLQ is a prefix swap. + // traffic (ingest.acme.>), which is what the tenant's streams hold. The + // same topic has the same tail on both kinds, so parking a message on the + // dead-letter queue is a prefix swap. ingestPrefix = "ingest." dlqPrefix = "dlq." - - ingestAll = ingestPrefix + ">" // every topic on the ingest stream - dlqAll = dlqPrefix + ">" // every topic on the DLQ stream ) +// ingestStreamName / dlqStreamName name tenant id's two streams. +func ingestStreamName(id tenant.ID) string { return ingestStreamPrefix + string(id) } +func dlqStreamName(id tenant.ID) string { return dlqStreamPrefix + string(id) } + +// tenantSubjects is every subject of tenant id's under prefix: what its +// stream of that kind holds. +func tenantSubjects(prefix string, id tenant.ID) string { return prefix + string(id) + ".>" } + +// streamTenant recovers the tenant a stream name carries under prefix, false +// for any other name: a stream of the other kind, a legacy one, or a name no +// tenant id could have produced. +func streamTenant(prefix, name string) (tenant.ID, bool) { + rest, ok := strings.CutPrefix(name, prefix) + if !ok { + return "", false + } + id, err := tenant.Parse(rest) + return id, err == nil +} + // encodeToken converts any table or scope name into a safe, single NATS // subject token. It preserves alphanumerics and underscores, but // percent-encodes everything else (so '.', ' ', '*' and '>' can never split @@ -66,28 +94,28 @@ func subject(prefix string, t Topic) (string, error) { } // topicKey is the tail of a subject carrying prefix — the key() of the topic -// it was published on, or a one-token tail written before the tenant led the -// subject (see parseTopicKey). A trim, no decoding. +// it was published on. A trim, no decoding. func topicKey(prefix, subj string) string { return strings.TrimPrefix(subj, prefix) } +// keyTenant is the tenant a topic key leads with — the token that decides +// which tenant's stream its subject lands in — whether or not the rest of +// the key parses. +func keyTenant(key string) (tenant.ID, bool) { + first, _, _ := strings.Cut(key, ".") + id, err := tenant.Parse(first) + return id, err == nil +} + // parseTopicKey recovers the Topic from a subject tail. Three tokens are -// tenant, table and scope; two are tenant and table. One token is the form -// this package wrote before the tenant led the subject (#583 story 5) and -// reads as tenant.Default's table: every event of that era was the default -// tenant's, and the durable consumers still deliver them after the upgrade, -// as the dead-letter queue still holds them. A tail this package could not -// have written — more tokens, a token that does not decode, a tenant outside -// the grammar — cannot be split reliably, so the whole of it becomes the -// table of no tenant rather than being dropped. +// tenant, table and scope; two are tenant and table. A tail this package +// could not have written — one token, more than three, a token that does not +// decode, a tenant outside the grammar — cannot be split reliably, so the +// whole of it becomes the table of no tenant rather than being dropped. func parseTopicKey(tail string) Topic { parts := strings.Split(tail, ".") switch len(parts) { - case 1: - if table, err := decodeToken(parts[0]); err == nil && table != "" { - return Topic{Tenant: tenant.Default, Table: table} - } case 2, 3: id, idErr := tenant.Parse(parts[0]) table, tableErr := decodeToken(parts[1]) diff --git a/internal/mq/subject_test.go b/internal/mq/subject_test.go index e4b8bb822..67536e4e0 100644 --- a/internal/mq/subject_test.go +++ b/internal/mq/subject_test.go @@ -1,8 +1,10 @@ package mq import ( + "strings" "testing" + "github.com/Wave-RF/WaveHouse/internal/tenant" "github.com/stretchr/testify/assert" "github.com/stretchr/testify/require" ) @@ -126,15 +128,6 @@ func TestTopicKey_IsInjective(t *testing.T) { assert.Equal(t, Topic{Tenant: "0", Table: "a", Scope: "b"}.key(), Topic{Tenant: "0", Table: "a", Scope: "b"}.key()) } -// The form written before the tenant led the subject (#583 story 5) is the -// default tenant's: it is what the durable consumers deliver across the -// upgrade, and what the dead-letter queue keeps holding after it. -func TestParseTopicKey_PreTenantTailIsTheDefaultTenants(t *testing.T) { - t.Parallel() - assert.Equal(t, Topic{Tenant: "0", Table: "events"}, parseTopicKey("events")) - assert.Equal(t, Topic{Tenant: "0", Table: "default.clicks"}, parseTopicKey("default%2Eclicks")) -} - func TestParseTopicKey_ForeignTailKeepsItself(t *testing.T) { t.Parallel() // Subjects this package did not write still yield one usable topic, of @@ -144,9 +137,50 @@ func TestParseTopicKey_ForeignTailKeepsItself(t *testing.T) { "0.bad%2Gtoken", // a token that does not decode "a%2Eb.events", // a tenant outside the grammar ".events", // a topic whose tenant was never set + "events", // one token: no tenant leads it "bad%2G", // one token that does not decode } { assert.Equal(t, Topic{Table: tail}, parseTopicKey(tail), tail) } assert.Equal(t, Topic{}, parseTopicKey("")) } + +// The tenant a key leads with picks the stream its subject lands in, so it +// is read off the first token whatever the rest of the key holds. +func TestKeyTenant(t *testing.T) { + t.Parallel() + for key, want := range map[string]tenant.ID{"acme.t": "acme", "a.b.c.d": "a", "0.bad%2G": "0"} { + id, ok := keyTenant(key) + assert.True(t, ok, key) + assert.Equal(t, want, id, key) + } + for _, key := range []string{"", ".events", "a%2Eb.events"} { + _, ok := keyTenant(key) + assert.False(t, ok, key) + } +} + +// Every tenant's two streams have names of their own: no id makes one +// kind's name another stream's, none is a stream an earlier build shared, +// and each name gives its tenant back. +func TestStreamNames_NeverCollide(t *testing.T) { + t.Parallel() + ids := []tenant.ID{"0", "acme", "DLQ", "DLQ_acme", "INGEST", "INGEST_acme", "_", "-", "WAVEHOUSE", tenant.ID(strings.Repeat("a", tenant.MaxLen))} + seen := map[string]tenant.ID{legacyIngestStream: "", legacyDLQStream: ""} + for _, id := range ids { + for prefix, name := range map[string]string{ingestStreamPrefix: ingestStreamName(id), dlqStreamPrefix: dlqStreamName(id)} { + other, dup := seen[name] + assert.False(t, dup, "%s names a stream of %q's too", name, other) + seen[name] = id + back, ok := streamTenant(prefix, name) + assert.True(t, ok, name) + assert.Equal(t, id, back, name) + } + } + for _, name := range []string{legacyIngestStream, legacyDLQStream, "INGEST_a.b", "DLQ_"} { + for _, prefix := range []string{ingestStreamPrefix, dlqStreamPrefix} { + _, ok := streamTenant(prefix, name) + assert.False(t, ok, "%s is no tenant's %s stream", name, prefix) + } + } +} diff --git a/internal/settings/settings.go b/internal/settings/settings.go index d0f836e5a..55ec089d3 100644 --- a/internal/settings/settings.go +++ b/internal/settings/settings.go @@ -163,10 +163,10 @@ type TableDedupe struct { } // DLQConfig gates the Dead Letter Queue: whether a row that still fails -// after the row-by-row isolation retry is parked on the WAVEHOUSE_DLQ stream -// (and its original acked) or left unacked to be redelivered indefinitely. -// The stream itself always exists — it is an empty limits-policy stream -// until something lands on it — so the switch is purely behavioral and +// after the row-by-row isolation retry is parked on the tenant's dead-letter +// queue (and its original acked) or left unacked to be redelivered +// indefinitely. The queue exists from the moment the tenant is first served — +// empty until something lands on it — so the switch is purely behavioral and // resolves per table through the same override cascade as dedupe. type DLQConfig struct { Enabled *bool `json:"enabled"` @@ -219,14 +219,15 @@ type StreamConfig struct { GapWindowMinutes *int `json:"gap_window_minutes"` } -// MQConfig sizes the embedded JetStream streams on disk. +// MQConfig sizes the tenant's message queue on disk. type MQConfig struct { - // MaxBytesGB caps the WAVEHOUSE ingest stream (the DLQ stream gets a - // tenth of it). Must be >= 1. A reload updates the live streams in - // place: growing takes effect immediately; shrinking below what is - // currently buffered makes the ingest stream refuse new publishes - // (DiscardNew → 503 backpressure) until the worker drains it — nothing - // already buffered is dropped. + // MaxBytesGB caps the tenant's ingest queue (its dead-letter queue gets a + // tenth of it). Must be >= 1. A reload updates the live queues in place: + // growing takes effect immediately; shrinking below what is currently + // buffered makes the ingest queue refuse new publishes (DiscardNew → 503 + // backpressure) until the worker drains it — nothing already buffered is + // dropped — and a dead-letter queue holding more than a tenth of the new + // budget keeps what it holds rather than dropping its oldest rows. MaxBytesGB *int `json:"max_bytes_gb"` } diff --git a/internal/settings/store.go b/internal/settings/store.go index d14939172..f68a2fbf3 100644 --- a/internal/settings/store.go +++ b/internal/settings/store.go @@ -193,7 +193,7 @@ func (s *Store) GapWindow() time.Duration { return time.Duration(*s.doc().Config.Stream.GapWindowMinutes) * time.Minute } -// MQMaxBytes returns the ingest stream's disk budget in bytes. +// MQMaxBytes returns the disk budget of the tenant's ingest queue in bytes. func (s *Store) MQMaxBytes() int64 { return int64(*s.doc().Config.MQ.MaxBytesGB) << 30 } diff --git a/internal/stream/subscriber.go b/internal/stream/subscriber.go index 548834928..5952aa1ad 100644 --- a/internal/stream/subscriber.go +++ b/internal/stream/subscriber.go @@ -55,15 +55,17 @@ type Subscriber struct { // which reads nothing here but may run alongside the fan-out. // // It does NOT make Hub.deliver's check→send→record sequence atomic, and - // deliver does not need it to be: Broadcast runs on ONE goroutine — the - // single jetstream Consume callback the hub bridge registers in - // internal/app, invoked inline per message — so no two events race - // to announce the same connection's columns. A future change that fans - // Broadcast out across goroutines must hold a lock across that whole - // sequence, or two events will both send an announcement (harmless) while a - // third slips a row between a check and its record (not harmless: the client - // zips it against the previous list). Replay does not touch this field at - // all — it tracks drift in its own closure; see Hub.ReplayProjector. + // deliver does not need it to be: the hub bridge registered in + // internal/app calls Broadcast inline per message, on one delivery + // goroutine per tenant (mq.Subscriber), and a connection subscribes to one + // tenant's topic — so all of its events come from that one goroutine, and + // no two events race to announce the same connection's columns. A future + // change that fans one tenant's Broadcasts out across goroutines must hold + // a lock across that whole sequence, or two events will both send an + // announcement (harmless) while a third slips a row between a check and its + // record (not harmless: the client zips it against the previous list). + // Replay does not touch this field at all — it tracks drift in its own + // closure; see Hub.ReplayProjector. schemaMu sync.Mutex // lastSchema is the signature of the column list most recently announced to // this connection ("" ⇒ none yet). Rows travel positionally, so a client that diff --git a/internal/testutil/mocks.go b/internal/testutil/mocks.go index 1b8cf358f..44f3ebe61 100644 --- a/internal/testutil/mocks.go +++ b/internal/testutil/mocks.go @@ -221,10 +221,10 @@ type MockPurger struct { // PurgeCall records one PurgeAcked call. type PurgeCall struct { Consumer string - OlderThan time.Time + OlderThan map[tenant.ID]time.Time } -func (m *MockPurger) PurgeAcked(_ context.Context, consumer string, olderThan time.Time) (bool, error) { +func (m *MockPurger) PurgeAcked(_ context.Context, consumer string, olderThan map[tenant.ID]time.Time) (bool, error) { m.mu.Lock() defer m.mu.Unlock() m.Calls = append(m.Calls, PurgeCall{Consumer: consumer, OlderThan: olderThan}) @@ -233,13 +233,21 @@ func (m *MockPurger) PurgeAcked(_ context.Context, consumer string, olderThan ti // ── Mock mq.DeadLetterStats ────────────────────────────────────── -// MockDeadLetterStats implements mq.DeadLetterStats with a canned answer. +// MockDeadLetterStats implements mq.DeadLetterStats with a canned answer, +// recording the tenant and table each call asked about. type MockDeadLetterStats struct { Counts mq.DeadLetterCounts Err error + + mu sync.Mutex + Tenant tenant.ID + Table string } -func (m *MockDeadLetterStats) DeadLetterCounts(context.Context, string) (mq.DeadLetterCounts, error) { +func (m *MockDeadLetterStats) DeadLetterCounts(_ context.Context, id tenant.ID, table string) (mq.DeadLetterCounts, error) { + m.mu.Lock() + defer m.mu.Unlock() + m.Tenant, m.Table = id, table return m.Counts, m.Err } diff --git a/internal/testutil/testutil.go b/internal/testutil/testutil.go index 371cdcd48..db1196869 100644 --- a/internal/testutil/testutil.go +++ b/internal/testutil/testutil.go @@ -14,6 +14,7 @@ import ( "github.com/stretchr/testify/require" "github.com/Wave-RF/WaveHouse/internal/discovery" + "github.com/Wave-RF/WaveHouse/internal/mq" "github.com/Wave-RF/WaveHouse/internal/tenant" ) @@ -39,6 +40,24 @@ func NewTestSchemaRegistry(t testing.TB, tables []*discovery.TableSchema) *disco // hardcoding the same literal twice. const TestServerVersion = "24.8.1.1" +// NewEmbeddedMQ starts the embedded broker over a temporary directory, closed +// by the test framework, with a queue open for each of tenants — +// tenant.Default when none is named — at maxBytes: a tenant has a queue once +// its budget is applied, as the wiring does for every tenant it serves. +func NewEmbeddedMQ(t testing.TB, maxBytes int64, tenants ...tenant.ID) *mq.EmbeddedNATS { + t.Helper() + emb, err := mq.NewEmbedded(t.TempDir()) + require.NoError(t, err) + t.Cleanup(func() { _ = emb.Close() }) + if len(tenants) == 0 { + tenants = []tenant.ID{tenant.Default} + } + for _, id := range tenants { + require.NoError(t, emb.SetMaxBytes(context.Background(), id, maxBytes)) + } + return emb +} + // schemaConn is a mock driver.Conn serving exactly the queries Refresh issues: // the SELECT timezone() (always "UTC") and SELECT version() probes, the // system.columns scan (rows synthesized from tables), and the system.tables DDL From 3022b926392f9cdfa16931ada447607b94fe09b7 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Thu, 24 Sep 2026 23:27:32 -0400 Subject: [PATCH 02/79] feat(config): choose each layer's implementation at boot mq.backend, cache.backend, dedupe.backend and coord.backend select each layer's implementation; only today's in-process one exists per layer and it is the default. Validate refuses an unknown value, internal/app picks the implementation in one switch per layer, data_dir is probed only when a selected backend keeps state there, and boot logs Config.Warnings. Part of #613. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 1 + cmd/wavehouse/main.go | 19 ++- config.yaml | 10 ++ docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/configuration.mdx | 31 +++- docs/src/content/docs/settings-directory.mdx | 2 +- internal/app/app.go | 4 + internal/app/app_test.go | 27 +++- internal/app/wire.go | 63 ++++++-- internal/config/backends.go | 143 ++++++++++++++++ internal/config/backends_test.go | 161 +++++++++++++++++++ internal/config/config.go | 12 +- internal/config/config_test.go | 10 +- tests/integration/setup_test.go | 5 +- tests/integration/tenants_test.go | 5 +- 15 files changed, 453 insertions(+), 42 deletions(-) create mode 100644 internal/config/backends.go create mode 100644 internal/config/backends_test.go diff --git a/CHANGELOG.md b/CHANGELOG.md index 23c0c7152..9f2e821bc 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,6 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added +- **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block, which until that backend lands is an unknown key and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name every backend: the zero value is not the default, and `app.New` refuses it. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/cmd/wavehouse/main.go b/cmd/wavehouse/main.go index 042257e2b..3a1304f0b 100644 --- a/cmd/wavehouse/main.go +++ b/cmd/wavehouse/main.go @@ -171,14 +171,17 @@ func run(ctx context.Context) int { return 1 } - // data_dir must be writable before anything dials out, so the refusal - // (and, for the typical cause — a bind mount owned by root rather than - // UID 65532 — the remediation) lands at the top of the log rather than - // after ClickHouse discovery. NATS and Pebble still fail loud on their - // own if the directory changes underneath us. - if err := config.CheckDataDir(cfg.DataDir); err != nil { - logger.Error("check data_dir", "error", err) - return 1 + // data_dir, when a selected backend keeps state there, must be writable + // before anything dials out, so the refusal (and, for the typical cause — + // a bind mount owned by root rather than UID 65532 — the remediation) + // lands at the top of the log rather than after ClickHouse discovery. + // NATS and Pebble still fail loud on their own if the directory changes + // underneath us. + if cfg.NeedsDataDir() { + if err := config.CheckDataDir(cfg.DataDir); err != nil { + logger.Error("check data_dir", "error", err) + return 1 + } } a, err := app.New(ctx, app.Options{ diff --git a/config.yaml b/config.yaml index 53a435029..65c384eaa 100644 --- a/config.yaml +++ b/config.yaml @@ -43,9 +43,19 @@ clickhouse: password: "" max_total_conns: 0 # ceiling on open native connections across pools; 0 = none +# Each layer's implementation, chosen at boot. Only the in-process backend +# exists for each today, and it is the default. +mq: + backend: embedded # NATS JetStream under /nats +dedupe: + backend: pebble # Pebble under /pebble +coord: + backend: local + # In-process L1 cache size. The query time-bucket # (query.timestamp_bucket_seconds) is a settings key. cache: + backend: local l1_max_cost: 67108864 # Auth has no on/off switch — the JWT middleware always runs. A request with no diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 6eaf3d54e..938751ff8 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -90,7 +90,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability, ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe` builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ` hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 8a709426f..8d311adc5 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -17,7 +17,7 @@ WaveHouse is configured via a YAML file with environment variable overrides. All 2. Environment variables override any values from the YAML file. 3. If no config file exists, all values are read from environment variables. Every key has a default except `settings.dir` (`WH_SETTINGS_DIR`), which must be set either way. 4. Both sources are **strict**. A YAML key this page doesn't list — a typo, or a tunable that has moved to the settings directory (`dlq.enabled`, `clickhouse.addr`, `stream.*`, a leftover `policy:` or `pipes:` block, …) — refuses to boot and names every offending key, so nothing is read, ignored, and believed. A `WH_*` environment variable that binds to no key on this page (`WH_DEDUPE_ENABLED`, `WH_CH_ADDR`, a misspelling) refuses to boot the same way. Two variables have no YAML key and are exempt because they are not config keys at all but process-level settings `main` reads directly: `WH_CONFIG` (below), which locates the file, and `WH_LOG_LEVEL`. Only the `WH_` prefix is checked, since the environment always carries names that aren't WaveHouse's. One outside source does share the prefix. Kubernetes injects `{SERVICE}_SERVICE_HOST`, `{SERVICE}_PORT`, and similar link variables into every pod in a Service's own namespace, for each Service with a cluster IP that existed before the pod started (a headless Service injects nothing, and a Service in another namespace is harmless). The name is uppercased with `-` mapped to `_`, so a Service named `wh` produces `WH_SERVICE_HOST` and `WH_PORT`, one named `wh-foo` produces `WH_FOO_SERVICE_HOST` and `WH_FOO_PORT`, and either way the pod refuses to boot on its next restart. Set `enableServiceLinks: false` on the pod spec, or name the Service something else. The error says so. -5. Before anything dials out, `data_dir` is probed, and boot refuses on any of these: the value is empty; the path exists but is not a directory; the path, or any component above it, is a dangling symlink (a mount that never came up); the directory exists but the process cannot write to it; the directory is absent and its nearest existing ancestor is not writable, so it could not be created. The probe runs before ClickHouse discovery, so the refusal lands at the top of the log, and a permission denial — on the write probe, or on reaching the path at all through a parent without search permission — carries the UID-65532 remediation, since a bind mount owned by root is the typical cause. +5. Before anything dials out, `data_dir` is probed — when a selected [backend](#backends) keeps state there, as the in-process `mq` and `dedupe` backends do — and boot refuses on any of these: the value is empty; the path exists but is not a directory; the path, or any component above it, is a dangling symlink (a mount that never came up); the directory exists but the process cannot write to it; the directory is absent and its nearest existing ancestor is not writable, so it could not be created. The probe runs before ClickHouse discovery, so the refusal lands at the top of the log, and a permission denial — on the write probe, or on reaching the path at all through a parent without search permission — carries the UID-65532 remediation, since a bind mount owned by root is the typical cause. Boot is the validator for this half of configuration: there is no dry run, and a refused boot with the offending key, variable, or path named in the error is the loud signal. The hot-reloadable half has a dry run — `wavehouse validate` — because it is edited under a running server; boot config only ever takes effect through a restart, so the restart is where it is checked. @@ -37,6 +37,19 @@ This page is boot config only — what the platform operator owns (wiring, lifec | --- | --- | ------- | ----------- | | `data_dir` | `WH_DATA_DIR` | `./data` | Root directory for embedded state. NATS JetStream lives at `/nats`; Pebble, holding every tenant's dedupe store while any tenant has dedupe enabled, at `/pebble`. Subdirectory names are conventions, not config — one knob, one mount. **In a container this MUST resolve to a host-backed volume**; the relative default is for local binary use. WaveHouse logs a startup `WARN` when the directory is missing or empty (no prior state). See [Persistent Storage](/deployment#persistent-storage-required-for-containers). | +### Backends + +Each layer's implementation is chosen once, at boot. Today every layer has one backend, the in-process one, and it is the default, so a config that sets none of these keys runs as it always has. A value this build has no backend for refuses boot and names the valid ones. + +| YAML Key | Env Var | Default | Description | +| --- | --- | ------- | ----------- | +| `mq.backend` | `WH_MQ_BACKEND` | `embedded` | The message queue. `embedded`: NATS JetStream inside this process, under `/nats`. It listens on no port, so no other process can reach its queue. | +| `cache.backend` | `WH_CACHE_BACKEND` | `local` | The query-result cache. `local`: in this process, sized by `cache.l1_max_cost`. | +| `dedupe.backend` | `WH_DEDUPE_BACKEND` | `pebble` | Where ingest dedupe keeps the event ids it has seen. `pebble`: in this process, under `/pebble`, open while any tenant has dedupe on. | +| `coord.backend` | `WH_COORD_BACKEND` | `local` | Where the leases for work only one process may do at a time are held. `local`: in this process. | + +Settings for one backend go in a sub-block named after it, `.`, read only when that backend is selected. A sub-block for a backend this build does not have is an unknown key and refuses boot. `mq` and `dedupe` also appear in the settings directory's `config.json`, with different keys (`mq.max_bytes_gb`, `dedupe.enabled`, …): those are per-tenant tunables and stay there, and one written in `config.yaml` refuses boot as an unknown key. + ### Server | YAML Key | Env Var | Default | Description | @@ -97,6 +110,8 @@ WaveHouse's per-role caps are sent as per-query `SETTINGS` on its connection, so ### Message Queue (NATS) +This section describes the `embedded` [backend](#backends), the only one today. + Each tenant's queue has its own disk budget, `mq.max_bytes_gb`, a hot-reloadable key in the [Settings Directory](/settings-directory#message-queue) — there is no boot-config knob for it. **Durability.** The embedded server runs with JetStream `SyncAlways`, so every event is `fsync`'d to disk before `POST /v1/ingest` returns `200`. This makes your storage's `fsync` latency your ingest latency floor — see [Durability & Storage](/durability) to check whether your substrate can sustain it. There is no knob to relax this today ([#139](https://github.com/Wave-RF/WaveHouse/issues/139) tracks a configurable group-commit interval). @@ -191,9 +206,19 @@ clickhouse: # headers and pool sizes are settings (config.json) max_total_conns: 0 # ceiling on open native connections; 0 = none +mq: + backend: embedded # in-process NATS JetStream under /nats + cache: + backend: local l1_max_cost: 67108864 +dedupe: + backend: pebble # in-process Pebble under /pebble + +coord: + backend: local + auth: jwt_secret: change-me-in-production # jwks_url and role_claim are settings (config.json) operator_key: "" # non-JWT full-access operator credential (Authorization: Operator , or X-Operator-Key); empty disables @@ -239,7 +264,11 @@ WH_SERVER_SHUTDOWN_TIMEOUT=10 WH_CH_PASSWORD= WH_CH_MAX_TOTAL_CONNS=0 +WH_MQ_BACKEND=embedded +WH_CACHE_BACKEND=local WH_CACHE_L1_MAX_COST=67108864 +WH_DEDUPE_BACKEND=pebble +WH_COORD_BACKEND=local WH_AUTH_JWT_SECRET=change-me-in-production WH_AUTH_OPERATOR_KEY= diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index c0e7a19b7..561354627 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -179,7 +179,7 @@ The tenant tunables. Every key is required (a missing one is a validation error) } ``` -What stays in boot config is only what cannot change under a running process — resource sizing (`data_dir`, `cache.l1_max_cost`, `clickhouse.max_total_conns`), the listeners, the observability exporters — and the **secrets**: `clickhouse.password`, `auth.jwt_secret`, `auth.operator_key`. Secrets never belong in a tracked JSON file, so they stay in the environment and are combined with the wiring here on every (re)connect; rotating one is a restart. See [Configuration](/configuration). Everything else lives here and reloads. +What stays in boot config is only what cannot change under a running process — the implementation each layer runs on (`mq.backend`, `cache.backend`, `dedupe.backend`, `coord.backend`), resource sizing (`data_dir`, `cache.l1_max_cost`, `clickhouse.max_total_conns`), the listeners, the observability exporters — and the **secrets**: `clickhouse.password`, `auth.jwt_secret`, `auth.operator_key`. Secrets never belong in a tracked JSON file, so they stay in the environment and are combined with the wiring here on every (re)connect; rotating one is a restart. See [Configuration](/configuration). Everything else lives here and reloads. ## Deduplication diff --git a/internal/app/app.go b/internal/app/app.go index dc15bf5d7..726dad19a 100644 --- a/internal/app/app.go +++ b/internal/app/app.go @@ -170,6 +170,10 @@ func New(ctx context.Context, opts Options) (app *App, err error) { return nil, err } a.wireObservability(ctx) + // After observability, so an OTLP log pipeline carries them too. + for _, w := range a.cfg.Warnings() { + slog.Warn(w) + } if err := a.wireClickHouse(); err != nil { return nil, err } diff --git a/internal/app/app_test.go b/internal/app/app_test.go index 2ecd77d98..10658ddc3 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -95,7 +95,10 @@ func testConfig(t *testing.T, settingsDir string) *config.Config { return &config.Config{ DataDir: t.TempDir(), Server: config.Server{Port: closedPort(t), ShutdownTimeout: 2}, - Cache: config.Cache{L1MaxCost: 1 << 20}, + MQ: config.MQ{Backend: config.MQEmbedded}, + Cache: config.Cache{Backend: config.CacheLocal, L1MaxCost: 1 << 20}, + Dedupe: config.Dedupe{Backend: config.DedupePebble}, + Coord: config.Coord{Backend: config.CoordLocal}, Auth: config.Auth{JWTSecret: "unit-test-secret"}, Settings: config.Settings{Dir: settingsDir}, } @@ -535,6 +538,28 @@ func TestNew_NestedDedupeStoreFollowsEachTenant(t *testing.T) { assert.False(t, restored.Open(), "Close releases every open store") } +// Validate refuses a backend no layer has a case for, so the switch's default +// is reached only by a Config built by hand; it must refuse boot, not wire +// nothing. +func TestNew_RefusesALayerWithoutABackend(t *testing.T) { + for _, tc := range []struct { + key string + unset func(*config.Config) + }{ + {"dedupe.backend", func(c *config.Config) { c.Dedupe.Backend = "" }}, + {"mq.backend", func(c *config.Config) { c.MQ.Backend = "" }}, + {"cache.backend", func(c *config.Config) { c.Cache.Backend = "" }}, + } { + t.Run(tc.key, func(t *testing.T) { + guardGlobals(t) + cfg := testConfig(t, writeSettings(t, nil)) + tc.unset(cfg) + _, err := New(t.Context(), Options{Config: cfg}) + require.ErrorContains(t, err, tc.key+` "" has no wiring`) + }) + } +} + // A Pebble instance that cannot open follows the registry's own rule for the // shape: a flat directory refuses boot, like every other store, and a nested // one fails closed for every tenant with dedupe on, since they share the diff --git a/internal/app/wire.go b/internal/app/wire.go index 60cbdcd82..d53720253 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -471,7 +471,18 @@ func (a *App) wireDiscovery(ctx context.Context) { a.add(component{name: "schema discovery", close: d.close}) } -// wireDedupe builds the dedupe stores: one per tenant (#583 story 7), each +// wireDedupe builds the dedupe stores — the one place the implementation is +// chosen. +func (a *App) wireDedupe() error { + switch b := a.cfg.Dedupe.Backend; b { + case config.DedupePebble: + return a.wirePebbleDedupe() + default: + return unreachableBackend("dedupe.backend", b) + } +} + +// wirePebbleDedupe builds the dedupe stores: one per tenant (#583 story 7), each // following its own tenant's hot-reloadable dedupe.enabled, over the // embedded Pebble implementation, which is handed data_dir and decides the // rest: every tenant's seen ids in one instance there, open while any @@ -488,7 +499,7 @@ func (a *App) wireDiscovery(ctx context.Context) { // than silently publishing un-deduped, since the files asked for dedupe; // nested fails closed the same way at boot too, for every tenant with // dedupe on, the next reload retrying, so it never costs the process. -func (a *App) wireDedupe() error { +func (a *App) wirePebbleDedupe() error { nested := a.tenants.Nested() embedded := dedupe.NewEmbedded(a.cfg.DataDir) stores := dedupe.NewStores(embedded.Tenant) @@ -530,10 +541,20 @@ func (a *App) wireDedupe() error { return nil } -// wireMQ starts the MQ — the embedded NATS under data_dir/nats, the one -// place the implementation is chosen; everything after it sees mq.Broker — -// and hands it each served tenant's mq.max_bytes_gb, which opens that -// tenant's queue the first time. The budget is hot-reloadable: after every +// wireMQ starts the MQ — the one place the implementation is chosen; +// everything after it sees mq.Broker. +func (a *App) wireMQ() error { + switch b := a.cfg.MQ.Backend; b { + case config.MQEmbedded: + return a.wireEmbeddedMQ() + default: + return unreachableBackend("mq.backend", b) + } +} + +// wireEmbeddedMQ starts the embedded NATS under data_dir/nats and hands it +// each served tenant's mq.max_bytes_gb, which opens that tenant's queue the +// first time. The budget is hot-reloadable: after every // reload the registry applies, each served tenant's is handed over again, // and the MQ owns how it is split across the tenant's queues and keeps them // consistent (see mq.Broker.SetMaxBytes). A tenant no longer served keeps @@ -543,7 +564,7 @@ func (a *App) wireDedupe() error { // previous budget; a nested directory logs it at boot too, so it never costs // the process — the tenant's ingest answers 503 until a reload opens its // queue. The hook is registered before the boot apply, as the dedupe one is. -func (a *App) wireMQ() error { +func (a *App) wireEmbeddedMQ() error { dir := filepath.Join(a.cfg.DataDir, "nats") config.WarnIfFreshDataDir("nats", dir) var broker mq.Broker @@ -591,16 +612,28 @@ func (a *App) wireMQ() error { return nil } -// wireCache opens the L1 cache — the only tier in standalone mode. +// wireCache opens the query-result cache — the one place the implementation +// is chosen. func (a *App) wireCache() error { - l1, err := cache.NewLocal(a.cfg.Cache.L1MaxCost) - if err != nil { - return fmt.Errorf("cache init: %w", err) + switch b := a.cfg.Cache.Backend; b { + case config.CacheLocal: + l1, err := cache.NewLocal(a.cfg.Cache.L1MaxCost) + if err != nil { + return fmt.Errorf("cache init: %w", err) + } + a.cache = l1 + a.add(component{name: "cache", close: withoutContext(l1.Close)}) + return nil + default: + return unreachableBackend("cache.backend", b) } - // TODO: eventually this is where we can switch between ristretto, redis, tiered (both), etc - a.cache = l1 - a.add(component{name: "cache", close: withoutContext(l1.Close)}) - return nil +} + +// unreachableBackend is each layer switch's default case. config.Validate +// refuses a backend with no case, so reaching it means a Config built by hand +// without one (the zero value is not the default), or a case missing here. +func unreachableBackend[T ~string](key string, got T) error { + return fmt.Errorf("%s %q has no wiring: a Config built without config.Load must name every backend", key, got) } // wireSweeper adds the active sweeper — purges messages that are both diff --git a/internal/config/backends.go b/internal/config/backends.go new file mode 100644 index 000000000..8328cea61 --- /dev/null +++ b/internal/config/backends.go @@ -0,0 +1,143 @@ +package config + +import ( + "fmt" + "slices" + "strings" +) + +// Each layer's implementation is chosen here, once, at boot: `.backend` +// names it, and the default is today's in-process one. Settings for one +// backend go in `.`, a sub-block read only when that backend +// is selected. Adding a backend is its constant in the layer's list, a case +// in the layer's validate for its sub-block, and a case in the layer's +// wire function in internal/app — nothing else in Validate changes. + +// MQBackend names the message queue implementation. +type MQBackend string + +// MQEmbedded is the NATS JetStream server inside this process, under +// /nats. +const MQEmbedded MQBackend = "embedded" + +var mqBackends = []MQBackend{MQEmbedded} + +// MQ selects the message queue. The per-tenant byte budget, mq.max_bytes_gb, +// is a settings-directory key, not this block's. +type MQ struct { + Backend MQBackend `yaml:"backend" env:"WH_MQ_BACKEND" env-default:"embedded"` +} + +func (m MQ) validate() error { + return checkBackend("mq.backend", "WH_MQ_BACKEND", m.Backend, mqBackends) +} + +// CacheBackend names the query-result cache implementation. +type CacheBackend string + +// CacheLocal is the in-process Ristretto cache, sized by cache.l1_max_cost. +const CacheLocal CacheBackend = "local" + +var cacheBackends = []CacheBackend{CacheLocal} + +// Cache selects and sizes the query-result cache. The time-range bucket +// structured queries normalize to is a settings-directory key +// (query.timestamp_bucket_seconds) — query shaping, not process memory. +type Cache struct { + Backend CacheBackend `yaml:"backend" env:"WH_CACHE_BACKEND" env-default:"local"` + L1MaxCost int64 `yaml:"l1_max_cost" env:"WH_CACHE_L1_MAX_COST" env-default:"67108864"` +} + +func (c Cache) validate() error { + return checkBackend("cache.backend", "WH_CACHE_BACKEND", c.Backend, cacheBackends) +} + +// DedupeBackend names where ingest dedupe keeps the ids it has seen. +type DedupeBackend string + +// DedupePebble is the Pebble instance inside this process, under +// /pebble, opened while any tenant has dedupe on. +const DedupePebble DedupeBackend = "pebble" + +var dedupeBackends = []DedupeBackend{DedupePebble} + +// Dedupe selects the dedupe store. Whether a tenant dedupes, and on which +// field, are settings-directory keys, not this block's. +type Dedupe struct { + Backend DedupeBackend `yaml:"backend" env:"WH_DEDUPE_BACKEND" env-default:"pebble"` +} + +func (d Dedupe) validate() error { + return checkBackend("dedupe.backend", "WH_DEDUPE_BACKEND", d.Backend, dedupeBackends) +} + +// CoordBackend names the lease implementation singleton work (the sweeper) +// is elected through. +type CoordBackend string + +// CoordLocal holds leases in this process: correct while no other process +// shares its queue. +const CoordLocal CoordBackend = "local" + +var coordBackends = []CoordBackend{CoordLocal} + +// Coord selects the coordination layer. +type Coord struct { + Backend CoordBackend `yaml:"backend" env:"WH_COORD_BACKEND" env-default:"local"` +} + +func (c Coord) validate() error { + return checkBackend("coord.backend", "WH_COORD_BACKEND", c.Backend, coordBackends) +} + +// checkBackend refuses a backend this build has no implementation for, +// listing the ones it has. env repeats the struct tag's literal: a tag can't +// reference a constant. +func checkBackend[T ~string](key, env string, got T, valid []T) error { + if slices.Contains(valid, got) { + return nil + } + names := make([]string, len(valid)) + for i, v := range valid { + names[i] = string(v) + } + return fmt.Errorf("%s (%s) %q is not a backend this build has; valid: %s", key, env, got, strings.Join(names, ", ")) +} + +// validateBackends checks every layer's backend and its sub-block. +func (c *Config) validateBackends() error { + for _, check := range []func() error{c.MQ.validate, c.Cache.validate, c.Dedupe.validate, c.Coord.validate} { + if err := check(); err != nil { + return err + } + } + return nil +} + +// Distributed reports whether the message queue is shared with other +// processes. The embedded one listens on no port, so while it is selected +// every process is an island: nothing else can reach its queue. +func (c *Config) Distributed() bool { return c.MQ.Backend != MQEmbedded } + +// NeedsDataDir reports whether a selected backend keeps state under data_dir, +// and so whether boot must probe it (CheckDataDir). +func (c *Config) NeedsDataDir() bool { + return c.MQ.Backend == MQEmbedded || c.Dedupe.Backend == DedupePebble +} + +// Warnings returns what a valid configuration is still likely to get wrong, +// one line each, for boot to log at WARN. They are not errors because each is +// correct for a single replica, and one process cannot count its replicas. +func (c *Config) Warnings() []string { + if !c.Distributed() { + return nil + } + var out []string + if c.Cache.Backend == CacheLocal { + out = append(out, "cache.backend=local with a shared mq.backend is correct for one replica only: an event ingested on another replica never invalidates this one's cache, so its reads stay stale until the cached entry expires") + } + if c.Dedupe.Backend == DedupePebble { + out = append(out, "dedupe.backend=pebble with a shared mq.backend dedupes per replica only: an id seen by another replica is not seen by this one") + } + return out +} diff --git a/internal/config/backends_test.go b/internal/config/backends_test.go new file mode 100644 index 000000000..0e70dc7a3 --- /dev/null +++ b/internal/config/backends_test.go @@ -0,0 +1,161 @@ +package config + +import ( + "os" + "path/filepath" + "testing" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" +) + +// withDefaultBackends sets what Load's env-defaults would: a literal Config +// names no backend, and Validate refuses that. +func withDefaultBackends(c Config) *Config { + c.MQ.Backend, c.Cache.Backend = MQEmbedded, CacheLocal + c.Dedupe.Backend, c.Coord.Backend = DedupePebble, CoordLocal + return &c +} + +func defaultBackends() Config { + return *withDefaultBackends(Config{Server: Server{Port: 8080}, Settings: Settings{Dir: "./settings"}}) +} + +func TestLoad_BackendDefaults(t *testing.T) { + t.Parallel() + cfg, err := Load("nonexistent.yaml") + require.NoError(t, err) + assert.Equal(t, MQEmbedded, cfg.MQ.Backend) + assert.Equal(t, CacheLocal, cfg.Cache.Backend) + assert.Equal(t, DedupePebble, cfg.Dedupe.Backend) + assert.Equal(t, CoordLocal, cfg.Coord.Backend) + assert.False(t, cfg.Distributed()) + assert.True(t, cfg.NeedsDataDir()) + assert.Empty(t, cfg.Warnings()) +} + +func TestLoad_BackendsFromEnv(t *testing.T) { + t.Setenv("WH_MQ_BACKEND", "embedded") + t.Setenv("WH_CACHE_BACKEND", "local") + t.Setenv("WH_DEDUPE_BACKEND", "pebble") + t.Setenv("WH_COORD_BACKEND", "local") + cfg, err := Load("nonexistent.yaml") + require.NoError(t, err) + assert.Equal(t, MQEmbedded, cfg.MQ.Backend) + assert.Equal(t, CoordLocal, cfg.Coord.Backend) +} + +func TestLoad_BackendFromEnvRefusesAnUnknownValue(t *testing.T) { + t.Setenv("WH_MQ_BACKEND", "nats") + _, err := Load("nonexistent.yaml") + require.Error(t, err) + assert.Contains(t, err.Error(), `mq.backend (WH_MQ_BACKEND) "nats" is not a backend this build has; valid: embedded`) +} + +func TestLoad_BackendsFromYAML(t *testing.T) { + t.Parallel() + path := filepath.Join(t.TempDir(), "config.yaml") + require.NoError(t, os.WriteFile(path, []byte(` +mq: + backend: embedded +cache: + backend: local + l1_max_cost: 1024 +dedupe: + backend: pebble +coord: + backend: local +`), 0o600)) + cfg, err := Load(path) + require.NoError(t, err) + assert.Equal(t, MQEmbedded, cfg.MQ.Backend) + assert.Equal(t, CacheLocal, cfg.Cache.Backend) + assert.Equal(t, int64(1024), cfg.Cache.L1MaxCost) + assert.Equal(t, DedupePebble, cfg.Dedupe.Backend) + assert.Equal(t, CoordLocal, cfg.Coord.Backend) +} + +// A sub-block written before its backend exists, and a settings-directory +// key under a block both files share, are unknown keys — not read and ignored. +func TestLoad_BackendBlocksRefuseUnknownKeys(t *testing.T) { + t.Parallel() + path := filepath.Join(t.TempDir(), "config.yaml") + require.NoError(t, os.WriteFile(path, []byte(` +mq: + backend: embedded + max_bytes_gb: 5 + nats: + urls: nats://localhost:4222 +dedupe: + enabled: true +`), 0o600)) + _, err := Load(path) + require.Error(t, err) + assert.Contains(t, err.Error(), "dedupe.enabled, mq.max_bytes_gb, mq.nats") + assert.Contains(t, err.Error(), EnvSettingsDir) +} + +func TestUnboundEnv_KnowsTheBackendVariables(t *testing.T) { + t.Parallel() + assert.Empty(t, unboundEnv([]string{ + "WH_MQ_BACKEND=embedded", "WH_CACHE_BACKEND=local", + "WH_DEDUPE_BACKEND=pebble", "WH_COORD_BACKEND=local", + })) +} + +func TestValidate_UnknownBackend(t *testing.T) { + t.Parallel() + cases := []struct { + name string + set func(*Config) + want string + }{ + {"mq", func(c *Config) { c.MQ.Backend = "kafka" }, `mq.backend (WH_MQ_BACKEND) "kafka" is not a backend this build has; valid: embedded`}, + {"cache", func(c *Config) { c.Cache.Backend = "redis" }, `cache.backend (WH_CACHE_BACKEND) "redis" is not a backend this build has; valid: local`}, + {"dedupe", func(c *Config) { c.Dedupe.Backend = "dynamodb" }, `dedupe.backend (WH_DEDUPE_BACKEND) "dynamodb" is not a backend this build has; valid: pebble`}, + {"coord", func(c *Config) { c.Coord.Backend = "nats" }, `coord.backend (WH_COORD_BACKEND) "nats" is not a backend this build has; valid: local`}, + // The zero value, which a Config built without Load carries. + {"empty", func(c *Config) { c.MQ.Backend = "" }, `mq.backend (WH_MQ_BACKEND) "" is not a backend`}, + } + for _, tc := range cases { + t.Run(tc.name, func(t *testing.T) { + t.Parallel() + cfg := defaultBackends() + require.NoError(t, cfg.Validate()) + tc.set(&cfg) + err := cfg.Validate() + require.Error(t, err) + assert.Contains(t, err.Error(), tc.want) + }) + } +} + +// Every warning keys on a shared queue, which no backend offers yet, so the +// value is set directly: Warnings reads the choice, it doesn't validate it. +func TestWarnings_SharedQueue(t *testing.T) { + t.Parallel() + cfg := defaultBackends() + assert.Empty(t, cfg.Warnings()) + + cfg.MQ.Backend = "shared" + require.True(t, cfg.Distributed()) + got := cfg.Warnings() + require.Len(t, got, 2) + assert.Contains(t, got[0], "cache.backend=local") + assert.Contains(t, got[1], "dedupe.backend=pebble") + + cfg.Cache.Backend, cfg.Dedupe.Backend = "shared", "shared" + assert.Empty(t, cfg.Warnings()) +} + +func TestNeedsDataDir(t *testing.T) { + t.Parallel() + cfg := defaultBackends() + assert.True(t, cfg.NeedsDataDir()) + cfg.MQ.Backend = "shared" + assert.True(t, cfg.NeedsDataDir(), "pebble dedupe still keeps state under data_dir") + cfg.Dedupe.Backend = "shared" + assert.False(t, cfg.NeedsDataDir()) + cfg.MQ.Backend = MQEmbedded + assert.True(t, cfg.NeedsDataDir(), "the embedded mq keeps state under data_dir") +} diff --git a/internal/config/config.go b/internal/config/config.go index 68b0314b6..cc7566965 100644 --- a/internal/config/config.go +++ b/internal/config/config.go @@ -19,7 +19,10 @@ type Config struct { DataDir string `yaml:"data_dir" env:"WH_DATA_DIR" env-default:"./data"` Server Server `yaml:"server"` ClickHouse ClickHouse `yaml:"clickhouse"` + MQ MQ `yaml:"mq"` Cache Cache `yaml:"cache"` + Dedupe Dedupe `yaml:"dedupe"` + Coord Coord `yaml:"coord"` Auth Auth `yaml:"auth"` OTel OTel `yaml:"otel"` Prometheus Prometheus `yaml:"prometheus"` @@ -132,13 +135,6 @@ type ClickHouse struct { MaxTotalConns int `yaml:"max_total_conns" env:"WH_CH_MAX_TOTAL_CONNS" env-default:"0"` } -// Cache sizes the in-process L1 cache. The time-range bucket structured -// queries normalize to is a settings-directory key -// (query.timestamp_bucket_seconds) — query shaping, not process memory. -type Cache struct { - L1MaxCost int64 `yaml:"l1_max_cost" env:"WH_CACHE_L1_MAX_COST" env-default:"67108864"` -} - // Auth holds the authentication secrets. The verifier wiring — `jwks_url`, // `role_claim` — is the settings directory's `auth` block (hot-reloadable: // a change rebuilds the verifier). There is no on/off switch: the middleware always runs. A request @@ -224,7 +220,7 @@ func (c *Config) Validate() error { } } - return nil + return c.validateBackends() } // Load reads config from a YAML file (if it exists) with env var overrides. diff --git a/internal/config/config_test.go b/internal/config/config_test.go index 822d639d1..ee8b07cc5 100644 --- a/internal/config/config_test.go +++ b/internal/config/config_test.go @@ -203,7 +203,7 @@ func TestValidate_SampleRatesIgnoredWhenObservabilityDisabled(t *testing.T) { Logs: OTelLogs{SampleRate: -1}, }, } - assert.NoError(t, cfg.Validate()) + assert.NoError(t, withDefaultBackends(cfg).Validate()) } func TestValidate_SampleRatesIgnoredWhenSignalDisabled(t *testing.T) { @@ -219,7 +219,7 @@ func TestValidate_SampleRatesIgnoredWhenSignalDisabled(t *testing.T) { Logs: OTelLogs{Enabled: false, SampleRate: -1}, }, } - assert.NoError(t, cfg.Validate()) + assert.NoError(t, withDefaultBackends(cfg).Validate()) } func TestLoad_Defaults_PrometheusDisabled(t *testing.T) { @@ -336,7 +336,7 @@ func TestValidate_PrometheusV1PathAllowedOnSidecarPort(t *testing.T) { Settings: Settings{Dir: "./settings"}, Prometheus: Prometheus{Enabled: true, Path: "/v1/metrics", Port: 9091}, } - assert.NoError(t, cfg.Validate()) + assert.NoError(t, withDefaultBackends(cfg).Validate()) } func TestValidate_PrometheusOnly_NoOTel(t *testing.T) { @@ -348,7 +348,7 @@ func TestValidate_PrometheusOnly_NoOTel(t *testing.T) { Settings: Settings{Dir: "./settings"}, Prometheus: Prometheus{Enabled: true, Path: "/metrics", Port: 0}, } - assert.NoError(t, cfg.Validate()) + assert.NoError(t, withDefaultBackends(cfg).Validate()) } func TestValidate_PrometheusIgnoredWhenDisabled(t *testing.T) { @@ -365,7 +365,7 @@ func TestValidate_PrometheusIgnoredWhenDisabled(t *testing.T) { Port: 8080, }, } - assert.NoError(t, cfg.Validate()) + assert.NoError(t, withDefaultBackends(cfg).Validate()) } // TestEnvSettingsDir_MatchesStructTag pins the exported constant to the diff --git a/tests/integration/setup_test.go b/tests/integration/setup_test.go index a01064a29..ade560f7c 100644 --- a/tests/integration/setup_test.go +++ b/tests/integration/setup_test.go @@ -161,7 +161,10 @@ func setup() (int, func()) { DataDir: dataDir, Server: config.Server{ShutdownTimeout: 10}, ClickHouse: config.ClickHouse{Password: testCHPassword}, - Cache: config.Cache{L1MaxCost: 1 << 30}, // 1 GB + MQ: config.MQ{Backend: config.MQEmbedded}, + Cache: config.Cache{Backend: config.CacheLocal, L1MaxCost: 1 << 30}, // 1 GB + Dedupe: config.Dedupe{Backend: config.DedupePebble}, + Coord: config.Coord{Backend: config.CoordLocal}, Settings: config.Settings{Dir: settingsDir}, } a, err := app.New(ctx, app.Options{Config: cfg, Listener: ln}) diff --git a/tests/integration/tenants_test.go b/tests/integration/tenants_test.go index 40c5a700c..ca00f42f3 100644 --- a/tests/integration/tenants_test.go +++ b/tests/integration/tenants_test.go @@ -62,7 +62,10 @@ func TestNestedDirectory_PerTenantPoolsAndDiscovery(t *testing.T) { Server: config.Server{ShutdownTimeout: 10}, ClickHouse: config.ClickHouse{Password: testCHPassword}, Auth: config.Auth{OperatorKey: operatorKey}, - Cache: config.Cache{L1MaxCost: 1 << 20}, + MQ: config.MQ{Backend: config.MQEmbedded}, + Cache: config.Cache{Backend: config.CacheLocal, L1MaxCost: 1 << 20}, + Dedupe: config.Dedupe{Backend: config.DedupePebble}, + Coord: config.Coord{Backend: config.CoordLocal}, Settings: config.Settings{Dir: root}, } a, err := app.New(ctx, app.Options{Config: cfg, Listener: ln}) From fde17ba4578a2a03521e085a77fd2d3326a75c7d Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Thu, 24 Sep 2026 23:30:43 -0400 Subject: [PATCH 03/79] docs(config): say coord.backend is reserved; sync the boot-config lists Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- config.yaml | 2 +- docs/src/content/docs/architecture.md | 7 ++++--- docs/src/content/docs/configuration.mdx | 4 ++-- internal/app/wire.go | 2 +- internal/config/backends.go | 8 ++++---- internal/settings/settings.go | 8 +++++--- 7 files changed, 18 insertions(+), 15 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 9f2e821bc..6564775a8 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block, which until that backend lands is an unknown key and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name every backend: the zero value is not the default, and `app.New` refuses it. +- **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block, which until that backend lands is an unknown key and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name the `mq`, `cache` and `dedupe` backends: the zero value is not the default, and `app.New` refuses it. `coord.backend` is reserved: nothing reads it until the lease layer lands, and the sweeper still runs in every process. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/config.yaml b/config.yaml index 65c384eaa..5519b78cd 100644 --- a/config.yaml +++ b/config.yaml @@ -50,7 +50,7 @@ mq: dedupe: backend: pebble # Pebble under /pebble coord: - backend: local + backend: local # reserved: nothing is elected yet # In-process L1 cache size. The query time-bucket # (query.timestamp_bucket_seconds) is a settings key. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 938751ff8..425918076 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -89,7 +89,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring -- **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability, ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. +- **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability (after which each `Config.Warnings` line is logged at `WARN`), ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. - **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out @@ -115,8 +115,9 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ ### `config/` — Configuration -- **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). -- **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load`, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. +- **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — the implementation each layer runs on, resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). +- **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load` when `Config.NeedsDataDir` says a selected backend keeps state there, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. +- **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` at the end of `Validate`), which refuses a value not on the list and names the ones that are. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns the valid combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), which `app.New` logs at `WARN`. - **strict.go** — `rejectUnknownKeys`, the YAML half: re-reads the file as a generic tree and walks it against the struct's `yaml` tags, listing every key the struct doesn't declare. cleanenv itself is lenient by design, which is exactly wrong for boot config once keys have moved to the settings directory. - **persistence.go** — `WarnIfFreshDataDir` logs the startup `WARN` when `data_dir` is missing or empty (on a redeploy, the sign that the volume didn't persist); `LogStorageInitError` attaches the UID-65532 `permissionHint` to a NATS or Pebble open failure that looks like a permission denial — the same hint string `CheckDataDir` uses. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 8d311adc5..ee264bf45 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -46,7 +46,7 @@ Each layer's implementation is chosen once, at boot. Today every layer has one b | `mq.backend` | `WH_MQ_BACKEND` | `embedded` | The message queue. `embedded`: NATS JetStream inside this process, under `/nats`. It listens on no port, so no other process can reach its queue. | | `cache.backend` | `WH_CACHE_BACKEND` | `local` | The query-result cache. `local`: in this process, sized by `cache.l1_max_cost`. | | `dedupe.backend` | `WH_DEDUPE_BACKEND` | `pebble` | Where ingest dedupe keeps the event ids it has seen. `pebble`: in this process, under `/pebble`, open while any tenant has dedupe on. | -| `coord.backend` | `WH_COORD_BACKEND` | `local` | Where the leases for work only one process may do at a time are held. `local`: in this process. | +| `coord.backend` | `WH_COORD_BACKEND` | `local` | Reserved for the leases that will elect work only one process may do at a time, such as the sweeper. Nothing is elected yet: every process runs its own sweeper, and `local`, the only value, changes nothing. | Settings for one backend go in a sub-block named after it, `.`, read only when that backend is selected. A sub-block for a backend this build does not have is an unknown key and refuses boot. `mq` and `dedupe` also appear in the settings directory's `config.json`, with different keys (`mq.max_bytes_gb`, `dedupe.enabled`, …): those are per-tenant tunables and stay there, and one written in `config.yaml` refuses boot as an unknown key. @@ -217,7 +217,7 @@ dedupe: backend: pebble # in-process Pebble under /pebble coord: - backend: local + backend: local # reserved: nothing is elected yet auth: jwt_secret: change-me-in-production # jwks_url and role_claim are settings (config.json) diff --git a/internal/app/wire.go b/internal/app/wire.go index d53720253..3cfcad4b5 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -633,7 +633,7 @@ func (a *App) wireCache() error { // refuses a backend with no case, so reaching it means a Config built by hand // without one (the zero value is not the default), or a case missing here. func unreachableBackend[T ~string](key string, got T) error { - return fmt.Errorf("%s %q has no wiring: a Config built without config.Load must name every backend", key, got) + return fmt.Errorf("%s %q has no wiring: a Config built without config.Load must name the backend of every layer it wires", key, got) } // wireSweeper adds the active sweeper — purges messages that are both diff --git a/internal/config/backends.go b/internal/config/backends.go index 8328cea61..f2ab93308 100644 --- a/internal/config/backends.go +++ b/internal/config/backends.go @@ -71,12 +71,12 @@ func (d Dedupe) validate() error { return checkBackend("dedupe.backend", "WH_DEDUPE_BACKEND", d.Backend, dedupeBackends) } -// CoordBackend names the lease implementation singleton work (the sweeper) -// is elected through. +// CoordBackend names where leases for singleton work (the sweeper) are held. +// Nothing reads it yet: the lease layer (#613) wires it. type CoordBackend string -// CoordLocal holds leases in this process: correct while no other process -// shares its queue. +// CoordLocal holds leases in this process, which is enough while no other +// process shares its queue. const CoordLocal CoordBackend = "local" var coordBackends = []CoordBackend{CoordLocal} diff --git a/internal/settings/settings.go b/internal/settings/settings.go index 55ec089d3..8db981b3a 100644 --- a/internal/settings/settings.go +++ b/internal/settings/settings.go @@ -59,9 +59,11 @@ type PipesFile struct { // TenantConfig is the shape of config.json: the behavioral tunables that // migrate out of boot config. Boot config (config.yaml/env) keeps only what -// cannot change under a running process — resource sizing (`data_dir`, -// `cache.l1_max_cost`), listeners, the observability -// exporters — and the secrets (`clickhouse.password`, `auth.jwt_secret`, +// cannot change under a running process — the implementation each layer +// runs on (`mq.backend`, `cache.backend`, `dedupe.backend`, +// `coord.backend`), resource sizing (`data_dir`, `cache.l1_max_cost`, +// `clickhouse.max_total_conns`), listeners, the observability exporters — +// and the secrets (`clickhouse.password`, `auth.jwt_secret`, // `auth.operator_key`), which never belong in a tracked JSON file. Every // block and every top-level key inside it is REQUIRED: the binary carries no // compiled defaults, so the adopted snapshot is exactly what the files say. From 431084eb9a89ebd0f9242888d7c6dcbdefea3328 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Thu, 24 Sep 2026 23:31:50 -0400 Subject: [PATCH 04/79] fix(cache): snapshot versions at lookup; tenant token on every key Cache.Get/Set becomes Lookup(ctx, tenant, sha, deps) -> (Entry, Snapshot, error) and Set(ctx, Snapshot, value, ttl): a fill is filed under the versions read before its query ran, so a bump landing mid-query orphans it instead of re-homing pre-write rows (#382). Every query key folds the tenant version, so InvalidateTenant now orphans pipe results too. A Lookup naming another tenant's namespace is ErrForeignDependency. Set errors only on backend failure. Adds internal/testutil/cachetest, the backend-agnostic conformance suite LocalCache runs and the Redis backend will. Fixes #382. Part of #613. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- AGENTS.md | 5 +- CHANGELOG.md | 1 + docs/src/content/docs/architecture.md | 6 +- docs/src/content/docs/deployment.md | 2 +- internal/api/cache_tenant_test.go | 50 ++++ internal/api/pipes.go | 16 +- internal/api/structured_query.go | 11 +- internal/app/wire.go | 10 +- internal/cache/cache.go | 53 +++- internal/cache/local.go | 48 ++-- internal/cache/local_test.go | 238 +--------------- internal/cache/version_manager.go | 33 ++- internal/cache/version_manager_test.go | 64 +++-- internal/ingest/worker_test.go | 21 +- internal/testutil/cachetest/cachetest.go | 349 +++++++++++++++++++++++ 15 files changed, 574 insertions(+), 333 deletions(-) create mode 100644 internal/testutil/cachetest/cachetest.go diff --git a/AGENTS.md b/AGENTS.md index 165957219..d3e34545e 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -31,7 +31,7 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers - **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` and the `mq.max_bytes_gb` reconcile handing the MQ each served tenant's own gap window and byte budget, and `defaultSetting`/`onDefaultAdopt` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, and the same hook's `Hub.Prune` ends the open streams of a tenant no longer served), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it - **`auth/`** — JWT auth middleware: HMAC **or** JWKS verification with `alg` pinned to the active verifier, role extraction from a configurable claim path; always runs, never rejects (bad token → empty role + stashed reason). One verifier per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 9): `Authenticator` keys them by `tenant.ID` — the request store's `settings.Store.Tenant()`, through an injected `TenantSource`; `tenant.Default` on the tenant-exempt routes — built from each tenant's `auth` block by `Reconfigure`, dropped by `Prune` once the tenant stops being served (rejected or removed), released by `Close`; the secrets (`Config`) are boot-level and shared. A JWKS key set is fetched off the boot and reload paths: until one has been stored the verifier is pending and a token-bearing request gets `503` + `Retry-After` from `api.refuseUnverifiable` (`auth.ErrVerifierPending`), never a `default_role` evaluation; refresh is library-managed (Eric, 2026-09-22), response capped at 1 MiB; the operator key's admin role is the request tenant's -- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, `..
.
.` for a namespace — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version that leads every namespace key of one tenant, orphaning every cached query keyed by its tables in one step (a pipe result names no table and keeps its TTL, [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the table-keyed cache (the structured-query results) of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) +- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, `..
.
.` for a namespace — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version folded into every namespace and query key of one tenant, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read and `Set` files the fill under it, so a write landing mid-query orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) - **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability — boot is the validator, there is no dry run @@ -127,6 +127,7 @@ Tooling notes (the non-obvious bits `make help` won't tell you): - **Shared mocks in `internal/testutil/`**: Use `MockPublisher` (records `Publish` and `DeadLetter`), `MockCache`, `MockDeduplicator`, `MockSubscriber`, `MockMessage`, `MockPurger`, `MockDeadLetterStats` instead of creating ad-hoc mocks. See `testutil/mocks.go`. - **JWT helpers**: Use `testutil.MakeJWT(t, claims)` and `testutil.MakeExpiredJWT(t, claims)` for auth tests. See `testutil/jwt.go`. - **Schema helpers**: Use `testutil.NewTestSchemaRegistry(t, tables)` for schema-aware tests — it builds the registry through the real discovery path (`Refresh` against a mock ClickHouse connection), so timestamp specs are precomputed like production. +- **Cache backends**: every `cache.Cache` implementation runs `cachetest.Run` (`internal/testutil/cachetest`), the backend-agnostic conformance suite; a behavior the contract promises goes there, not in one backend's tests. - **Policy helpers**: Use `policy.Static(p)` for a fixed `policy.Source` in tests. - **Pipes helpers**: Use `pipes.Static(queries...)` for a fixed `pipes.Source` in tests. - **Response assertions**: Use `testutil.AssertJSONResponse(t, rec, status, expected)` and `testutil.AssertJSONContains(t, rec, status, substring)`. @@ -442,7 +443,7 @@ internal/query/ → Structured query AST + SQL builder internal/settings/ → Settings directory (validate, adopted snapshot + reload, watcher, embedded seed) internal/stream/ → SSE fan-out (event Hub: project once per role, Subscriber outbound queue, Bucket fan-out, keepalive Heartbeater wheel) internal/tenant/ → Tenant id (type, grammar, reserved default, request header name) -internal/testutil/ → Shared test helpers (mocks, JWT + schema helpers; logtest/ captures or silences the default logger) +internal/testutil/ → Shared test helpers (mocks, JWT + schema helpers; logtest/ captures or silences the default logger; cachetest/ is the conformance suite every cache.Cache backend runs) tests/ → Integration & E2E tests tests/integration/ → Go integration tests (//go:build integration; ClickHouse testcontainer) tests/e2e/ → E2E test stack (scripts/orchestrator boots a ClickHouse testcontainer + the wavehouse-cov binary) diff --git a/CHANGELOG.md b/CHANGELOG.md index 23c0c7152..83f2fd832 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -78,6 +78,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed +- **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/app/wire.go`, `docs/src/content/docs/{architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` and pipe execution rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL. The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 6eaf3d54e..3ba9a9eb0 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -90,7 +90,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability, ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe` builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ` hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the whole cache — structured-query and pipe results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe` builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ` hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out @@ -109,9 +109,9 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ ### `cache/` — Query Cache -- **cache.go** — `Cache` interface: `Get`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is keyed by the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — folded with the `Namespace`s the result depends on, each naming its tenant, table and scope: one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)). +- **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on, each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope: one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)). `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. -- **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: a namespace key is `..
.
.`, and a query key is folded with each dependency's namespace key and namespace version, so bumping a table (a scopeless write) or one scope — scope is reserved and empty today, so every write is the whole-table bump — orphans every dependent entry without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results, and the flat directory's single tenant simply carries the `0` prefix. `BumpTenant` (behind `InvalidateTenant`) advances the tenant version that leads every namespace key of one tenant, orphaning its every namespace, and every cached query keyed by one, in one step (a pipe result names no table and keeps its TTL) — a table no bump ever keyed included, which an enumeration of the index would miss — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. +- **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: a namespace key is `..
.
.`, and a query key is folded with the tenant's version and each dependency's namespace key and namespace version, so bumping a table (a scopeless write) or one scope — scope is reserved and empty today, so every write is the whole-table bump — orphans every dependent entry without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results, and the flat directory's single tenant simply carries the `0` prefix. `BumpTenant` (behind `InvalidateTenant`) advances the tenant version that leads every namespace key of one tenant, orphaning its every namespace and every cached query in one step (the tenant version is folded into every query key, so a pipe result with no dependencies is orphaned too) — a table no bump ever keyed included, which an enumeration of the index would miss — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. ### `config/` — Configuration diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 095b49909..79dae7d6a 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -389,7 +389,7 @@ The folder name is the tenant id, and each folder is a complete settings directo **The admin routes take the operator key only.** `/v1/ops/*` reaches every tenant, so over a nested directory no tenant's admin role opens it: the [operator key](/api#authentication) alone does, and a token carrying an admin role gets `403`. Boot a nested directory without `auth.operator_key` and no caller can reach these routes at all, which leaves `SIGHUP` as the only reload; the server warns about it at boot. `GET /v1/ops/pipes`, `GET /v1/ops/pipes/{name}`, `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the same `?tenant=`, and address tenant `0` without it; `GET /v1/ops/dlq/stats` takes it too, and reads a rejected or removed tenant's dead-letter queue like a served one's, since the queue is kept; a tenant that has none is a `404`. On the routes that take it the parameter is parsed strictly — a query string that does not parse, an empty or repeated `tenant`, or a malformed id is a `400`, never a silent read of the default tenant or, on the reload route, a reload of every tenant. The SDK sends it as the [`tenant` option](/sdk/admin#settings--whsettings). -**What a tenant's folder decides, and what tenant `0`'s does.** A request is evaluated against its own tenant's `policies.json` and `pipes.json` (ingest, structured queries, pipes), its `query.*` keys, its `cors.allowed_origins`, and its `dedupe` block: whether its records are deduplicated, by which id, against the tenant's own store, which that folder's `dedupe.enabled` opens and closes on reload exactly as [the single-tenant one](/settings-directory#deduplication) does (every tenant's store is a share of the one Pebble instance at `/pebble`, each key led by its tenant), so `wavehouse_ingest_dedupe_disabled_total` ticks only across a tenant's own reload, whatever the other tenants' switches say. A tenant's seen ids are its own: the same event id is first seen under each tenant that sends it. Its `auth` block is its own too: each tenant's folder wires that tenant's token verifier (`jwks_url`, `role_claim`), built when the folder is adopted and rebuilt when its wiring changes, so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its provider's key set. Under another tenant's header a token is treated as invalid, and the request falls back to that tenant's `default_role` like any other unverifiable token, possibly after a rate-limited key refetch (see [Authentication](/settings-directory#authentication)). Keep `X-Tenant-ID` pinned at the proxy so a token is never presented under the wrong tenant. Tenants can still accept each other's tokens: those that leave `jwks_url` empty share the boot HMAC secret when `auth.jwt_secret` is set, so a token verifies under any of them (with no secret they validate no token at all), and those whose `jwks_url` names the same key set accept each other's tokens; isolate them by provider, or scope rows by a signed claim ([row-level security](/access-control#row-level-security)). A tenant whose `jwks_url` has not been fetched yet answers `503` with `Retry-After` to its token-bearing requests alone. A tenant that stops being served — its folder rejected or removed — loses its verifier and the JWKS refresh with it, and gets a fresh one when its folder is adopted again. The HMAC secret and the operator key stay boot config, shared by every tenant; the operator key is stamped with the request tenant's `admin_role`. A tenant's `clickhouse` and `schema` blocks are its own as well: each tenant reads and writes its own ClickHouse — one native pool per distinct address, database, user, password and `tls` tuple, shared by the tenants naming it, under the process-wide [connection ceiling](/settings-directory#clickhouse) — and discovers its own tables from its own database on its own `schema.refresh_interval`. Its message queue is its own as well: its events are queued on a stream of their own, capped at its own `mq.max_bytes_gb` — at that budget its ingest answers `503` while every other tenant's keeps publishing — beside a dead-letter stream of its own at a tenth of it, and the history gap-fill replays from it is kept for its own `stream.gap_window_minutes`. Nothing checks what the tenants' budgets add up to against the disk, so size them together ([Message Queue](/settings-directory#message-queue)). An event is published on its tenant's subject (`ingest.{tenant}.{table}`), so a `GET /v1/stream` connection is authorized by its own tenant's `policies.json` and receives its own tenant's rows alone, the ingest worker inserts a row into its own tenant's ClickHouse, a failed row is parked under its own tenant's `dlq.enabled` and subject (`dlq.{tenant}.{table}`), and two tenants' tables of one name never share a batch. The query cache is one pool, but its entries are keyed by tenant: identical `POST /v1/query` and pipe requests from two tenants are two entries and two queries to ClickHouse, and a tenant is never served another's cached rows. An insert invalidates the table's cached results under every tenant on the same ClickHouse address and database as the tenant it was ingested for, whatever their user or `tls` block, since they read the same tables; a tenant on no pool — its folder rejected or removed, or no pool could be opened for it, such as by the ceiling — is out of that fan-out while it is, and has its cached `POST /v1/query` results dropped the moment it is back on one, so a repaired or restored folder never serves query rows cached before the inserts it missed, and so does a tenant whose folder moves it to another address or database, whose cached rows came from other tables; a cached pipe result is left alone by all of this — no insert invalidates one, since it names no table — and stays until its TTL expires. One setting weighs every tenant: the SSE keepalive, where the wheel runs at the shortest `stream.keepalive_interval` among the tenants being served, with that tenant's `stream.keepalive_buckets`. +**What a tenant's folder decides, and what tenant `0`'s does.** A request is evaluated against its own tenant's `policies.json` and `pipes.json` (ingest, structured queries, pipes), its `query.*` keys, its `cors.allowed_origins`, and its `dedupe` block: whether its records are deduplicated, by which id, against the tenant's own store, which that folder's `dedupe.enabled` opens and closes on reload exactly as [the single-tenant one](/settings-directory#deduplication) does (every tenant's store is a share of the one Pebble instance at `/pebble`, each key led by its tenant), so `wavehouse_ingest_dedupe_disabled_total` ticks only across a tenant's own reload, whatever the other tenants' switches say. A tenant's seen ids are its own: the same event id is first seen under each tenant that sends it. Its `auth` block is its own too: each tenant's folder wires that tenant's token verifier (`jwks_url`, `role_claim`), built when the folder is adopted and rebuilt when its wiring changes, so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its provider's key set. Under another tenant's header a token is treated as invalid, and the request falls back to that tenant's `default_role` like any other unverifiable token, possibly after a rate-limited key refetch (see [Authentication](/settings-directory#authentication)). Keep `X-Tenant-ID` pinned at the proxy so a token is never presented under the wrong tenant. Tenants can still accept each other's tokens: those that leave `jwks_url` empty share the boot HMAC secret when `auth.jwt_secret` is set, so a token verifies under any of them (with no secret they validate no token at all), and those whose `jwks_url` names the same key set accept each other's tokens; isolate them by provider, or scope rows by a signed claim ([row-level security](/access-control#row-level-security)). A tenant whose `jwks_url` has not been fetched yet answers `503` with `Retry-After` to its token-bearing requests alone. A tenant that stops being served — its folder rejected or removed — loses its verifier and the JWKS refresh with it, and gets a fresh one when its folder is adopted again. The HMAC secret and the operator key stay boot config, shared by every tenant; the operator key is stamped with the request tenant's `admin_role`. A tenant's `clickhouse` and `schema` blocks are its own as well: each tenant reads and writes its own ClickHouse — one native pool per distinct address, database, user, password and `tls` tuple, shared by the tenants naming it, under the process-wide [connection ceiling](/settings-directory#clickhouse) — and discovers its own tables from its own database on its own `schema.refresh_interval`. Its message queue is its own as well: its events are queued on a stream of their own, capped at its own `mq.max_bytes_gb` — at that budget its ingest answers `503` while every other tenant's keeps publishing — beside a dead-letter stream of its own at a tenth of it, and the history gap-fill replays from it is kept for its own `stream.gap_window_minutes`. Nothing checks what the tenants' budgets add up to against the disk, so size them together ([Message Queue](/settings-directory#message-queue)). An event is published on its tenant's subject (`ingest.{tenant}.{table}`), so a `GET /v1/stream` connection is authorized by its own tenant's `policies.json` and receives its own tenant's rows alone, the ingest worker inserts a row into its own tenant's ClickHouse, a failed row is parked under its own tenant's `dlq.enabled` and subject (`dlq.{tenant}.{table}`), and two tenants' tables of one name never share a batch. The query cache is one pool, but its entries are keyed by tenant: identical `POST /v1/query` and pipe requests from two tenants are two entries and two queries to ClickHouse, and a tenant is never served another's cached rows. An insert invalidates the table's cached results under every tenant on the same ClickHouse address and database as the tenant it was ingested for, whatever their user or `tls` block, since they read the same tables; a tenant on no pool — its folder rejected or removed, or no pool could be opened for it, such as by the ceiling — is out of that fan-out while it is, and has its cached `POST /v1/query` results dropped the moment it is back on one, so a repaired or restored folder never serves query rows cached before the inserts it missed, and so does a tenant whose folder moves it to another address or database, whose cached rows came from other tables. Both drop the tenant's cached pipe results too, but no insert invalidates one, since a pipe names no table: between those, a cached pipe result stays until its TTL expires. One setting weighs every tenant: the SSE keepalive, where the wheel runs at the shortest `stream.keepalive_interval` among the tenants being served, with that tenant's `stream.keepalive_buckets`. **What a lost tenant `0` costs.** A `0` folder that a reload rejects or removes stops tenant `0` being served like any other, and what becomes of the shared settings depends on how they are read. Tenant `0` leaves its ClickHouse pool (closed only once no served tenant names its tuple), and its schema registry and verifier are released with the folder, like any other tenant's; the `/v1/ops/*` routes, which resolve no tenant, verify against it, so a token there reads as invalid (`401`) rather than merely non-admin (`403`) until tenant `0` is served again — the operator key, which never consults a verifier, is unaffected. CORS does not stay either: the responses that read tenant `0`'s list — the tenant-exempt routes, the refusals, a preflight naming no tenant — carry no CORS headers until the folder is served again, while every other tenant's routes keep their own list. Tenant `0`'s own dedupe store closes, as any rejected or removed tenant's does, its seen ids kept for the folder that restores it. What is read per event follows the event's tenant, so tenant `0`'s events are the ones affected: with no ClickHouse to insert into, its rows fail and are parked on the DLQ whatever its switch said, and its open `GET /v1/stream` connections are ended, as any tenant's are when it stops being served — the other tenants' events are untouched. A nested directory that has never served a tenant `0` — no `0` folder, or one rejected at boot — serves every other tenant from its own ClickHouse. Outside `/v1/ops/*`, a `/v1` request that sends no `X-Tenant-ID` resolves to tenant `0`, so with no `0` folder it answers `404 unknown tenant: 0` (`503` with a rejected one) — the SDK's `/v1/health` reachability ping included. diff --git a/internal/api/cache_tenant_test.go b/internal/api/cache_tenant_test.go index 3c86f8b07..2ab2fa9ce 100644 --- a/internal/api/cache_tenant_test.go +++ b/internal/api/cache_tenant_test.go @@ -18,6 +18,7 @@ import ( "github.com/Wave-RF/WaveHouse/internal/cache" "github.com/Wave-RF/WaveHouse/internal/pipes" "github.com/Wave-RF/WaveHouse/internal/policy" + "github.com/Wave-RF/WaveHouse/internal/query" "github.com/Wave-RF/WaveHouse/internal/settings" "github.com/Wave-RF/WaveHouse/internal/stream" "github.com/Wave-RF/WaveHouse/internal/tenant" @@ -203,3 +204,52 @@ func TestCachedRoutes_SingleflightIsPerTenant(t *testing.T) { } } } + +// bumpingConn runs bump inside the first query only, as an insert that lands +// while ClickHouse is still reading would. +type bumpingConn struct { + driver.Conn + bump func() + queries atomic.Int32 +} + +func (c *bumpingConn) Query(context.Context, string, ...any) (driver.Rows, error) { + if c.queries.Add(1) == 1 { + c.bump() + } + return &chainEmptyRows{}, nil +} + +// #382: a result is filed under the versions read before its query ran, so +// a bump landing mid-query orphans the fill — the next request misses and +// reads the post-write rows — rather than serving pre-write rows until TTL. +// The structured query is bumped the way the ingest worker bumps it; a pipe, +// which names no table yet, by InvalidateTenant. +func TestCachedRoutes_BumpDuringQueryOrphansTheFill(t *testing.T) { + bumps := map[string]func(ctx context.Context, c cache.Cache) error{ + "structured query": func(ctx context.Context, c cache.Cache) error { + _, err := c.Invalidate(ctx, []cache.Namespace{{Tenant: tenant.Default, Table: query.SafeEncodeToken("clicks")}}) + return err + }, + "pipe execute": func(ctx context.Context, c cache.Cache) error { return c.InvalidateTenant(ctx, tenant.Default) }, + } + for _, route := range cachedRoutes { + t.Run(route.name, func(t *testing.T) { + l1, err := cache.NewLocal(1 << 20) + require.NoError(t, err) + t.Cleanup(func() { _ = l1.Close() }) + conn := &bumpingConn{bump: func() { require.NoError(t, bumps[route.name](t.Context(), l1)) }} + router := cachedRouter(t, testTenants(), conn, l1) + xcache := func() string { + w := serveAs(t, router, route.path, route.body, "") + require.Equal(t, http.StatusOK, w.Code, "body: %s", w.Body.String()) + l1.Wait() + return w.Header().Get("X-Cache") + } + assert.Equal(t, "MISS", xcache()) + assert.Equal(t, "MISS", xcache(), "the fill of a query a bump overtook is orphaned") + assert.Equal(t, "HIT", xcache()) + assert.Equal(t, int32(2), conn.queries.Load()) + }) + } +} diff --git a/internal/api/pipes.go b/internal/api/pipes.go index 66c2a249b..68df3b307 100644 --- a/internal/api/pipes.go +++ b/internal/api/pipes.go @@ -162,18 +162,20 @@ func (h *PipesHandler) Execute(w http.ResponseWriter, r *http.Request) { } // Cache. A pipe can read several tables, but the current pipe impl doesn't - // expose its table/scope dependencies, so we pass no deps: the result is keyed - // by the tenant and sha alone (TTL-only) and the ingest worker cannot - // version-invalidate it. The tenant on the key is what keeps one tenant's - // pipe result from answering another until then (#583 story 8). + // expose its table/scope dependencies, so we pass no deps: the result folds + // the tenant's version alone, so InvalidateTenant orphans it but no insert + // does (TTL-bound until #343). The snapshot is of the versions before the + // query runs, so a bump landing mid-query orphans the fill (#382). // TODO: once pipes expose their tables/scopes, pass them as deps here so writes // invalidate cached pipe results. cacheKey := queryCacheKey(store.Tenant(), sql, params) + var snap cache.Snapshot if h.Cache != nil { - if data, _, err := h.Cache.Get(r.Context(), cacheKey, nil); err == nil && data != nil { + var entry cache.Entry + if entry, snap, _ = h.Cache.Lookup(r.Context(), store.Tenant(), cacheKey, nil); entry.Value != nil { w.Header().Set("Content-Type", "application/json") w.Header().Set("X-Cache", "HIT") - _, _ = w.Write(data) //nolint:gosec // G705: the tenant id on the key only selects the entry; the bytes are JSON the handler marshalled from ClickHouse rows + _, _ = w.Write(entry.Value) return } } @@ -201,7 +203,7 @@ func (h *PipesHandler) Execute(w http.ResponseWriter, r *http.Request) { ttl := cache.QueryTimeToTTL(queryDuration) if h.Cache != nil { - _ = h.Cache.Set(r.Context(), cacheKey, nil, data, ttl) + _ = h.Cache.Set(r.Context(), snap, data, ttl) } return data, nil }) diff --git a/internal/api/structured_query.go b/internal/api/structured_query.go index 55fd8f357..921bc8f08 100644 --- a/internal/api/structured_query.go +++ b/internal/api/structured_query.go @@ -182,12 +182,15 @@ func (h *StructuredQueryHandler) Handle(w http.ResponseWriter, r *http.Request) // SafeEncodeToken("") is "", so this is a no-op while scope is empty. deps := []cache.Namespace{{Tenant: store.Tenant(), Table: safeTableName, Scope: query.SafeEncodeToken(scope)}} - // Try cache. + // Try cache. The snapshot is of the versions before the query runs, so a + // write landing mid-query orphans the fill (#382). + var snap cache.Snapshot if h.Cache != nil { - if data, _, err := h.Cache.Get(r.Context(), cacheKey, deps); err == nil && data != nil { + var entry cache.Entry + if entry, snap, _ = h.Cache.Lookup(r.Context(), store.Tenant(), cacheKey, deps); entry.Value != nil { w.Header().Set("Content-Type", "application/json") w.Header().Set("X-Cache", "HIT") - _, _ = w.Write(data) //nolint:gosec // G705: the tenant id on the key only selects the entry; the bytes are JSON the handler marshalled from ClickHouse rows + _, _ = w.Write(entry.Value) return } } @@ -244,7 +247,7 @@ func (h *StructuredQueryHandler) Handle(w http.ResponseWriter, r *http.Request) ttl := cache.QueryTimeToTTL(queryDuration) if h.Cache != nil { - _ = h.Cache.Set(r.Context(), cacheKey, deps, data, ttl) + _ = h.Cache.Set(r.Context(), snap, data, ttl) } return data, nil }) diff --git a/internal/app/wire.go b/internal/app/wire.go index 60cbdcd82..4c9f1da21 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -199,11 +199,11 @@ func dlqFor(tenants *settings.Registry) func(tenant.ID, string) bool { // its tables (chconn.Pools.SharingTables), the named one included. Reads are // untouched: a tenant's cached results stay its own. A tenant on no pool — // rejected, removed, or one no pool could be opened for, such as by the -// connection ceiling — is out of the fan-out, and its table-keyed cache is -// orphaned when it gets one (wireClickHouse, Cache.InvalidateTenant), so a -// folder repaired or restored inside a TTL never serves pre-insert -// structured-query rows; a pipe result names no table, so no insert -// invalidates it and it stays until its TTL expires (#343). +// connection ceiling — is out of the fan-out, and its cache is orphaned +// when it gets one (wireClickHouse, Cache.InvalidateTenant), so a folder +// repaired or restored inside a TTL never serves pre-insert rows; a pipe +// result names no table, so no insert invalidates it and between those it +// stays until its TTL expires (#343). type sharedTables struct { cache.Cache sharing func(tenant.ID) []tenant.ID diff --git a/internal/cache/cache.go b/internal/cache/cache.go index be62bbb42..c9e66a96b 100644 --- a/internal/cache/cache.go +++ b/internal/cache/cache.go @@ -2,22 +2,51 @@ package cache import ( "context" + "errors" "time" "github.com/Wave-RF/WaveHouse/internal/tenant" ) +// Entry is what a Lookup found. A nil Value is a miss. +type Entry struct { + Value []byte + TTL time.Duration // remaining +} + +// Snapshot is the dependency versions a Lookup observed. Set files a result +// under the snapshot taken before its query ran, so a bump that lands while +// the query runs orphans the fill rather than re-homing pre-write rows under +// the post-bump versions (#382). The zero Snapshot makes Set a no-op. +type Snapshot struct { + key string // the backend's key for the entry at the observed versions +} + +// ErrForeignDependency is a Lookup whose dependencies name a tenant other +// than the one it is for: a cached result is one tenant's, and so is every +// version it is filed under. +var ErrForeignDependency = errors.New("cache: dependency names another tenant") + // Cache provides versioned query-result storage with TTL support. +// +// Every entry is one tenant's and folds that tenant's version, so +// InvalidateTenant orphans all of it — a result with no dependencies (a pipe) +// included. A backend that cannot be reached is a miss on Lookup and a no-op +// on Set; the caller runs its query either way. type Cache interface { - // Get retrieves a cached query result and its remaining TTL. sha is the - // caller's key for the SQL+params, led by the tenant it was built for; deps - // are the namespaces the result depends on (one for a structured query, - // several for a pipe), each naming its tenant. Returns nil, 0, nil on miss. - Get(ctx context.Context, sha string, deps []Namespace) ([]byte, time.Duration, error) + // Lookup reads the entry for sha under tenant id at deps' current + // versions, and returns the snapshot of those versions for the Set that + // fills it on a miss. sha is the caller's key for the SQL and params; + // deps are the namespaces the result reads (one for a structured query, + // none yet for a pipe), each of tenant id — any other is + // ErrForeignDependency. An error is a miss with a zero Snapshot. + Lookup(ctx context.Context, id tenant.ID, sha string, deps []Namespace) (Entry, Snapshot, error) - // TODO: TTL should be set based on query execution time - // Set stores a query result keyed by sha + its dependency namespaces. - Set(ctx context.Context, sha string, deps []Namespace, value []byte, ttl time.Duration) error + // Set stores value under snap, the Snapshot a Lookup returned before the + // value was computed. It returns an error only when the backend failed; + // a value the cache declines to keep — too large, refused admission, a + // non-positive ttl, or a zero snap — is not an error. + Set(ctx context.Context, snap Snapshot, value []byte, ttl time.Duration) error // TODO: option to prefetch pipes when invalidated? // TODO: AST query builder needs to give us a deterministic key or bypass cache entirely @@ -29,18 +58,14 @@ type Cache interface { // another tenant keeps its versions. Returns the number of namespaces processed. Invalidate(ctx context.Context, namespaces []Namespace) (uint64, error) - // InvalidateTenant orphans every cached query of one tenant that is keyed - // by its tables in one step — every table and scope, bumped or not; a - // pipe result names no table, so neither this nor any insert - // invalidates it and it stays until its TTL expires (#343) — for a + // InvalidateTenant orphans every cached result of one tenant in one step + // — every table and scope, bumped or not, and every pipe result — for a // tenant that comes back after an absence from the invalidation fan-out // (its settings folder rejected or removed, #583 story 6), stale by every // insert it missed, or that moved to another ClickHouse address or // database, whose cached results were read from other tables. InvalidateTenant(ctx context.Context, id tenant.ID) error - // TODO: for local cache, we can just store the versions in memory, but for distributed/L2 cache, we will need to be able to either have stored procedures/pipelines etc to query them and attach them to a query, or sync them to each edge api server. - // Close releases resources. Close() error } diff --git a/internal/cache/local.go b/internal/cache/local.go index 4959242c7..1a755e058 100644 --- a/internal/cache/local.go +++ b/internal/cache/local.go @@ -15,6 +15,7 @@ import ( // tenant leading every key, so no entry is shared across tenants. type LocalCache struct { cache *ristretto.Cache[string, []byte] + maxCost int64 versionManager *VersionManager } @@ -29,31 +30,36 @@ func NewLocal(maxCost int64) (*LocalCache, error) { return nil, err } vm := NewVersionManager() - return &LocalCache{cache: cache, versionManager: vm}, nil + return &LocalCache{cache: cache, maxCost: maxCost, versionManager: vm}, nil } -// Get looks up a cached query RESULT by its sha (hash of SQL+params) and the -// namespaces it depends on. Used by BOTH structured queries (which pass one -// Namespace) and pipes (which pass several). Returns nil, 0, nil on miss. -func (l *LocalCache) Get(_ context.Context, sha string, deps []Namespace) ([]byte, time.Duration, error) { - cacheKey := l.versionManager.QueryKey(sha, deps) - - val, found := l.cache.Get(cacheKey) +// Lookup reads a cached query RESULT by its sha (hash of SQL+params) and the +// namespaces it depends on, and snapshots the key at their current versions. +// Used by BOTH structured queries (which pass one Namespace) and pipes (none +// yet). +func (l *LocalCache) Lookup(_ context.Context, id tenant.ID, sha string, deps []Namespace) (Entry, Snapshot, error) { + for _, d := range deps { + if d.Tenant != id { + return Entry{}, Snapshot{}, fmt.Errorf("%w: %q under %q", ErrForeignDependency, d.Tenant, id) + } + } + key := l.versionManager.QueryKey(id, sha, deps) + snap := Snapshot{key: key} + val, found := l.cache.Get(key) if !found { - return nil, 0, nil + return Entry{}, snap, nil } - remaining, _ := l.cache.GetTTL(cacheKey) - return val, remaining, nil + remaining, _ := l.cache.GetTTL(key) + return Entry{Value: val, TTL: remaining}, snap, nil } -// Set stores a query result under the folded key for its dependency namespaces. -// Used by both structured queries and pipes. -func (l *LocalCache) Set(_ context.Context, sha string, deps []Namespace, value []byte, ttl time.Duration) error { - cacheKey := l.versionManager.QueryKey(sha, deps) - - if ok := l.cache.SetWithTTL(cacheKey, value, int64(len(value)), ttl); !ok { - return fmt.Errorf("cache admission rejected for key %q", cacheKey) +// Set stores a query result under the key its Lookup snapshotted. Admission +// is asynchronous (see Wait), and Ristretto may still decline the value. +func (l *LocalCache) Set(_ context.Context, snap Snapshot, value []byte, ttl time.Duration) error { + if snap.key == "" || ttl <= 0 || int64(len(value)) > l.maxCost { + return nil } + l.cache.SetWithTTL(snap.key, value, int64(len(value)), ttl) return nil } @@ -78,9 +84,9 @@ func (l *LocalCache) Invalidate(_ context.Context, namespaces []Namespace) (uint return uint64(len(namespaces)), nil } -// InvalidateTenant orphans every cached query of tenant id keyed by its -// tables (a pipe result names none and keeps its TTL): one version -// bump, nothing enumerated (see VersionManager.BumpTenant). +// InvalidateTenant orphans every cached result of tenant id, pipe results +// included: one version bump, nothing enumerated (see +// VersionManager.BumpTenant). func (l *LocalCache) InvalidateTenant(_ context.Context, id tenant.ID) error { l.versionManager.BumpTenant(id) return nil diff --git a/internal/cache/local_test.go b/internal/cache/local_test.go index ed3bc14c9..2fdb3237e 100644 --- a/internal/cache/local_test.go +++ b/internal/cache/local_test.go @@ -1,241 +1,25 @@ -package cache +package cache_test import ( - "context" "testing" - "time" - "github.com/stretchr/testify/assert" "github.com/stretchr/testify/require" - "github.com/Wave-RF/WaveHouse/internal/tenant" + "github.com/Wave-RF/WaveHouse/internal/cache" + "github.com/Wave-RF/WaveHouse/internal/testutil/cachetest" ) -func TestLocalCache_GetMiss(t *testing.T) { - t.Parallel() - c, err := NewLocal(1 << 20) // 1 MB - require.NoError(t, err) - defer func() { _ = c.Close() }() - - val, ttl, err := c.Get(context.Background(), "missing", []Namespace{{Tenant: tenant.Default, Table: "table"}}) - assert.NoError(t, err) - assert.Nil(t, val) - assert.Zero(t, ttl) -} - -func TestLocalCache_SetAndGet(t *testing.T) { - t.Parallel() - c, err := NewLocal(1 << 20) - require.NoError(t, err) - defer func() { _ = c.Close() }() - - ctx := context.Background() - deps := []Namespace{{Tenant: tenant.Default, Table: "table", Scope: "scope"}} - err = c.Set(ctx, "key1", deps, []byte("hello"), 10*time.Second) - assert.NoError(t, err) - - // Ristretto uses async admission — wait briefly for it to be admitted. - c.Wait() +const localMaxCost = 1 << 20 - val, ttl, err := c.Get(ctx, "key1", deps) - assert.NoError(t, err) - assert.Equal(t, []byte("hello"), val) - assert.True(t, ttl > 0, "expected positive remaining TTL") -} - -func TestLocalCache_ExpiredKey(t *testing.T) { - t.Parallel() - c, err := NewLocal(1 << 20) +func newLocal(t *testing.T) cache.Cache { + t.Helper() + c, err := cache.NewLocal(localMaxCost) require.NoError(t, err) - defer func() { _ = c.Close() }() - - ctx := context.Background() - deps := []Namespace{{Tenant: tenant.Default, Table: "table"}} - // Set with very short TTL. - err = c.Set(ctx, "expires", deps, []byte("data"), 1*time.Millisecond) - assert.NoError(t, err) - - // Ensure async admission completes, then wait for expiry. - c.Wait() - time.Sleep(50 * time.Millisecond) - - val, _, err := c.Get(ctx, "expires", deps) - assert.NoError(t, err) - assert.Nil(t, val, "expected nil for expired key") + t.Cleanup(func() { _ = c.Close() }) + return c } -func TestLocalCache_Overwrite(t *testing.T) { +func TestLocalCache_Conformance(t *testing.T) { t.Parallel() - c, err := NewLocal(1 << 20) - require.NoError(t, err) - defer func() { _ = c.Close() }() - - ctx := context.Background() - deps := []Namespace{{Tenant: tenant.Default, Table: "table"}} - require.NoError(t, c.Set(ctx, "key", deps, []byte("v1"), 10*time.Second)) - c.Wait() - require.NoError(t, c.Set(ctx, "key", deps, []byte("v2"), 10*time.Second)) - c.Wait() - - val, _, err := c.Get(ctx, "key", deps) - assert.NoError(t, err) - assert.Equal(t, []byte("v2"), val) -} - -func TestLocalCache_ZeroTTL(t *testing.T) { - t.Parallel() - c, err := NewLocal(1 << 20) - require.NoError(t, err) - defer func() { _ = c.Close() }() - - ctx := context.Background() - deps := []Namespace{{Tenant: tenant.Default, Table: "table"}} - err = c.Set(ctx, "notimed", deps, []byte("data"), 0) - assert.NoError(t, err) - - c.Wait() - time.Sleep(10 * time.Millisecond) // arbitrary tiny sleep to see its still here after - - val, ttl, err := c.Get(ctx, "notimed", deps) - assert.NoError(t, err) - if val != nil { - assert.Equal(t, []byte("data"), val) - assert.Zero(t, ttl, "expected zero remaining TTL for key without TTL") - } -} - -func TestLocalCache_Invalidate(t *testing.T) { - t.Parallel() - c, err := NewLocal(1 << 20) - require.NoError(t, err) - defer func() { _ = c.Close() }() - - ctx := context.Background() - deps := []Namespace{{Tenant: tenant.Default, Table: "users", Scope: "org_1"}} - - // Set value - err = c.Set(ctx, "queryHash", deps, []byte("my_data"), 10*time.Second) - assert.NoError(t, err) - c.Wait() - - // Ensure readable - val, _, err := c.Get(ctx, "queryHash", deps) - assert.NoError(t, err) - assert.Equal(t, []byte("my_data"), val) - - // Invalidate the (users, org_1) namespace. - count, err := c.Invalidate(ctx, deps) - assert.NoError(t, err) - assert.Equal(t, uint64(1), count) - - // The folded key embeds the namespace version, which was just bumped, so this - // must now miss. - valAfter, ttlAfter, errAfter := c.Get(ctx, "queryHash", deps) - assert.NoError(t, errAfter) - assert.Nil(t, valAfter) - assert.Zero(t, ttlAfter) -} - -// Invalidate with an empty-scope namespace bumps the whole table, which must -// orphan that table's scoped entries too — not just the whole-table view. -func TestLocalCache_Invalidate_WholeTable(t *testing.T) { - t.Parallel() - c, err := NewLocal(1 << 20) - require.NoError(t, err) - defer func() { _ = c.Close() }() - - ctx := context.Background() - scoped := []Namespace{{Tenant: tenant.Default, Table: "events", Scope: "org_1"}} - - require.NoError(t, c.Set(ctx, "q", scoped, []byte("v1"), 10*time.Second)) - c.Wait() - val, _, err := c.Get(ctx, "q", scoped) - require.NoError(t, err) - require.Equal(t, []byte("v1"), val) - - // Whole-table invalidation (empty scope) must orphan the scoped entry. - _, err = c.Invalidate(ctx, []Namespace{{Tenant: tenant.Default, Table: "events"}}) - require.NoError(t, err) - - after, _, err := c.Get(ctx, "q", scoped) - assert.NoError(t, err) - assert.Nil(t, after, "whole-table bump must invalidate the scoped entry") -} - -// The same sha and table under two tenants are two entries: a result cached -// for one tenant never answers the other, and invalidating one tenant's table -// leaves the other's entry in place — whole-table and per-scope bumps alike. -func TestLocalCache_KeyedByTenant(t *testing.T) { - t.Parallel() - c, err := NewLocal(1 << 20) - require.NoError(t, err) - defer func() { _ = c.Close() }() - - ctx := context.Background() - acme := []Namespace{{Tenant: "acme", Table: "events", Scope: "org_1"}} - globex := []Namespace{{Tenant: "globex", Table: "events", Scope: "org_1"}} - - require.NoError(t, c.Set(ctx, "q", acme, []byte("acme rows"), 10*time.Second)) - require.NoError(t, c.Set(ctx, "q", globex, []byte("globex rows"), 10*time.Second)) - c.Wait() - - val, _, err := c.Get(ctx, "q", acme) - require.NoError(t, err) - assert.Equal(t, []byte("acme rows"), val) - val, _, err = c.Get(ctx, "q", globex) - require.NoError(t, err) - assert.Equal(t, []byte("globex rows"), val, "a tenant must never be served another tenant's entry") - - // A per-scope bump for acme orphans acme's entry only. - _, err = c.Invalidate(ctx, acme) - require.NoError(t, err) - val, _, err = c.Get(ctx, "q", acme) - require.NoError(t, err) - assert.Nil(t, val) - val, _, err = c.Get(ctx, "q", globex) - require.NoError(t, err) - assert.Equal(t, []byte("globex rows"), val, "an invalidation must not reach another tenant's entry") - - // So does a whole-table bump. - _, err = c.Invalidate(ctx, []Namespace{{Tenant: "acme", Table: "events"}}) - require.NoError(t, err) - val, _, err = c.Get(ctx, "q", globex) - require.NoError(t, err) - assert.Equal(t, []byte("globex rows"), val) -} - -// A tenant back after an absence from the invalidation fan-out has its every -// entry orphaned at once — every table, bumped before or not — and the other -// tenants keep theirs. -func TestLocalCache_InvalidateTenant(t *testing.T) { - t.Parallel() - c, err := NewLocal(1 << 20) - require.NoError(t, err) - defer func() { _ = c.Close() }() - - ctx := context.Background() - acmeEvents := []Namespace{{Tenant: "acme", Table: "events"}} - acmeOrders := []Namespace{{Tenant: "acme", Table: "orders", Scope: "org_1"}} - globex := []Namespace{{Tenant: "globex", Table: "events"}} - require.NoError(t, c.Set(ctx, "q", acmeEvents, []byte("acme events"), 10*time.Second)) - require.NoError(t, c.Set(ctx, "q", acmeOrders, []byte("acme orders"), 10*time.Second)) - require.NoError(t, c.Set(ctx, "q", globex, []byte("globex events"), 10*time.Second)) - c.Wait() - - require.NoError(t, c.InvalidateTenant(ctx, "acme")) - for name, deps := range map[string][]Namespace{"events": acmeEvents, "orders": acmeOrders} { - val, _, err := c.Get(ctx, "q", deps) - require.NoError(t, err) - assert.Nil(t, val, "acme's %s entry is orphaned", name) - } - val, _, err := c.Get(ctx, "q", globex) - require.NoError(t, err) - assert.Equal(t, []byte("globex events"), val, "another tenant's entry stays") - - // Entries cached after the bump are served: it is a generation, not a lock. - require.NoError(t, c.Set(ctx, "q", acmeEvents, []byte("acme again"), 10*time.Second)) - c.Wait() - val, _, err = c.Get(ctx, "q", acmeEvents) - require.NoError(t, err) - assert.Equal(t, []byte("acme again"), val) + cachetest.Run(t, newLocal, cachetest.Options{MaxValueBytes: localMaxCost}) } diff --git a/internal/cache/version_manager.go b/internal/cache/version_manager.go index 933fa81d6..d97bc70aa 100644 --- a/internal/cache/version_manager.go +++ b/internal/cache/version_manager.go @@ -65,18 +65,20 @@ func (vm *VersionManager) NamespaceKey(ns Namespace) string { return vm.namespaceKeyLocked(ns) } -// QueryKey builds the queries-table key for a result that depends on deps: the -// query's sha (hash of SQL+params) folded with every dependency's namespace key -// AND its namespace version, so a bump of any dependency misses the key. A -// structured query passes one Namespace; a pipe passes several. Deps are sorted -// so their order never changes the key. -func (vm *VersionManager) QueryKey(sha string, deps []Namespace) string { +// QueryKey builds the queries-table key for tenant id's result that depends +// on deps: the query's sha (hash of SQL+params) folded with the tenant's +// version and every dependency's namespace key AND its namespace version, so +// a bump of the tenant or of any dependency misses the key — a result with no +// deps (a pipe) is orphaned by BumpTenant too. A structured query passes one +// Namespace; a pipe passes several. Deps are sorted so their order never +// changes the key. +func (vm *VersionManager) QueryKey(id tenant.ID, sha string, deps []Namespace) string { segs := make([]string, len(deps)) // Lock per dependency rather than across the whole loop: each dep's table + // namespace versions are read together (consistent for that dep), but we don't - // hold the lock across all deps. A concurrent bump can land between deps, but the - // key is already a racy snapshot (versions can move between building it and using - // it), so cross-dep consistency buys nothing. Crucially, the sort/join run with + // hold the lock across all deps. A concurrent bump can land between deps; the + // caller files its fill under this key (a Snapshot), so a bump that lands + // anywhere after the read of a version orphans it. The sort/join run with // no lock held. for i, d := range deps { vm.mu.RLock() @@ -84,8 +86,11 @@ func (vm *VersionManager) QueryKey(sha string, deps []Namespace) string { segs[i] = fmt.Sprintf("%s.%d", nsKey, vm.namespaceVersions[nsKey]) vm.mu.RUnlock() } + vm.mu.RLock() + tv := vm.tenantVersions[id] + vm.mu.RUnlock() sort.Strings(segs) - return sha + "|" + strings.Join(segs, "|") + return fmt.Sprintf("%s|%s.%d|%s", sha, id, tv, strings.Join(segs, "|")) } // BumpTable advances a tenant's table version, orphaning every namespace — and @@ -98,10 +103,10 @@ func (vm *VersionManager) BumpTable(id tenant.ID, table string) { } // BumpTenant advances a tenant's version, orphaning its every namespace — -// and every cached query keyed by one — in one step (the whole-tenant -// nuke): every namespace key of the tenant carries the version, so nothing -// has to be enumerated, and a table no bump ever keyed is orphaned like the -// rest. Other tenants are untouched. +// and every cached query, whatever its deps — in one step (the whole-tenant +// nuke): every namespace and query key of the tenant carries the version, so +// nothing has to be enumerated, and a table no bump ever keyed is orphaned +// like the rest. Other tenants are untouched. func (vm *VersionManager) BumpTenant(id tenant.ID) { vm.mu.Lock() defer vm.mu.Unlock() diff --git a/internal/cache/version_manager_test.go b/internal/cache/version_manager_test.go index 36da8bdff..e1a1c330b 100644 --- a/internal/cache/version_manager_test.go +++ b/internal/cache/version_manager_test.go @@ -40,19 +40,25 @@ func TestVersionManager_QueryKey(t *testing.T) { t.Parallel() vm := NewVersionManager() - // One dependency at default versions: sha | .
.... - key := vm.QueryKey("hash123", []Namespace{{Tenant: "acme", Table: "users", Scope: "org_1"}}) - assert.Equal(t, "hash123|acme.0.users.0.org_1.0", key) + // One dependency at default versions: + // sha | . | ..
.... + key := vm.QueryKey("acme", "hash123", []Namespace{{Tenant: "acme", Table: "users", Scope: "org_1"}}) + assert.Equal(t, "hash123|acme.0|acme.0.users.0.org_1.0", key) + + // No deps (a pipe) still folds the tenant version. + assert.Equal(t, "hash123|acme.0|", vm.QueryKey("acme", "hash123", nil)) // Dependency order must not change the key (segments are sorted). deps1 := []Namespace{{Tenant: "acme", Table: "a"}, {Tenant: "acme", Table: "b"}} deps2 := []Namespace{{Tenant: "acme", Table: "b"}, {Tenant: "acme", Table: "a"}} - assert.Equal(t, vm.QueryKey("h", deps1), vm.QueryKey("h", deps2)) + assert.Equal(t, vm.QueryKey("acme", "h", deps1), vm.QueryKey("acme", "h", deps2)) - // The same sha and table under two tenants fold to two keys. + // The same sha and table under two tenants fold to two keys; so does the + // same sha with no deps. assert.NotEqual(t, - vm.QueryKey("h", []Namespace{{Tenant: "acme", Table: "users"}}), - vm.QueryKey("h", []Namespace{{Tenant: "globex", Table: "users"}})) + vm.QueryKey("acme", "h", []Namespace{{Tenant: "acme", Table: "users"}}), + vm.QueryKey("globex", "h", []Namespace{{Tenant: "globex", Table: "users"}})) + assert.NotEqual(t, vm.QueryKey("acme", "h", nil), vm.QueryKey("globex", "h", nil)) } func TestVersionManager_BumpTable(t *testing.T) { @@ -63,16 +69,16 @@ func TestVersionManager_BumpTable(t *testing.T) { orders := []Namespace{{Tenant: "acme", Table: "orders", Scope: "org_1"}} globexUsers := []Namespace{{Tenant: "globex", Table: "users", Scope: "org_1"}} - usersBefore := vm.QueryKey("h", users) - ordersBefore := vm.QueryKey("h", orders) - globexBefore := vm.QueryKey("h", globexUsers) + usersBefore := vm.QueryKey(users[0].Tenant, "h", users) + ordersBefore := vm.QueryKey(orders[0].Tenant, "h", orders) + globexBefore := vm.QueryKey(globexUsers[0].Tenant, "h", globexUsers) // Bumping a table changes the key for that tenant's table but leaves other // tables — and the same table under another tenant — alone. vm.BumpTable("acme", "users") - assert.NotEqual(t, usersBefore, vm.QueryKey("h", users)) - assert.Equal(t, ordersBefore, vm.QueryKey("h", orders)) - assert.Equal(t, globexBefore, vm.QueryKey("h", globexUsers)) + assert.NotEqual(t, usersBefore, vm.QueryKey(users[0].Tenant, "h", users)) + assert.Equal(t, ordersBefore, vm.QueryKey(orders[0].Tenant, "h", orders)) + assert.Equal(t, globexBefore, vm.QueryKey(globexUsers[0].Tenant, "h", globexUsers)) } func TestVersionManager_BumpNamespace(t *testing.T) { @@ -84,19 +90,19 @@ func TestVersionManager_BumpNamespace(t *testing.T) { otherScope := []Namespace{{Tenant: "acme", Table: "users", Scope: "org_2"}} otherTenant := []Namespace{{Tenant: "globex", Table: "users", Scope: "org_1"}} - scopedBefore := vm.QueryKey("h", scoped) - wholeBefore := vm.QueryKey("h", wholeTable) - otherBefore := vm.QueryKey("h", otherScope) - otherTenantBefore := vm.QueryKey("h", otherTenant) + scopedBefore := vm.QueryKey(scoped[0].Tenant, "h", scoped) + wholeBefore := vm.QueryKey(wholeTable[0].Tenant, "h", wholeTable) + otherBefore := vm.QueryKey(otherScope[0].Tenant, "h", otherScope) + otherTenantBefore := vm.QueryKey(otherTenant[0].Tenant, "h", otherTenant) // Bumping (acme, users, org_1) changes that scope AND the whole-table view, // but leaves every other scope — and the same scope under another tenant — // valid. vm.BumpNamespace(Namespace{Tenant: "acme", Table: "users", Scope: "org_1"}) - assert.NotEqual(t, scopedBefore, vm.QueryKey("h", scoped)) - assert.NotEqual(t, wholeBefore, vm.QueryKey("h", wholeTable)) - assert.Equal(t, otherBefore, vm.QueryKey("h", otherScope)) - assert.Equal(t, otherTenantBefore, vm.QueryKey("h", otherTenant)) + assert.NotEqual(t, scopedBefore, vm.QueryKey(scoped[0].Tenant, "h", scoped)) + assert.NotEqual(t, wholeBefore, vm.QueryKey(wholeTable[0].Tenant, "h", wholeTable)) + assert.Equal(t, otherBefore, vm.QueryKey(otherScope[0].Tenant, "h", otherScope)) + assert.Equal(t, otherTenantBefore, vm.QueryKey(otherTenant[0].Tenant, "h", otherTenant)) } // TestVersionManager_BumpTenant: a tenant's every namespace is orphaned in @@ -111,12 +117,16 @@ func TestVersionManager_BumpTenant(t *testing.T) { globexUsers := []Namespace{{Tenant: "globex", Table: "users", Scope: "org_1"}} vm.BumpTable("acme", "users") - usersBefore := vm.QueryKey("h", users) - ordersBefore := vm.QueryKey("h", orders) - globexBefore := vm.QueryKey("h", globexUsers) + usersBefore := vm.QueryKey(users[0].Tenant, "h", users) + ordersBefore := vm.QueryKey(orders[0].Tenant, "h", orders) + globexBefore := vm.QueryKey(globexUsers[0].Tenant, "h", globexUsers) + + vm.BumpTenant("acme") + assert.NotEqual(t, usersBefore, vm.QueryKey(users[0].Tenant, "h", users)) + assert.NotEqual(t, ordersBefore, vm.QueryKey(orders[0].Tenant, "h", orders), "a table no bump ever keyed is orphaned too") + assert.Equal(t, globexBefore, vm.QueryKey(globexUsers[0].Tenant, "h", globexUsers)) + pipeBefore := vm.QueryKey("acme", "h", nil) vm.BumpTenant("acme") - assert.NotEqual(t, usersBefore, vm.QueryKey("h", users)) - assert.NotEqual(t, ordersBefore, vm.QueryKey("h", orders), "a table no bump ever keyed is orphaned too") - assert.Equal(t, globexBefore, vm.QueryKey("h", globexUsers)) + assert.NotEqual(t, pipeBefore, vm.QueryKey("acme", "h", nil), "a result with no deps is orphaned too") } diff --git a/internal/ingest/worker_test.go b/internal/ingest/worker_test.go index 7fe9130e2..dc744ae6e 100644 --- a/internal/ingest/worker_test.go +++ b/internal/ingest/worker_test.go @@ -577,18 +577,23 @@ func TestInvalidate_ReachesOneTenantsEntries(t *testing.T) { ctx := context.Background() acme := []cache.Namespace{{Tenant: "acme", Table: "events", Scope: "org_1"}} globex := []cache.Namespace{{Tenant: "globex", Table: "events", Scope: "org_1"}} - require.NoError(t, l1.Set(ctx, "q", acme, []byte("acme rows"), time.Minute)) - require.NoError(t, l1.Set(ctx, "q", globex, []byte("globex rows"), time.Minute)) + get := func(id tenant.ID, deps []cache.Namespace) (cache.Entry, cache.Snapshot) { + e, snap, err := l1.Lookup(ctx, id, "q", deps) + require.NoError(t, err) + return e, snap + } + _, snap := get("acme", acme) + require.NoError(t, l1.Set(ctx, snap, []byte("acme rows"), time.Minute)) + _, snap = get("globex", globex) + require.NoError(t, l1.Set(ctx, snap, []byte("globex rows"), time.Minute)) l1.Wait() w.invalidate(ctx, "acme", "events", []parsedMsg{{scope: "org_1"}}) - val, _, err := l1.Get(ctx, "q", acme) - require.NoError(t, err) - assert.Nil(t, val, "acme's entry is orphaned by acme's insert") - val, _, err = l1.Get(ctx, "q", globex) - require.NoError(t, err) - assert.Equal(t, []byte("globex rows"), val, "globex's entry survives acme's insert") + e, _ := get("acme", acme) + assert.Nil(t, e.Value, "acme's entry is orphaned by acme's insert") + e, _ = get("globex", globex) + assert.Equal(t, []byte("globex rows"), e.Value, "globex's entry survives acme's insert") } // --------------------------------------------------------------------------- diff --git a/internal/testutil/cachetest/cachetest.go b/internal/testutil/cachetest/cachetest.go new file mode 100644 index 000000000..10a74b393 --- /dev/null +++ b/internal/testutil/cachetest/cachetest.go @@ -0,0 +1,349 @@ +// Package cachetest is the conformance suite every cache.Cache backend runs: +// what a hit, a miss and an invalidation mean, independent of where the +// entries and versions live. A backend's own tests call Run with a factory. +package cachetest + +import ( + "context" + "fmt" + "sync" + "testing" + "time" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + + "github.com/Wave-RF/WaveHouse/internal/cache" + "github.com/Wave-RF/WaveHouse/internal/tenant" +) + +// Options describes what a backend can do beyond the Cache contract. +type Options struct { + // MaxValueBytes is the largest value the backend keeps; 0 skips the + // oversize case. + MaxValueBytes int + + // NewPair returns two instances over one shared store, as two processes + // see it; nil skips the cross-instance cases. + NewPair func(t *testing.T) (a, b cache.Cache) +} + +// Run runs the suite, each case on a fresh cache from newCache. +func Run(t *testing.T, newCache func(t *testing.T) cache.Cache, opts Options) { + t.Helper() + cases := []struct { + name string + run func(t *testing.T, c cache.Cache) + }{ + {"miss", testMiss}, + {"set then hit", testSetThenHit}, + {"overwrite", testOverwrite}, + {"ttl expiry", testTTLExpiry}, + {"non-positive ttl stores nothing", testNonPositiveTTL}, + {"zero snapshot stores nothing", testZeroSnapshot}, + {"deps order does not matter", testDepsOrder}, + {"deps are part of the key", testDepsKeyed}, + {"tenant isolation", testTenantIsolation}, + {"foreign dependency refused", testForeignDependency}, + {"scope lattice", testScopeLattice}, + {"invalidate counts namespaces", testInvalidateCount}, + {"invalidate tenant orphans queries and pipes", testInvalidateTenant}, + {"bump during the query orphans the fill", testBumpDuringQuery}, + {"concurrent use", testConcurrent}, + } + if opts.MaxValueBytes > 0 { + cases = append(cases, struct { + name string + run func(t *testing.T, c cache.Cache) + }{"oversize value is not stored", func(t *testing.T, c cache.Cache) { testOversize(t, c, opts.MaxValueBytes) }}) + } + for _, tc := range cases { + t.Run(tc.name, func(t *testing.T) { + t.Parallel() + tc.run(t, newCache(t)) + }) + } + if opts.NewPair != nil { + t.Run("shared across instances", func(t *testing.T) { + t.Parallel() + a, b := opts.NewPair(t) + testShared(t, a, b) + }) + } +} + +const ttl = time.Minute + +var ( + acme tenant.ID = "acme" + globex tenant.ID = "globex" +) + +func ns(id tenant.ID, table, scope string) cache.Namespace { + return cache.Namespace{Tenant: id, Table: table, Scope: scope} +} + +// settle waits out asynchronous admission, for a backend that has any. +func settle(c cache.Cache) { + if w, ok := c.(interface{ Wait() }); ok { + w.Wait() + } +} + +func lookup(t *testing.T, c cache.Cache, id tenant.ID, sha string, deps ...cache.Namespace) (cache.Entry, cache.Snapshot) { + t.Helper() + e, snap, err := c.Lookup(context.Background(), id, sha, deps) + require.NoError(t, err) + return e, snap +} + +// fill stores value the way a handler does — Lookup, then Set under its +// snapshot — and checks it is then served. +func fill(t *testing.T, c cache.Cache, id tenant.ID, sha string, value string, deps ...cache.Namespace) { + t.Helper() + _, snap := lookup(t, c, id, sha, deps...) + require.NoError(t, c.Set(context.Background(), snap, []byte(value), ttl)) + settle(c) + requireHit(t, c, value, id, sha, deps...) +} + +func requireHit(t *testing.T, c cache.Cache, want string, id tenant.ID, sha string, deps ...cache.Namespace) { + t.Helper() + e, _ := lookup(t, c, id, sha, deps...) + require.Equal(t, want, string(e.Value), "%s %s %v", id, sha, deps) +} + +func assertMiss(t *testing.T, c cache.Cache, id tenant.ID, sha string, deps ...cache.Namespace) { + t.Helper() + e, _ := lookup(t, c, id, sha, deps...) + assert.Nil(t, e.Value, "%s %s %v: want a miss", id, sha, deps) + assert.Zero(t, e.TTL) +} + +func invalidate(t *testing.T, c cache.Cache, nss ...cache.Namespace) { + t.Helper() + _, err := c.Invalidate(context.Background(), nss) + require.NoError(t, err) +} + +func testMiss(t *testing.T, c cache.Cache) { + assertMiss(t, c, acme, "q", ns(acme, "events", "")) + assertMiss(t, c, acme, "q") +} + +func testSetThenHit(t *testing.T, c cache.Cache) { + deps := []cache.Namespace{ns(acme, "events", "org_1")} + fill(t, c, acme, "q", "rows", deps...) + e, _ := lookup(t, c, acme, "q", deps...) + assert.Positive(t, e.TTL) + assert.LessOrEqual(t, e.TTL, ttl) +} + +func testOverwrite(t *testing.T, c cache.Cache) { + deps := []cache.Namespace{ns(acme, "events", "")} + fill(t, c, acme, "q", "v1", deps...) + fill(t, c, acme, "q", "v2", deps...) +} + +func testTTLExpiry(t *testing.T, c cache.Cache) { + deps := []cache.Namespace{ns(acme, "events", "")} + _, snap := lookup(t, c, acme, "q", deps...) + require.NoError(t, c.Set(context.Background(), snap, []byte("rows"), time.Second)) + settle(c) + requireHit(t, c, "rows", acme, "q", deps...) + require.Eventually(t, func() bool { + e, _ := lookup(t, c, acme, "q", deps...) + return e.Value == nil + }, 5*time.Second, 50*time.Millisecond) +} + +func testNonPositiveTTL(t *testing.T, c cache.Cache) { + for i, d := range []time.Duration{0, -time.Second} { + sha := fmt.Sprintf("q%d", i) + _, snap := lookup(t, c, acme, sha) + require.NoError(t, c.Set(context.Background(), snap, []byte("rows"), d)) + settle(c) + assertMiss(t, c, acme, sha) + } +} + +func testZeroSnapshot(t *testing.T, c cache.Cache) { + require.NoError(t, c.Set(context.Background(), cache.Snapshot{}, []byte("rows"), ttl)) + settle(c) + assertMiss(t, c, acme, "") +} + +func testDepsOrder(t *testing.T, c cache.Cache) { + a, b := ns(acme, "events", ""), ns(acme, "orders", "org_1") + fill(t, c, acme, "q", "rows", a, b) + requireHit(t, c, "rows", acme, "q", b, a) +} + +func testDepsKeyed(t *testing.T, c cache.Cache) { + a, b := ns(acme, "events", ""), ns(acme, "orders", "") + fill(t, c, acme, "q", "rows", a) + assertMiss(t, c, acme, "q", a, b) + assertMiss(t, c, acme, "q", b) + assertMiss(t, c, acme, "q") +} + +// The same sha and table under two tenants are two entries, and a bump +// under one — scoped or whole-table — leaves the other's in place. +func testTenantIsolation(t *testing.T, c cache.Cache) { + fill(t, c, acme, "q", "acme rows", ns(acme, "events", "org_1")) + fill(t, c, globex, "q", "globex rows", ns(globex, "events", "org_1")) + fill(t, c, acme, "pipe", "acme pipe") + fill(t, c, globex, "pipe", "globex pipe") + + invalidate(t, c, ns(acme, "events", "org_1")) + assertMiss(t, c, acme, "q", ns(acme, "events", "org_1")) + requireHit(t, c, "globex rows", globex, "q", ns(globex, "events", "org_1")) + + invalidate(t, c, ns(acme, "events", "")) + requireHit(t, c, "globex rows", globex, "q", ns(globex, "events", "org_1")) + + require.NoError(t, c.InvalidateTenant(context.Background(), acme)) + assertMiss(t, c, acme, "pipe") + requireHit(t, c, "globex pipe", globex, "pipe") +} + +// Every version an entry is filed under is its own tenant's, so a Lookup +// naming another tenant's namespace is refused, and its snapshot stores +// nothing. +func testForeignDependency(t *testing.T, c cache.Cache) { + e, snap, err := c.Lookup(context.Background(), acme, "q", []cache.Namespace{ns(globex, "events", "")}) + require.ErrorIs(t, err, cache.ErrForeignDependency) + assert.Nil(t, e.Value) + require.NoError(t, c.Set(context.Background(), snap, []byte("rows"), ttl)) + settle(c) + assertMiss(t, c, acme, "q", ns(acme, "events", "")) + assertMiss(t, c, globex, "q", ns(globex, "events", "")) +} + +// A scoped bump orphans that scope and the whole-table view; a scopeless +// (whole-table) bump orphans every scope of the table; neither reaches +// another table. +func testScopeLattice(t *testing.T, c cache.Cache) { + org1, org2, whole, orders := ns(acme, "events", "org_1"), ns(acme, "events", "org_2"), ns(acme, "events", ""), ns(acme, "orders", "") + for _, d := range []cache.Namespace{org1, org2, whole, orders} { + fill(t, c, acme, "q", d.Table+"/"+d.Scope, d) + } + + invalidate(t, c, org1) + assertMiss(t, c, acme, "q", org1) + assertMiss(t, c, acme, "q", whole) + requireHit(t, c, "events/org_2", acme, "q", org2) + requireHit(t, c, "orders/", acme, "q", orders) + + invalidate(t, c, whole) + assertMiss(t, c, acme, "q", org2) + requireHit(t, c, "orders/", acme, "q", orders) +} + +func testInvalidateCount(t *testing.T, c cache.Cache) { + n, err := c.Invalidate(context.Background(), []cache.Namespace{ns(acme, "events", ""), ns(globex, "events", "org_1")}) + require.NoError(t, err) + assert.Equal(t, uint64(2), n) + n, err = c.Invalidate(context.Background(), nil) + require.NoError(t, err) + assert.Zero(t, n) +} + +// InvalidateTenant orphans every entry of the tenant — tables it never +// bumped, and results with no deps (pipes) — in one step, and entries +// filled after it are served again: a generation, not a lock. +func testInvalidateTenant(t *testing.T, c cache.Cache) { + events, orders := ns(acme, "events", ""), ns(acme, "orders", "org_1") + fill(t, c, acme, "q", "events", events) + fill(t, c, acme, "q", "orders", orders) + fill(t, c, acme, "pipe", "pipe") + fill(t, c, globex, "pipe", "globex pipe") + + require.NoError(t, c.InvalidateTenant(context.Background(), acme)) + assertMiss(t, c, acme, "q", events) + assertMiss(t, c, acme, "q", orders) + assertMiss(t, c, acme, "pipe") + requireHit(t, c, "globex pipe", globex, "pipe") + + fill(t, c, acme, "q", "events again", events) + fill(t, c, acme, "pipe", "pipe again") +} + +// #382: a fill is filed under the versions read before its query ran, so a +// bump that lands while the query runs orphans it instead of re-homing the +// pre-write rows under the post-bump versions. +func testBumpDuringQuery(t *testing.T, c cache.Cache) { + ctx := context.Background() + bumps := []struct { + name string + deps []cache.Namespace + bump func() + }{ + {"table", []cache.Namespace{ns(acme, "events", "")}, func() { invalidate(t, c, ns(acme, "events", "")) }}, + {"scope", []cache.Namespace{ns(acme, "events", "org_1")}, func() { invalidate(t, c, ns(acme, "events", "org_1")) }}, + {"tenant", nil, func() { require.NoError(t, c.InvalidateTenant(ctx, acme)) }}, + } + for _, b := range bumps { + sha := "q/" + b.name + _, snap := lookup(t, c, acme, sha, b.deps...) + b.bump() + require.NoError(t, c.Set(ctx, snap, []byte("pre-write rows"), ttl)) + settle(c) + assertMiss(t, c, acme, sha, b.deps...) + } +} + +func testOversize(t *testing.T, c cache.Cache, maxValue int) { + _, snap := lookup(t, c, acme, "big") + require.NoError(t, c.Set(context.Background(), snap, make([]byte, maxValue+1), ttl)) + settle(c) + assertMiss(t, c, acme, "big") +} + +// Lookups, fills and bumps from many goroutines at once, for -race; the +// last bump still orphans whatever was filled before it. +func testConcurrent(t *testing.T, c cache.Cache) { + ctx := context.Background() + deps := []cache.Namespace{ns(acme, "events", "")} + var wg sync.WaitGroup + for i := range 8 { + wg.Go(func() { + for j := range 50 { + _, snap, err := c.Lookup(ctx, acme, "q", deps) + assert.NoError(t, err) + assert.NoError(t, c.Set(ctx, snap, []byte("rows"), ttl)) + switch (i + j) % 10 { + case 0: + _, err = c.Invalidate(ctx, deps) + assert.NoError(t, err) + case 5: + assert.NoError(t, c.InvalidateTenant(ctx, acme)) + } + } + }) + } + wg.Wait() + settle(c) + invalidate(t, c, deps...) + assertMiss(t, c, acme, "q", deps...) +} + +// Two instances over one store are one cache: a fill on one is served by the +// other, and a bump on either orphans it for both. +func testShared(t *testing.T, a, b cache.Cache) { + deps := []cache.Namespace{ns(acme, "events", "")} + fill(t, a, acme, "q", "rows", deps...) + requireHit(t, b, "rows", acme, "q", deps...) + invalidate(t, b, deps...) + assertMiss(t, a, acme, "q", deps...) + + fill(t, a, acme, "pipe", "pipe") + require.NoError(t, b.InvalidateTenant(context.Background(), acme)) + assertMiss(t, a, acme, "pipe") + + _, snap := lookup(t, a, acme, "q", deps...) + invalidate(t, b, deps...) + require.NoError(t, a.Set(context.Background(), snap, []byte("pre-write rows"), ttl)) + settle(a) + assertMiss(t, b, acme, "q", deps...) +} From f129d5775f2ba700c4997d37e3999e5b7f062849 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Thu, 24 Sep 2026 23:36:56 -0400 Subject: [PATCH 05/79] docs(config): no backend has a sub-block yet; index backends.go in AGENTS.md Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- AGENTS.md | 2 +- CHANGELOG.md | 2 +- docs/src/content/docs/configuration.mdx | 2 +- 3 files changed, 3 insertions(+), 3 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 165957219..a7e4a28e9 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -34,7 +34,7 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, `..
.
.` for a namespace — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version that leads every namespace key of one tenant, orphaning every cached query keyed by its tables in one step (a pipe result names no table and keeps its TTL, [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the table-keyed cache (the structured-query results) of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) -- **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability — boot is the validator, there is no dry run +- **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability when a selected backend keeps state there (`NeedsDataDir`); `backends.go` holds each layer's `.backend` (only the in-process value today; `coord.backend` reserved) — boot is the validator, there is no dry run - **`dedupe/`** — `Deduplicator` interface → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) diff --git a/CHANGELOG.md b/CHANGELOG.md index 6564775a8..0aac12595 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block, which until that backend lands is an unknown key and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name the `mq`, `cache` and `dedupe` backends: the zero value is not the default, and `app.New` refuses it. `coord.backend` is reserved: nothing reads it until the lease layer lands, and the sweeper still runs in every process. +- **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block; no backend has settings yet, so every such sub-block is an unknown key for now and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name the `mq`, `cache` and `dedupe` backends: the zero value is not the default, and `app.New` refuses it. `coord.backend` is reserved: nothing reads it until the lease layer lands, and the sweeper still runs in every process. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index ee264bf45..193a6c21b 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -48,7 +48,7 @@ Each layer's implementation is chosen once, at boot. Today every layer has one b | `dedupe.backend` | `WH_DEDUPE_BACKEND` | `pebble` | Where ingest dedupe keeps the event ids it has seen. `pebble`: in this process, under `/pebble`, open while any tenant has dedupe on. | | `coord.backend` | `WH_COORD_BACKEND` | `local` | Reserved for the leases that will elect work only one process may do at a time, such as the sweeper. Nothing is elected yet: every process runs its own sweeper, and `local`, the only value, changes nothing. | -Settings for one backend go in a sub-block named after it, `.`, read only when that backend is selected. A sub-block for a backend this build does not have is an unknown key and refuses boot. `mq` and `dedupe` also appear in the settings directory's `config.json`, with different keys (`mq.max_bytes_gb`, `dedupe.enabled`, …): those are per-tenant tunables and stay there, and one written in `config.yaml` refuses boot as an unknown key. +Settings for one backend will go in a sub-block named after it, `.`, read only when that backend is selected. No backend has settings yet, so today any such sub-block, `mq.embedded` included, is an unknown key and refuses boot. `mq` and `dedupe` also appear in the settings directory's `config.json`, with different keys (`mq.max_bytes_gb`, `dedupe.enabled`, …): those are per-tenant tunables and stay there, and one written in `config.yaml` refuses boot as an unknown key. ### Server From 67df3173993f30d18cf92c00810fa1e886346ad6 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Thu, 24 Sep 2026 23:56:52 -0400 Subject: [PATCH 06/79] perf(cache): flat version index, pruned per tenant The local version index nested each table under its tenant's version and each scope under its table's, and was never pruned: every InvalidateTenant left the tenant's whole index behind. It now holds one version per tenant, (tenant, table) and (tenant, table, scope), bumped in place. A tenant bump drops the tenant's index and its next key gets a process-unique generation; a table bump drops the table's scopes. After each reload the wiring prunes the index to the tenants served. Fixes #262 for the local backend. Part of #613. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- AGENTS.md | 4 +- CHANGELOG.md | 1 + docs/src/content/docs/architecture.md | 2 +- internal/app/app_test.go | 42 +++++ internal/app/wire.go | 16 +- internal/cache/local.go | 19 +- internal/cache/local_test.go | 33 ++++ internal/cache/version_manager.go | 232 +++++++++++++++++-------- internal/cache/version_manager_test.go | 192 ++++++++++++++------ 9 files changed, 406 insertions(+), 135 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index d3e34545e..6207f60a7 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -29,9 +29,9 @@ One binary: Eighteen internal packages under `internal/` (plus `internal/testutil/` for shared test helpers): - **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers -- **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` and the `mq.max_bytes_gb` reconcile handing the MQ each served tenant's own gap window and byte budget, and `defaultSetting`/`onDefaultAdopt` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, and the same hook's `Hub.Prune` ends the open streams of a tenant no longer served), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it +- **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` and the `mq.max_bytes_gb` reconcile handing the MQ each served tenant's own gap window and byte budget, and `defaultSetting`/`onDefaultAdopt` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, the same hook's `Hub.Prune` ends the open streams of a tenant no longer served, and `LocalCache.Prune` drops its cache version index), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it - **`auth/`** — JWT auth middleware: HMAC **or** JWKS verification with `alg` pinned to the active verifier, role extraction from a configurable claim path; always runs, never rejects (bad token → empty role + stashed reason). One verifier per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 9): `Authenticator` keys them by `tenant.ID` — the request store's `settings.Store.Tenant()`, through an injected `TenantSource`; `tenant.Default` on the tenant-exempt routes — built from each tenant's `auth` block by `Reconfigure`, dropped by `Prune` once the tenant stops being served (rejected or removed), released by `Close`; the secrets (`Config`) are boot-level and shared. A JWKS key set is fetched off the boot and reload paths: until one has been stored the verifier is pending and a token-bearing request gets `503` + `Retry-After` from `api.refuseUnverifiable` (`auth.ErrVerifierPending`), never a `default_role` evaluation; refresh is library-managed (Eric, 2026-09-22), response capped at 1 MiB; the operator key's admin role is the request tenant's -- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, `..
.
.` for a namespace — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version folded into every namespace and query key of one tenant, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read and `Set` files the fill under it, so a write landing mid-query orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) +- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, a version per tenant, per (tenant, table) and per (tenant, table, scope), keyed by name and bumped in place (one entry per live namespace however often it is bumped, [#262](https://github.com/Wave-RF/WaveHouse/issues/262)) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` drops the tenant's index so its next key gets a process-unique generation, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read and `Set` files the fill under it, so a write landing mid-query orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) - **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability — boot is the validator, there is no dry run diff --git a/CHANGELOG.md b/CHANGELOG.md index 83f2fd832..50ea8f2ef 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -79,6 +79,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed - **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/app/wire.go`, `docs/src/content/docs/{architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` and pipe execution rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL. The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. +- **The cache's version index no longer grows with every bump, and forgets a tenant no longer served** (`internal/cache/{local,version_manager}.go` (+ tests), `internal/app/wire.go` (+ tests), `docs/src/content/docs/architecture.md`, `AGENTS.md`): fixes [#262](https://github.com/Wave-RF/WaveHouse/issues/262) for the in-process cache, part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The index nested each table under its tenant's version and each scope under its table's, and never pruned, so every tenant invalidation left the tenant's whole index behind, and it grew with every tenant ever served. It now holds one version per tenant, per (tenant, table) and per (tenant, table, scope), bumped in place. A tenant invalidation drops the tenant's index and hands its next key a generation unique within the process, so nothing cached before it can match again, and a table bump drops the table's scope versions. After each settings reload the index of every tenant no longer served, removed or rejected, is dropped the same way; its cached results are orphaned with it, as they already were when such a tenant came back on a pool. No change to what is cached or served. The Redis-compatible backend (#613) bounds its versions with a TTL instead. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 3ba9a9eb0..de4f6dfc1 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -111,7 +111,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on, each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope: one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)). `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. -- **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: a namespace key is `..
.
.`, and a query key is folded with the tenant's version and each dependency's namespace key and namespace version, so bumping a table (a scopeless write) or one scope — scope is reserved and empty today, so every write is the whole-table bump — orphans every dependent entry without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results, and the flat directory's single tenant simply carries the `0` prefix. `BumpTenant` (behind `InvalidateTenant`) advances the tenant version that leads every namespace key of one tenant, orphaning its every namespace and every cached query in one step (the tenant version is folded into every query key, so a pipe result with no dependencies is orphaned too) — a table no bump ever keyed included, which an enumeration of the index would miss — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. +- **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: one version per tenant, per (tenant, table) and per (tenant, table, scope), each keyed by its name alone and bumped in place, so the index holds one entry per live tenant, table and scope however often each is bumped ([#262](https://github.com/Wave-RF/WaveHouse/issues/262)). A query key folds the tenant's version and, for each dependency, its tenant's, table's and scope's, so bumping a table (a scopeless write) orphans every scope of it, and bumping one scope orphans that scope and the whole-table view — scope is reserved and empty today, so every write is the whole-table bump — all without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results. A tenant's version is a *generation*, unique within the process and handed out by the first key built for the tenant; `BumpTenant` (behind `InvalidateTenant`) drops the tenant's whole index, so the next key gets a fresh generation no cached entry folds, orphaning every cached result of the tenant in one step — a pipe result with no dependencies, and a table no bump ever keyed, included — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). `LocalCache.Prune` does the same for every tenant no longer served, which `internal/app` runs after each settings reload, so a tenant removed or rejected stops holding its index. A table bump drops the table's scope versions with it, since every key they were folded into also folds the old table version; and a bump of a tenant with no index is a no-op, since no key folds its next generation yet. The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. ### `config/` — Configuration diff --git a/internal/app/app_test.go b/internal/app/app_test.go index 2ecd77d98..1f665b6a1 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -16,6 +16,7 @@ import ( "os" "path/filepath" "strings" + "sync" "sync/atomic" "syscall" "testing" @@ -635,6 +636,47 @@ func TestReload_ReadmittedTenantCacheIsOrphaned(t *testing.T) { assert.Equal(t, []tenant.ID{"globex", "acme"}, mock.GetTenants(), "restored: the same") } +// pruneRecorder is a cache that records, at each Prune, which of the tenants +// it is asked about are still served. +type pruneRecorder struct { + testutil.MockCache + mu sync.Mutex + served []map[tenant.ID]bool +} + +func (p *pruneRecorder) Prune(served func(tenant.ID) bool) { + p.mu.Lock() + defer p.mu.Unlock() + p.served = append(p.served, map[tenant.ID]bool{"acme": served("acme"), "globex": served("globex")}) +} + +func (p *pruneRecorder) last() map[tenant.ID]bool { + p.mu.Lock() + defer p.mu.Unlock() + if len(p.served) == 0 { + return nil + } + return p.served[len(p.served)-1] +} + +// Every reload prunes the cache's version index down to the tenants served, +// so a tenant rejected or removed stops holding it (#262). +func TestReload_PrunesCacheIndexToServedTenants(t *testing.T) { + root := writeNestedSettings(t, map[string]map[string]any{"acme": nil, "globex": nil}) + a := newApp(t, testConfig(t, root), Options{}) + rec := &pruneRecorder{} + a.cache = rec + + rewriteSettings(t, filepath.Join(root, "globex"), invalidQuery) + a.tenants.Reload("test") + assert.Equal(t, map[tenant.ID]bool{"acme": true, "globex": false}, rec.last(), "rejected") + + rewriteSettings(t, filepath.Join(root, "globex"), nil) + require.NoError(t, os.RemoveAll(filepath.Join(root, "acme"))) + a.tenants.Reload("test") + assert.Equal(t, map[tenant.ID]bool{"acme": false, "globex": true}, rec.last(), "removed; the repaired one served again") +} + // keepalive is a config.json patch setting the stream block's keepalive pair. func keepalive(interval, buckets int) map[string]any { return map[string]any{"stream": map[string]any{"keepalive_interval": interval, "keepalive_buckets": buckets, "gap_window_minutes": 15}} diff --git a/internal/app/wire.go b/internal/app/wire.go index 4c9f1da21..8811ba63e 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -591,7 +591,16 @@ func (a *App) wireMQ() error { return nil } -// wireCache opens the L1 cache — the only tier in standalone mode. +// pruner is a cache whose version index lives in the process and would +// otherwise keep a tenant that stopped being served (cache.LocalCache). +type pruner interface { + Prune(served func(tenant.ID) bool) +} + +// wireCache opens the L1 cache — the only tier in standalone mode. After +// every reload a tenant no longer served, removed or rejected alike, has its +// version index dropped (#262); its cache is orphaned with it, as it would +// be anyway when it came back (wireClickHouse). func (a *App) wireCache() error { l1, err := cache.NewLocal(a.cfg.Cache.L1MaxCost) if err != nil { @@ -600,6 +609,11 @@ func (a *App) wireCache() error { // TODO: eventually this is where we can switch between ristretto, redis, tiered (both), etc a.cache = l1 a.add(component{name: "cache", close: withoutContext(l1.Close)}) + a.tenants.AfterAdopt(func([]tenant.ID) { + if p, ok := a.cache.(pruner); ok { + p.Prune(a.served) + } + }) return nil } diff --git a/internal/cache/local.go b/internal/cache/local.go index 1a755e058..3e0e14b5e 100644 --- a/internal/cache/local.go +++ b/internal/cache/local.go @@ -69,10 +69,11 @@ func (l *LocalCache) Set(_ context.Context, snap Snapshot, value []byte, ttl tim // view. Returns the number of namespaces processed. // // This bumps exactly what it's given. A whole-table bump already subsumes every -// per-scope bump for the same table (the table version is embedded in every -// namespace key), so a caller that knows a whole-table bump is coming should drop -// the now-redundant scope entries itself — the ingest worker does this as it -// builds the batch, where it already loops once and knows it's a single table. +// per-scope bump for the same table (every key that folds a scope version +// folds the table version too), so a caller that knows a whole-table bump is +// coming should drop the now-redundant scope entries itself — the ingest +// worker does this as it builds the batch, where it already loops once and +// knows it's a single table. func (l *LocalCache) Invalidate(_ context.Context, namespaces []Namespace) (uint64, error) { for _, ns := range namespaces { if ns.Scope == "" { @@ -85,13 +86,21 @@ func (l *LocalCache) Invalidate(_ context.Context, namespaces []Namespace) (uint } // InvalidateTenant orphans every cached result of tenant id, pipe results -// included: one version bump, nothing enumerated (see +// included: its version index is dropped, nothing enumerated (see // VersionManager.BumpTenant). func (l *LocalCache) InvalidateTenant(_ context.Context, id tenant.ID) error { l.versionManager.BumpTenant(id) return nil } +// Prune drops the version index of every tenant served rejects, orphaning +// its entries as InvalidateTenant would, so a tenant removed or rejected at +// a reload stops holding memory (#262). The entries themselves go with +// their TTL or Ristretto's eviction. +func (l *LocalCache) Prune(served func(tenant.ID) bool) { + l.versionManager.Prune(served) +} + // Wait blocks until all buffered writes have been applied. // Exposed for testing; production callers rarely need this. func (l *LocalCache) Wait() { diff --git a/internal/cache/local_test.go b/internal/cache/local_test.go index 2fdb3237e..ad77a40db 100644 --- a/internal/cache/local_test.go +++ b/internal/cache/local_test.go @@ -2,10 +2,13 @@ package cache_test import ( "testing" + "time" + "github.com/stretchr/testify/assert" "github.com/stretchr/testify/require" "github.com/Wave-RF/WaveHouse/internal/cache" + "github.com/Wave-RF/WaveHouse/internal/tenant" "github.com/Wave-RF/WaveHouse/internal/testutil/cachetest" ) @@ -23,3 +26,33 @@ func TestLocalCache_Conformance(t *testing.T) { t.Parallel() cachetest.Run(t, newLocal, cachetest.Options{MaxValueBytes: localMaxCost}) } + +// A tenant that stops being served has its index dropped: what it cached is +// orphaned — it misses when served again — and a tenant still served keeps +// its entries. +func TestLocalCache_Prune(t *testing.T) { + t.Parallel() + c, err := cache.NewLocal(localMaxCost) + require.NoError(t, err) + t.Cleanup(func() { _ = c.Close() }) + ctx := t.Context() + fill := func(id tenant.ID) { + _, snap, err := c.Lookup(ctx, id, "q", []cache.Namespace{{Tenant: id, Table: "events"}}) + require.NoError(t, err) + require.NoError(t, c.Set(ctx, snap, []byte("rows"), time.Minute)) + } + get := func(id tenant.ID) []byte { + e, _, err := c.Lookup(ctx, id, "q", []cache.Namespace{{Tenant: id, Table: "events"}}) + require.NoError(t, err) + return e.Value + } + fill("acme") + fill("globex") + c.Wait() + require.NotNil(t, get("acme")) + require.NotNil(t, get("globex")) + + c.Prune(func(id tenant.ID) bool { return id == "acme" }) + assert.NotNil(t, get("acme"), "still served") + assert.Nil(t, get("globex"), "pruned: orphaned, never revived") +} diff --git a/internal/cache/version_manager.go b/internal/cache/version_manager.go index d97bc70aa..dfe18b8c6 100644 --- a/internal/cache/version_manager.go +++ b/internal/cache/version_manager.go @@ -9,27 +9,44 @@ import ( "github.com/Wave-RF/WaveHouse/internal/tenant" ) -// VersionManager handles the safe tracking of table + scope versioning. -// It uses a standard map because versions must NEVER be evicted under memory pressure. -// TODO: this potentially could be bad/dangerous with a low amount of RAM available/high memory pressure AND a TON of tables/scopes per table... will need to work out eventually +// VersionManager is the invalidation index: one version per tenant, per +// (tenant, table) and per (tenant, table, scope), in maps keyed by name +// alone, never by another version (#262). A bump overwrites a version in +// place, so the index holds one entry per live tenant, table and scope +// however often each is bumped, and forgetting a tenant releases all of it. +// +// A query key folds all three versions of each dependency, which gives the +// lattice: a table bump orphans every scope, a scope bump that scope and the +// whole-table view, and a tenant bump everything of the tenant's. type VersionManager struct { mu sync.RWMutex - // tenantVersions leads every key of a tenant, so BumpTenant orphans the - // tenant's every namespace and query in one step — the ones no bump ever - // keyed included, which is what an enumeration of the maps would miss. - tenantVersions map[tenant.ID]uint64 // -> tenant_version - tableVersions map[string]uint64 // ..
-> table_version - namespaceVersions map[string]uint64 // ..
.. -> namespace_version + // tenants holds each tenant's index from the first query key built for + // it until the tenant is bumped or pruned. + tenants map[tenant.ID]*tenantVersions + + // lastGen is the last generation handed to a tenant; see tenantVersions.gen. + lastGen uint64 } -// NewVersionManager initializes the thread-safe version store. -func NewVersionManager() *VersionManager { - return &VersionManager{ - tenantVersions: make(map[tenant.ID]uint64), - tableVersions: make(map[string]uint64), - namespaceVersions: make(map[string]uint64), - } +// tenantVersions is one tenant's slice of the index. +type tenantVersions struct { + // gen is the tenant's version: unique within the process, so a tenant + // forgotten and recreated can never fold a generation an entry was + // cached under. That is what makes dropping the tenant's whole index a + // safe bump. + gen uint64 + tables map[string]*tableVersions +} + +// tableVersions is one table's version and its scopes'. A missing table or +// scope reads as 0: an entry is only ever removed together with a bump of +// the version above it (a table bump clears the scopes, a tenant bump +// drops the tables), so a 0 read after a removal never matches an entry +// cached before it. +type tableVersions struct { + version uint64 + scopes map[string]uint64 } // Namespace is one (tenant, table, scope) a cached result depends on. The @@ -41,76 +58,101 @@ type Namespace struct { Scope string } -// tableKeyLocked renders the table-versions key, -// "..
"; caller must hold vm.mu. A tenant id -// cannot contain a dot and callers encode the table dot-free, so the tokens -// can never run together. -func (vm *VersionManager) tableKeyLocked(id tenant.ID, table string) string { - return fmt.Sprintf("%s.%d.%s", id, vm.tenantVersions[id], table) -} - -// namespaceKeyLocked builds the namespace-table key; caller must hold vm.mu. -func (vm *VersionManager) namespaceKeyLocked(ns Namespace) string { - tk := vm.tableKeyLocked(ns.Tenant, ns.Table) - return fmt.Sprintf("%s.%d.%s", tk, vm.tableVersions[tk], ns.Scope) -} - -// NamespaceKey renders the namespace-table key for ns at its tenant's and -// table's current versions: -// "..
.." (scopeless -// scope is "", so e.g. ".0.
.."). -func (vm *VersionManager) NamespaceKey(ns Namespace) string { - vm.mu.RLock() - defer vm.mu.RUnlock() - return vm.namespaceKeyLocked(ns) +// NewVersionManager initializes the thread-safe version store. +func NewVersionManager() *VersionManager { + return &VersionManager{tenants: make(map[tenant.ID]*tenantVersions)} } // QueryKey builds the queries-table key for tenant id's result that depends // on deps: the query's sha (hash of SQL+params) folded with the tenant's -// version and every dependency's namespace key AND its namespace version, so -// a bump of the tenant or of any dependency misses the key — a result with no -// deps (a pipe) is orphaned by BumpTenant too. A structured query passes one -// Namespace; a pipe passes several. Deps are sorted so their order never -// changes the key. +// version and, for every dependency, its tenant's, table's and scope's +// versions, so a bump of the tenant or of any dependency misses the key — a +// result with no deps (a pipe) is orphaned by BumpTenant too. A structured +// query passes one Namespace, a pipe none yet (#343). Deps are sorted so +// their order never changes the key. Every version is read under one lock, +// so the key is one consistent snapshot. +// +// The first key built for a tenant creates its index at a fresh generation. func (vm *VersionManager) QueryKey(id tenant.ID, sha string, deps []Namespace) string { + vm.mu.RLock() + key, ok := vm.queryKeyLocked(id, sha, deps, false) + vm.mu.RUnlock() + if ok { + return key + } + vm.mu.Lock() + defer vm.mu.Unlock() + key, _ = vm.queryKeyLocked(id, sha, deps, true) + return key +} + +// queryKeyLocked renders QueryKey with vm.mu held — for writing when create +// is set, which creates the index of each tenant the key names that has +// none; otherwise such a tenant reports !ok. +func (vm *VersionManager) queryKeyLocked(id tenant.ID, sha string, deps []Namespace, create bool) (string, bool) { + index := func(id tenant.ID) (*tenantVersions, bool) { + tv := vm.tenants[id] + if tv == nil && create { + tv = vm.newTenantLocked(id) + } + return tv, tv != nil + } + own, ok := index(id) + if !ok { + return "", false + } segs := make([]string, len(deps)) - // Lock per dependency rather than across the whole loop: each dep's table + - // namespace versions are read together (consistent for that dep), but we don't - // hold the lock across all deps. A concurrent bump can land between deps; the - // caller files its fill under this key (a Snapshot), so a bump that lands - // anywhere after the read of a version orphans it. The sort/join run with - // no lock held. for i, d := range deps { - vm.mu.RLock() - nsKey := vm.namespaceKeyLocked(d) - segs[i] = fmt.Sprintf("%s.%d", nsKey, vm.namespaceVersions[nsKey]) - vm.mu.RUnlock() + tv, ok := index(d.Tenant) + if !ok { + return "", false + } + var table, scope uint64 + if t := tv.tables[d.Table]; t != nil { + table, scope = t.version, t.scopes[d.Scope] + } + segs[i] = fmt.Sprintf("%s.%d.%s.%d.%s.%d", d.Tenant, tv.gen, d.Table, table, d.Scope, scope) } - vm.mu.RLock() - tv := vm.tenantVersions[id] - vm.mu.RUnlock() sort.Strings(segs) - return fmt.Sprintf("%s|%s.%d|%s", sha, id, tv, strings.Join(segs, "|")) + return fmt.Sprintf("%s|%s.%d|%s", sha, id, own.gen, strings.Join(segs, "|")), true +} + +func (vm *VersionManager) newTenantLocked(id tenant.ID) *tenantVersions { + vm.lastGen++ + tv := &tenantVersions{gen: vm.lastGen, tables: make(map[string]*tableVersions)} + vm.tenants[id] = tv + return tv +} + +// tableLocked is the entry for a tenant's table, created at version 0, or +// nil when the tenant has no index: no key folds its current generation +// yet, so there is nothing a bump could orphan. Caller holds vm.mu for +// writing. +func (vm *VersionManager) tableLocked(id tenant.ID, table string) *tableVersions { + tv := vm.tenants[id] + if tv == nil { + return nil + } + t := tv.tables[table] + if t == nil { + t = &tableVersions{} + tv.tables[table] = t + } + return t } // BumpTable advances a tenant's table version, orphaning every namespace — and // every cached query — that depends on the table, in one step (the whole-table -// nuke). The same table under another tenant is untouched. +// nuke). The table's scope versions are dropped with it: every key they were +// folded into also folds the old table version. The same table under another +// tenant is untouched. func (vm *VersionManager) BumpTable(id tenant.ID, table string) { vm.mu.Lock() defer vm.mu.Unlock() - vm.tableVersions[vm.tableKeyLocked(id, table)]++ -} - -// BumpTenant advances a tenant's version, orphaning its every namespace — -// and every cached query, whatever its deps — in one step (the whole-tenant -// nuke): every namespace and query key of the tenant carries the version, so -// nothing has to be enumerated, and a table no bump ever keyed is orphaned -// like the rest. Other tenants are untouched. -func (vm *VersionManager) BumpTenant(id tenant.ID) { - vm.mu.Lock() - defer vm.mu.Unlock() - vm.tenantVersions[id]++ + if t := vm.tableLocked(id, table); t != nil { + t.version++ + t.scopes = nil + } } // BumpNamespace advances one (tenant, table, scope) namespace plus the table's @@ -119,8 +161,54 @@ func (vm *VersionManager) BumpTenant(id tenant.ID) { func (vm *VersionManager) BumpNamespace(ns Namespace) { vm.mu.Lock() defer vm.mu.Unlock() - vm.namespaceVersions[vm.namespaceKeyLocked(ns)]++ + t := vm.tableLocked(ns.Tenant, ns.Table) + if t == nil { + return + } + if t.scopes == nil { + t.scopes = make(map[string]uint64) + } + t.scopes[ns.Scope]++ if ns.Scope != "" { - vm.namespaceVersions[vm.namespaceKeyLocked(Namespace{Tenant: ns.Tenant, Table: ns.Table})]++ + t.scopes[""]++ + } +} + +// BumpTenant orphans every cached query of a tenant, whatever its deps, in +// one step (the whole-tenant nuke), by dropping the tenant's index: the next +// key built for it gets a fresh generation, which no cached entry folds. +// Nothing has to be enumerated, a table no bump ever keyed is orphaned like +// the rest, and the index the tenant held is released. Other tenants are +// untouched. +func (vm *VersionManager) BumpTenant(id tenant.ID) { + vm.mu.Lock() + defer vm.mu.Unlock() + delete(vm.tenants, id) +} + +// Prune drops the index of every tenant keep rejects, as BumpTenant would, +// so a tenant that stops being served stops holding memory; one served again +// starts over at a fresh generation. +func (vm *VersionManager) Prune(keep func(tenant.ID) bool) { + vm.mu.Lock() + defer vm.mu.Unlock() + for id := range vm.tenants { + if !keep(id) { + delete(vm.tenants, id) + } + } +} + +// size is the number of versions the index holds, for tests. +func (vm *VersionManager) size() int { + vm.mu.RLock() + defer vm.mu.RUnlock() + n := len(vm.tenants) + for _, tv := range vm.tenants { + n += len(tv.tables) + for _, t := range tv.tables { + n += len(t.scopes) + } } + return n } diff --git a/internal/cache/version_manager_test.go b/internal/cache/version_manager_test.go index e1a1c330b..c9278577e 100644 --- a/internal/cache/version_manager_test.go +++ b/internal/cache/version_manager_test.go @@ -8,45 +8,17 @@ import ( "github.com/Wave-RF/WaveHouse/internal/tenant" ) -func TestVersionManager_NamespaceKey(t *testing.T) { - t.Parallel() - vm := NewVersionManager() - - // The tenant leads at its default version (0), then the table at its - // default version (0); a scopeless namespace renders a trailing dot. The - // flat directory's tenant is "0". - assert.Equal(t, "acme.0.users.0.", vm.NamespaceKey(Namespace{Tenant: "acme", Table: "users"})) - assert.Equal(t, "acme.0.users.0.org_1", vm.NamespaceKey(Namespace{Tenant: "acme", Table: "users", Scope: "org_1"})) - assert.Equal(t, "0.0.users.0.", vm.NamespaceKey(Namespace{Tenant: tenant.Default, Table: "users"})) - - // The table version is embedded in every namespace key for that tenant's - // table, so a BumpTable is reflected across all its scopes at once — and - // nowhere else: the same table under another tenant keeps its version. - vm.BumpTable("acme", "users") - assert.Equal(t, "acme.0.users.1.", vm.NamespaceKey(Namespace{Tenant: "acme", Table: "users"})) - assert.Equal(t, "acme.0.users.1.org_1", vm.NamespaceKey(Namespace{Tenant: "acme", Table: "users", Scope: "org_1"})) - assert.Equal(t, "globex.0.users.0.", vm.NamespaceKey(Namespace{Tenant: "globex", Table: "users"})) - - // The tenant version leads every key of the tenant, so a BumpTenant moves - // every table of acme's — the never-bumped orders table included — to a - // fresh key space, at table version 0 again, and no other tenant's. - vm.BumpTenant("acme") - assert.Equal(t, "acme.1.users.0.", vm.NamespaceKey(Namespace{Tenant: "acme", Table: "users"})) - assert.Equal(t, "acme.1.orders.0.", vm.NamespaceKey(Namespace{Tenant: "acme", Table: "orders"})) - assert.Equal(t, "globex.0.users.0.", vm.NamespaceKey(Namespace{Tenant: "globex", Table: "users"})) -} - func TestVersionManager_QueryKey(t *testing.T) { t.Parallel() vm := NewVersionManager() - // One dependency at default versions: - // sha | . | ..
.... + // sha | . | ..
...; + // acme's index is created by its first key, at generation 1. key := vm.QueryKey("acme", "hash123", []Namespace{{Tenant: "acme", Table: "users", Scope: "org_1"}}) - assert.Equal(t, "hash123|acme.0|acme.0.users.0.org_1.0", key) + assert.Equal(t, "hash123|acme.1|acme.1.users.0.org_1.0", key) // No deps (a pipe) still folds the tenant version. - assert.Equal(t, "hash123|acme.0|", vm.QueryKey("acme", "hash123", nil)) + assert.Equal(t, "hash123|acme.1|", vm.QueryKey("acme", "hash123", nil)) // Dependency order must not change the key (segments are sorted). deps1 := []Namespace{{Tenant: "acme", Table: "a"}, {Tenant: "acme", Table: "b"}} @@ -59,6 +31,9 @@ func TestVersionManager_QueryKey(t *testing.T) { vm.QueryKey("acme", "h", []Namespace{{Tenant: "acme", Table: "users"}}), vm.QueryKey("globex", "h", []Namespace{{Tenant: "globex", Table: "users"}})) assert.NotEqual(t, vm.QueryKey("acme", "h", nil), vm.QueryKey("globex", "h", nil)) + + // Reading keys is stable: nothing but a bump moves a version. + assert.Equal(t, key, vm.QueryKey("acme", "hash123", []Namespace{{Tenant: "acme", Table: "users", Scope: "org_1"}})) } func TestVersionManager_BumpTable(t *testing.T) { @@ -69,16 +44,16 @@ func TestVersionManager_BumpTable(t *testing.T) { orders := []Namespace{{Tenant: "acme", Table: "orders", Scope: "org_1"}} globexUsers := []Namespace{{Tenant: "globex", Table: "users", Scope: "org_1"}} - usersBefore := vm.QueryKey(users[0].Tenant, "h", users) - ordersBefore := vm.QueryKey(orders[0].Tenant, "h", orders) - globexBefore := vm.QueryKey(globexUsers[0].Tenant, "h", globexUsers) + usersBefore := vm.QueryKey("acme", "h", users) + ordersBefore := vm.QueryKey("acme", "h", orders) + globexBefore := vm.QueryKey("globex", "h", globexUsers) // Bumping a table changes the key for that tenant's table but leaves other // tables — and the same table under another tenant — alone. vm.BumpTable("acme", "users") - assert.NotEqual(t, usersBefore, vm.QueryKey(users[0].Tenant, "h", users)) - assert.Equal(t, ordersBefore, vm.QueryKey(orders[0].Tenant, "h", orders)) - assert.Equal(t, globexBefore, vm.QueryKey(globexUsers[0].Tenant, "h", globexUsers)) + assert.NotEqual(t, usersBefore, vm.QueryKey("acme", "h", users)) + assert.Equal(t, ordersBefore, vm.QueryKey("acme", "h", orders)) + assert.Equal(t, globexBefore, vm.QueryKey("globex", "h", globexUsers)) } func TestVersionManager_BumpNamespace(t *testing.T) { @@ -90,24 +65,44 @@ func TestVersionManager_BumpNamespace(t *testing.T) { otherScope := []Namespace{{Tenant: "acme", Table: "users", Scope: "org_2"}} otherTenant := []Namespace{{Tenant: "globex", Table: "users", Scope: "org_1"}} - scopedBefore := vm.QueryKey(scoped[0].Tenant, "h", scoped) - wholeBefore := vm.QueryKey(wholeTable[0].Tenant, "h", wholeTable) - otherBefore := vm.QueryKey(otherScope[0].Tenant, "h", otherScope) - otherTenantBefore := vm.QueryKey(otherTenant[0].Tenant, "h", otherTenant) + scopedBefore := vm.QueryKey("acme", "h", scoped) + wholeBefore := vm.QueryKey("acme", "h", wholeTable) + otherBefore := vm.QueryKey("acme", "h", otherScope) + otherTenantBefore := vm.QueryKey("globex", "h", otherTenant) // Bumping (acme, users, org_1) changes that scope AND the whole-table view, // but leaves every other scope — and the same scope under another tenant — // valid. vm.BumpNamespace(Namespace{Tenant: "acme", Table: "users", Scope: "org_1"}) - assert.NotEqual(t, scopedBefore, vm.QueryKey(scoped[0].Tenant, "h", scoped)) - assert.NotEqual(t, wholeBefore, vm.QueryKey(wholeTable[0].Tenant, "h", wholeTable)) - assert.Equal(t, otherBefore, vm.QueryKey(otherScope[0].Tenant, "h", otherScope)) - assert.Equal(t, otherTenantBefore, vm.QueryKey(otherTenant[0].Tenant, "h", otherTenant)) + assert.NotEqual(t, scopedBefore, vm.QueryKey("acme", "h", scoped)) + assert.NotEqual(t, wholeBefore, vm.QueryKey("acme", "h", wholeTable)) + assert.Equal(t, otherBefore, vm.QueryKey("acme", "h", otherScope)) + assert.Equal(t, otherTenantBefore, vm.QueryKey("globex", "h", otherTenant)) +} + +// A table bump drops the table's scope versions, which then read as 0 again +// — safe only because every key a scope version was folded into also folds +// the table version the bump moved. Pinned so a table bump that forgot to +// advance the table version would revive the scoped entry. +func TestVersionManager_BumpTableDropsScopes(t *testing.T) { + t.Parallel() + vm := NewVersionManager() + scoped := []Namespace{{Tenant: "acme", Table: "users", Scope: "org_1"}} + + fresh := vm.QueryKey("acme", "h", scoped) + vm.BumpNamespace(scoped[0]) + bumped := vm.QueryKey("acme", "h", scoped) + vm.BumpTable("acme", "users") + after := vm.QueryKey("acme", "h", scoped) + + assert.NotEqual(t, fresh, after) + assert.NotEqual(t, bumped, after) + assert.Equal(t, 2, vm.size(), "the tenant and its table; the scopes went with the table bump") } // TestVersionManager_BumpTenant: a tenant's every namespace is orphaned in -// one step — a table that was never bumped (so has no key of its own to bump) -// included — and no other tenant's is touched. +// one step — a table that was never bumped included — and no other tenant's +// is touched. func TestVersionManager_BumpTenant(t *testing.T) { t.Parallel() vm := NewVersionManager() @@ -115,18 +110,107 @@ func TestVersionManager_BumpTenant(t *testing.T) { users := []Namespace{{Tenant: "acme", Table: "users", Scope: "org_1"}} orders := []Namespace{{Tenant: "acme", Table: "orders"}} globexUsers := []Namespace{{Tenant: "globex", Table: "users", Scope: "org_1"}} + vm.QueryKey("acme", "h", users) vm.BumpTable("acme", "users") - usersBefore := vm.QueryKey(users[0].Tenant, "h", users) - ordersBefore := vm.QueryKey(orders[0].Tenant, "h", orders) - globexBefore := vm.QueryKey(globexUsers[0].Tenant, "h", globexUsers) + usersBefore := vm.QueryKey("acme", "h", users) + ordersBefore := vm.QueryKey("acme", "h", orders) + globexBefore := vm.QueryKey("globex", "h", globexUsers) vm.BumpTenant("acme") - assert.NotEqual(t, usersBefore, vm.QueryKey(users[0].Tenant, "h", users)) - assert.NotEqual(t, ordersBefore, vm.QueryKey(orders[0].Tenant, "h", orders), "a table no bump ever keyed is orphaned too") - assert.Equal(t, globexBefore, vm.QueryKey(globexUsers[0].Tenant, "h", globexUsers)) + assert.NotEqual(t, usersBefore, vm.QueryKey("acme", "h", users)) + assert.NotEqual(t, ordersBefore, vm.QueryKey("acme", "h", orders), "a table no bump ever keyed is orphaned too") + assert.Equal(t, globexBefore, vm.QueryKey("globex", "h", globexUsers)) pipeBefore := vm.QueryKey("acme", "h", nil) vm.BumpTenant("acme") assert.NotEqual(t, pipeBefore, vm.QueryKey("acme", "h", nil), "a result with no deps is orphaned too") } + +// Dropping a tenant's index is a bump only because the index it gets back +// never repeats a generation: every key built before any of these drops must +// differ from every key built after it. A counter per tenant restarting at 0 +// fails this, reviving the first entry. +func TestVersionManager_GenerationsNeverRepeat(t *testing.T) { + t.Parallel() + vm := NewVersionManager() + deps := []Namespace{{Tenant: "acme", Table: "users"}} + seen := map[string]bool{} + for i := range 100 { + key := vm.QueryKey("acme", "h", deps) + assert.False(t, seen[key], "round %d revived %s", i, key) + seen[key] = true + if i%2 == 0 { + vm.BumpTenant("acme") + } else { + vm.Prune(func(tenant.ID) bool { return false }) + } + } +} + +// A bump of a tenant with no index is a no-op: no key folds its next +// generation yet, so nothing needs orphaning — and an insert still in flight +// for a tenant just pruned does not bring its index back. +func TestVersionManager_BumpWithoutIndex(t *testing.T) { + t.Parallel() + vm := NewVersionManager() + vm.BumpTable("acme", "users") + vm.BumpNamespace(Namespace{Tenant: "acme", Table: "users", Scope: "org_1"}) + vm.BumpTenant("acme") + assert.Zero(t, vm.size()) +} + +func TestVersionManager_Prune(t *testing.T) { + t.Parallel() + vm := NewVersionManager() + acme := []Namespace{{Tenant: "acme", Table: "users"}} + globex := []Namespace{{Tenant: "globex", Table: "users"}} + acmeBefore := vm.QueryKey("acme", "h", acme) + globexBefore := vm.QueryKey("globex", "h", globex) + vm.BumpTable("acme", "users") + vm.BumpTable("globex", "users") + acmeBumped := vm.QueryKey("acme", "h", acme) + globexBumped := vm.QueryKey("globex", "h", globex) + + vm.Prune(func(id tenant.ID) bool { return id == "globex" }) + assert.Equal(t, 2, vm.size(), "globex and its table; acme released") + assert.Equal(t, globexBumped, vm.QueryKey("globex", "h", globex), "a kept tenant is untouched") + + back := vm.QueryKey("acme", "h", acme) + assert.NotEqual(t, acmeBefore, back, "a pruned tenant never revives what it cached") + assert.NotEqual(t, acmeBumped, back) + assert.NotEqual(t, globexBefore, globexBumped) +} + +// The index holds one version per live tenant, table and scope, however often +// each is bumped (#262): the nested index this replaced kept every table and +// scope under every tenant version it had seen. +func TestVersionManager_SizeDoesNotGrowWithBumps(t *testing.T) { + t.Parallel() + vm := NewVersionManager() + touch := func() { + for _, id := range []tenant.ID{"acme", "globex"} { + for _, table := range []string{"users", "orders"} { + for _, scope := range []string{"", "org_1", "org_2"} { + vm.QueryKey(id, "h", []Namespace{{Tenant: id, Table: table, Scope: scope}}) + vm.BumpNamespace(Namespace{Tenant: id, Table: table, Scope: scope}) + } + } + } + } + touch() + settled := vm.size() + assert.Equal(t, 2+2*2+2*2*3, settled, "two tenants, two tables each, three scopes each") + + for i := range 10_000 { + switch i % 3 { + case 0: + vm.BumpTable("acme", []string{"users", "orders"}[i%2]) + case 1: + vm.BumpTenant("globex") + } + touch() + assert.LessOrEqual(t, vm.size(), settled) + } + assert.Equal(t, settled, vm.size()) +} From ed8022bf34127e5cdd751d916f7d7dde162971fc Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 00:06:14 -0400 Subject: [PATCH 07/79] docs(cache): name the cache prune hook; pin LocalCache as a pruner Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- AGENTS.md | 2 +- docs/src/content/docs/architecture.md | 2 +- internal/app/wire.go | 9 +++++++-- 3 files changed, 9 insertions(+), 4 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 6207f60a7..adbec1d81 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -29,7 +29,7 @@ One binary: Eighteen internal packages under `internal/` (plus `internal/testutil/` for shared test helpers): - **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers -- **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` and the `mq.max_bytes_gb` reconcile handing the MQ each served tenant's own gap window and byte budget, and `defaultSetting`/`onDefaultAdopt` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, the same hook's `Hub.Prune` ends the open streams of a tenant no longer served, and `LocalCache.Prune` drops its cache version index), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it +- **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` and the `mq.max_bytes_gb` reconcile handing the MQ each served tenant's own gap window and byte budget, and `defaultSetting`/`onDefaultAdopt` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, the same hook's `Hub.Prune` ends the open streams of a tenant no longer served, and `wireCache`'s hook drops, through `LocalCache.Prune`, the cache version index of a tenant no longer served ([#262](https://github.com/Wave-RF/WaveHouse/issues/262))), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it - **`auth/`** — JWT auth middleware: HMAC **or** JWKS verification with `alg` pinned to the active verifier, role extraction from a configurable claim path; always runs, never rejects (bad token → empty role + stashed reason). One verifier per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 9): `Authenticator` keys them by `tenant.ID` — the request store's `settings.Store.Tenant()`, through an injected `TenantSource`; `tenant.Default` on the tenant-exempt routes — built from each tenant's `auth` block by `Reconfigure`, dropped by `Prune` once the tenant stops being served (rejected or removed), released by `Close`; the secrets (`Config`) are boot-level and shared. A JWKS key set is fetched off the boot and reload paths: until one has been stored the verifier is pending and a token-bearing request gets `503` + `Retry-After` from `api.refuseUnverifiable` (`auth.ErrVerifierPending`), never a `default_role` evaluation; refresh is library-managed (Eric, 2026-09-22), response capped at 1 MiB; the operator key's admin role is the request tenant's - **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, a version per tenant, per (tenant, table) and per (tenant, table, scope), keyed by name and bumped in place (one entry per live namespace however often it is bumped, [#262](https://github.com/Wave-RF/WaveHouse/issues/262)) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` drops the tenant's index so its next key gets a process-unique generation, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read and `Set` files the fill under it, so a write landing mid-query orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index de4f6dfc1..368c93680 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -90,7 +90,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability, ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the whole cache — structured-query and pipe results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe` builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ` hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the whole cache — structured-query and pipe results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth, dedupe and cache hooks use too), ending the open streams of a tenant no longer served, and `wireCache`'s hook prunes the cache's version index the same way (`LocalCache.Prune`, [#262](https://github.com/Wave-RF/WaveHouse/issues/262)), so a tenant no longer served stops holding it. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe` builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ` hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out diff --git a/internal/app/wire.go b/internal/app/wire.go index 8811ba63e..900c19992 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -143,8 +143,9 @@ func gapWindows(tenants *settings.Registry) map[tenant.ID]time.Duration { } // served reports whether the registry is serving tenant id: what the -// per-tenant resources — verifiers, dedupe stores, open streams — are pruned -// by once a reload removes or rejects their tenant. +// per-tenant resources — verifiers, dedupe stores, open streams, the cache +// version index — are pruned by once a reload removes or rejects their +// tenant. func (a *App) served(id tenant.ID) bool { _, ok := a.tenants.For(id) return ok @@ -597,6 +598,10 @@ type pruner interface { Prune(served func(tenant.ID) bool) } +// The hook below asserts pruner at run time; this keeps LocalCache from +// silently dropping out of it. +var _ pruner = (*cache.LocalCache)(nil) + // wireCache opens the L1 cache — the only tier in standalone mode. After // every reload a tenant no longer served, removed or rejected alike, has its // version index dropped (#262); its cache is orphaned with it, as it would From 50b7e0c3727eb4fb76eda458c2d5dd05c3aaa12c Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 00:10:10 -0400 Subject: [PATCH 08/79] feat(cache): Redis-compatible shared cache backend RedisCache keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process. Versions are random tokens under the tenant's hash tag; a value carries the tokens it was filed under, so a lookup is one pipelined MGET+GET and a lost token can only cause misses. Failures bypass the cache behind a circuit breaker, and undelivered invalidations are retried until they land. Built and tested against Redis, Valkey, Dragonfly and a Redis Cluster node; not yet selectable by config (E4). Part of #613. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- AGENTS.md | 6 +- CHANGELOG.md | 1 + Makefile | 2 +- docs/src/content/docs/architecture.md | 8 +- docs/src/content/docs/development.md | 4 +- go.mod | 7 +- go.sum | 4 + internal/cache/breaker.go | 65 +++ internal/cache/breaker_test.go | 59 +++ internal/cache/cache.go | 3 +- internal/cache/export_test.go | 12 + internal/cache/metrics.go | 106 ++++ internal/cache/pending.go | 78 +++ internal/cache/pending_test.go | 45 ++ internal/cache/redis.go | 644 +++++++++++++++++++++++ internal/cache/redis_codec.go | 194 +++++++ internal/cache/redis_codec_test.go | 198 +++++++ internal/cache/redis_integration_test.go | 418 +++++++++++++++ internal/cache/redis_test.go | 231 ++++++++ 19 files changed, 2074 insertions(+), 11 deletions(-) create mode 100644 internal/cache/breaker.go create mode 100644 internal/cache/breaker_test.go create mode 100644 internal/cache/export_test.go create mode 100644 internal/cache/metrics.go create mode 100644 internal/cache/pending.go create mode 100644 internal/cache/pending_test.go create mode 100644 internal/cache/redis.go create mode 100644 internal/cache/redis_codec.go create mode 100644 internal/cache/redis_codec_test.go create mode 100644 internal/cache/redis_integration_test.go create mode 100644 internal/cache/redis_test.go diff --git a/AGENTS.md b/AGENTS.md index d3e34545e..9455b6fe6 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -31,7 +31,7 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers - **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` and the `mq.max_bytes_gb` reconcile handing the MQ each served tenant's own gap window and byte budget, and `defaultSetting`/`onDefaultAdopt` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, and the same hook's `Hub.Prune` ends the open streams of a tenant no longer served), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it - **`auth/`** — JWT auth middleware: HMAC **or** JWKS verification with `alg` pinned to the active verifier, role extraction from a configurable claim path; always runs, never rejects (bad token → empty role + stashed reason). One verifier per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 9): `Authenticator` keys them by `tenant.ID` — the request store's `settings.Store.Tenant()`, through an injected `TenantSource`; `tenant.Default` on the tenant-exempt routes — built from each tenant's `auth` block by `Reconfigure`, dropped by `Prune` once the tenant stops being served (rejected or removed), released by `Close`; the secrets (`Config`) are boot-level and shared. A JWKS key set is fetched off the boot and reload paths: until one has been stored the verifier is pending and a token-bearing request gets `503` + `Retry-After` from `api.refuseUnverifiable` (`auth.ErrVerifierPending`), never a `default_role` evaluation; refresh is library-managed (Eric, 2026-09-22), response capped at 1 MiB; the operator key's admin role is the request tenant's -- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, `..
.
.` for a namespace — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version folded into every namespace and query key of one tenant, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read and `Set` files the fill under it, so a write landing mid-query orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) +- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index), and `RedisCache`, the Redis-compatible shared backend (random version tokens under the tenant's hash tag, one-round-trip lookups, bypass on failure behind a circuit breaker, deferred invalidations retried; built and tested, not yet selectable by config — [#613](https://github.com/Wave-RF/WaveHouse/issues/613) E4). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, `..
.
.` for a namespace — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version folded into every namespace and query key of one tenant, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read and `Set` files the fill under it, so a write landing mid-query orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) - **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability — boot is the validator, there is no dry run @@ -127,7 +127,7 @@ Tooling notes (the non-obvious bits `make help` won't tell you): - **Shared mocks in `internal/testutil/`**: Use `MockPublisher` (records `Publish` and `DeadLetter`), `MockCache`, `MockDeduplicator`, `MockSubscriber`, `MockMessage`, `MockPurger`, `MockDeadLetterStats` instead of creating ad-hoc mocks. See `testutil/mocks.go`. - **JWT helpers**: Use `testutil.MakeJWT(t, claims)` and `testutil.MakeExpiredJWT(t, claims)` for auth tests. See `testutil/jwt.go`. - **Schema helpers**: Use `testutil.NewTestSchemaRegistry(t, tables)` for schema-aware tests — it builds the registry through the real discovery path (`Refresh` against a mock ClickHouse connection), so timestamp specs are precomputed like production. -- **Cache backends**: every `cache.Cache` implementation runs `cachetest.Run` (`internal/testutil/cachetest`), the backend-agnostic conformance suite; a behavior the contract promises goes there, not in one backend's tests. +- **Cache backends**: every `cache.Cache` implementation runs `cachetest.Run` (`internal/testutil/cachetest`), the backend-agnostic conformance suite; a behavior the contract promises goes there, not in one backend's tests. `RedisCache` runs it from `internal/cache/redis_integration_test.go` (`//go:build integration`, pinned Redis, Valkey, Dragonfly and Redis Cluster containers), which `make test-integration` includes. - **Policy helpers**: Use `policy.Static(p)` for a fixed `policy.Source` in tests. - **Pipes helpers**: Use `pipes.Static(queries...)` for a fixed `pipes.Source` in tests. - **Response assertions**: Use `testutil.AssertJSONResponse(t, rec, status, expected)` and `testutil.AssertJSONContains(t, rec, status, substring)`. @@ -428,7 +428,7 @@ cmd/ → Binary entry points (thin — argv, logger, config, internal/api/ → HTTP layer (handlers, router, middleware, schema/DLQ/pipes endpoints) internal/app/ → Process wiring (build every component, run them under one errgroup, release in reverse) internal/auth/ → JWT/JWKS authentication middleware (HMAC or JWKS, role extraction from claims) -internal/cache/ → Query cache (interface, Ristretto L1, tenant-led version index) +internal/cache/ → Query cache (interface, Ristretto L1, tenant-led version index, Redis-compatible shared backend) internal/chconn/ → ClickHouse pools, one per connection tuple among the served tenants (reconciled on settings reload) internal/chsql/ → Shared ClickHouse SQL helpers (identifier quoting + bind-safety) internal/config/ → Configuration structs + loader diff --git a/CHANGELOG.md b/CHANGELOG.md index 83f2fd832..6a39bc8df 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,6 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added +- **A Redis-compatible shared cache backend, built but not yet selectable** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values over 1 KiB are zstd-compressed and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row open a circuit breaker that skips the server until a probe succeeds, and deferred invalidations are retried until they land, collapsing to one tenant-wide bump per tenant past 100,000 keys. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending`, `wavehouse_cache_value_bytes`, `wavehouse_cache_oversize_total`, `wavehouse_cache_set_failures_total{backend,reason}`. Nothing selects it yet: the `cache.backend` switch and the `cache.redis.*` settings arrive with the wiring (E4), so every deployment still runs `LocalCache`. Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression and server-stops-answering cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/Makefile b/Makefile index 8dcbe35d0..5fff09f2b 100644 --- a/Makefile +++ b/Makefile @@ -764,7 +764,7 @@ test-integration: go-mod-download ## Run Go integration tests + render coverage @rm -rf $(COV_INT)/data && mkdir -p $(COV_INT)/data @GOCOVERDIR="$(CURDIR)/$(COV_INT)/data" go tool gotestsum --format $(GOTESTSUM_FMT) -- \ -tags="integration $(TAGS)" -timeout 240s -coverpkg=./... -race -count=1 \ - ./tests/integration/... $(ARGS) \ + ./tests/integration/... ./internal/cache/... $(ARGS) \ -args -test.gocoverdir="$(CURDIR)/$(COV_INT)/data" @if [ -z "$(COV_DEFER)" ]; then go run ./scripts/cov render integration; fi diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 3ba9a9eb0..53850f49f 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -54,7 +54,7 @@ internal/ ├── api/ HTTP layer (Chi router, handlers, middleware, the cached read paths' singleflight) ├── app/ Process wiring: build every component, run them under one errgroup, release in reverse ├── auth/ JWT/JWKS authentication middleware (HMAC or JWKS, role extraction) -├── cache/ Query cache: Ristretto L1 + the tenant-led version index +├── cache/ Query cache: Ristretto L1 + the tenant-led version index; the Redis-compatible shared backend ├── chconn/ One ClickHouse pool per connection tuple among the served tenants, reconciled on reload under the ceiling ├── chsql/ Shared ClickHouse SQL helpers (identifier quoting, bind-safety) ├── config/ YAML + env var configuration loading @@ -112,6 +112,11 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on, each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope: one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)). `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. - **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: a namespace key is `..
.
.`, and a query key is folded with the tenant's version and each dependency's namespace key and namespace version, so bumping a table (a scopeless write) or one scope — scope is reserved and empty today, so every write is the whole-table bump — orphans every dependent entry without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results, and the flat directory's single tenant simply carries the `0` prefix. `BumpTenant` (behind `InvalidateTenant`) advances the tenant version that leads every namespace key of one tenant, orphaning its every namespace and every cached query in one step (the tenant version is folded into every query key, so a pipe result with no dependencies is orphaned too) — a table no bump ever keyed included, which an enumeration of the index would miss — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. +- **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET` and `MGET`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, so a token lost to eviction, expiry or a restart is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing. Nothing selects this backend yet — the boot config and wiring arrive with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s E4. +- **redis_codec.go** — the key schema (`tokenKeys`, `bumpKeys`, `valueKey`) and the value frame: format, flags, expiry, the token list, the payload, zstd-compressed from `CompressMinBytes` when that is smaller. A stored value is capped at `MaxValueBytes` and a decoded one at eight times that, which refuses a zip bomb planted in a shared server; an unknown format, as a newer process writes during a rolling upgrade, is a miss. +- **breaker.go** — the circuit breaker: `BreakerThreshold` (5) consecutive failures open it, and while open every operation skips the server at once; after `BreakerOpenFor` (5 s) one background `PING` decides whether it closes. A reply from the server, error replies included, counts as a success; a caller that gave up first counts as nothing. +- **pending.go** — invalidations the server did not take, kept per token key (repeats coalesce) and retried by a background loop with backoff from 100 ms to 10 s until they land. Landing late is still correct: a fresh token orphans the pre-write entries and any fill made meanwhile. Past `PendingMax` (100,000) keys the set collapses to one tenant bump per affected tenant. Bumps still owed when a process stops are lost after one last attempt, and the entries they would orphan are served until their TTL — the same failure as a worker stopping between an insert and its invalidation. +- **metrics.go** — the shared backend's instruments (meter `wavehouse-cache`, `backend="redis"`, no tenant attribute): `wavehouse_cache_lookups_total` by `result` (`hit`, `miss`, `stale` — filed under since-bumped tokens, `bypass` — server skipped, `error`), `wavehouse_cache_op_duration_seconds` by `op` (`lookup`, `set`, `invalidate`), `wavehouse_cache_breaker_open` (1 while bypassed, including before the first connection), `wavehouse_cache_invalidations_total` by `result` (`ok`, `deferred`), `wavehouse_cache_invalidations_pending` (bumps owed: entries they would orphan may be served stale meanwhile), `wavehouse_cache_value_bytes` (stored size), `wavehouse_cache_oversize_total` and `wavehouse_cache_set_failures_total` by `reason` (`oom`, `timeout`, `other`). ### `config/` — Configuration @@ -360,6 +365,7 @@ Client GET /v1/stream | Analytics DB | ClickHouse | Primary data store + schema source of truth | | Message Queue | NATS + JetStream | Durable event streaming | | L1 Cache | Ristretto v2 | In-process memory cache | +| Shared cache | [rueidis](https://github.com/redis/rueidis) | Redis-compatible client for the shared backend (not yet selectable) | | Embedded KV | Pebble | Optional deduplication | | Config | cleanenv | YAML + env var config loading | | Release | GoReleaser | Cross-platform binary builds | diff --git a/docs/src/content/docs/development.md b/docs/src/content/docs/development.md index 01f82e735..8e6a85036 100644 --- a/docs/src/content/docs/development.md +++ b/docs/src/content/docs/development.md @@ -341,7 +341,7 @@ Each test target writes `covdata` to `tmp/coverage//data/`, renders a tex | -------- | -------- | ------- | ------- | | Unit tests | `internal/*/_test.go` | No | `make test` | | SDK unit tests | `clients/ts/src/**/*.test.ts` | No | `make test-ts` (always includes coverage + gate) | -| Integration tests (Go) | `tests/integration/*_test.go` | Yes | `make test-integration` | +| Integration tests (Go) | `tests/integration/*_test.go`, `internal/cache/*_integration_test.go` | Yes | `make test-integration` | | E2E tests (SDK) | `tests/e2e/sdk/*.test.ts` | Yes | `make test-e2e` | - **Unit tests** live beside the code they test (e.g., `internal/discovery/discovery_test.go`). They use mocks or embedded NATS (in-process, no Docker needed). @@ -352,7 +352,7 @@ Shared test utilities live in `internal/testutil/`. The packages log through `sl ### Adding New Tests - **Unit test for `internal/foo/`** → create `internal/foo/foo_test.go` (same package). -- **Integration test needing Docker** → add a subtest under `tests/integration/` (e.g. a new file with `//go:build integration`). +- **Integration test needing Docker** → add a subtest under `tests/integration/` (e.g. a new file with `//go:build integration`). A test of one package against its own external server — the shared cache backend against Redis, Valkey and Dragonfly containers — lives beside the package instead (`internal/cache/redis_integration_test.go`, same build tag), and the package is listed in the `test-integration` target. - **E2E test via SDK** → add a `tests/e2e/sdk/*.test.ts` file. These tests exercise the full pipeline (ingest → ClickHouse → query) through the TypeScript SDK. Run with `make test-e2e`. - **Test helpers** → add to `internal/testutil/` (Go) or `tests/e2e/sdk/helpers.ts` (E2E). diff --git a/go.mod b/go.mod index ca418e893..d75355341 100644 --- a/go.mod +++ b/go.mod @@ -26,9 +26,13 @@ require ( github.com/golang-jwt/jwt/v5 v5.3.1 github.com/google/uuid v1.6.0 github.com/ilyakaznacheev/cleanenv v1.5.0 + github.com/klauspost/compress v1.19.2 + github.com/moby/moby/api v1.55.0 + github.com/moby/moby/client v0.5.1 github.com/nats-io/nats-server/v2 v2.14.6 github.com/nats-io/nats.go v1.53.1 github.com/prometheus/client_golang v1.24.1 + github.com/redis/rueidis v1.0.78 github.com/samber/slog-multi v1.8.0 github.com/samber/slog-sampling v1.7.0 github.com/stretchr/testify v1.12.1 @@ -132,7 +136,6 @@ require ( github.com/jedib0t/go-pretty/v6 v6.7.10 // indirect github.com/jmespath/go-jmespath v0.4.0 // indirect github.com/joho/godotenv v1.5.1 // indirect - github.com/klauspost/compress v1.19.2 // indirect github.com/knadh/profiler v0.2.0 // indirect github.com/kr/pretty v0.3.1 // indirect github.com/kr/text v0.2.0 // indirect @@ -147,8 +150,6 @@ require ( github.com/minio/highwayhash v1.0.4 // indirect github.com/moby/docker-image-spec v1.3.1 // indirect github.com/moby/go-archive v0.3.0 // indirect - github.com/moby/moby/api v1.55.0 // indirect - github.com/moby/moby/client v0.5.1 // indirect github.com/moby/patternmatcher v0.6.1 // indirect github.com/moby/sys/sequential v0.7.0 // indirect github.com/moby/sys/user v0.4.1 // indirect diff --git a/go.sum b/go.sum index 71dc27272..3b203e221 100644 --- a/go.sum +++ b/go.sum @@ -286,6 +286,8 @@ github.com/nats-io/nuid v1.0.1 h1:5iA8DT8V7q8WK2EScv2padNa/rTESc1KdnPw4TC2paw= github.com/nats-io/nuid v1.0.1/go.mod h1:19wcPz3Ph3q0Jbyiqsd0kePYG7A95tJPxeL+1OSON2c= github.com/nikolaydubina/treemap v1.2.5 h1:oSC5z/qnsGLbkU2IihSrh2pS7uDjUq7ipGj8aw8bfII= github.com/nikolaydubina/treemap v1.2.5/go.mod h1:8+wLGh917AyeJqBN1D5KM26tv6W/XfvsY+nfJd04/u8= +github.com/onsi/gomega v1.42.1 h1:iN1rCUX+44NZ1Dc97MPoeFYbFR0vh8zxoxMFwKdyZ6I= +github.com/onsi/gomega v1.42.1/go.mod h1:REff/hsDsodHoKlWsP2mAPhu1+5/6hVYNf9rIEBpeSg= github.com/opencontainers/go-digest v1.0.0 h1:apOUWs51W5PlhuyGyz9FCeeBIOUDA/6nW8Oi/yOhh5U= github.com/opencontainers/go-digest v1.0.0/go.mod h1:0JzlMkj0TRzQZfJkVvzbP0HBR3IKzErnv2BNG4W4MAM= github.com/opencontainers/image-spec v1.1.1 h1:y0fUlFfIZhPF1W537XOLg0/fcx6zcHCJwooC2xJA040= @@ -320,6 +322,8 @@ github.com/prometheus/procfs v0.21.1 h1:GljZCt+zSTS+NZq88cyQ1LjZ+RCHp3uVuabBWA5+ github.com/prometheus/procfs v0.21.1/go.mod h1:aB55Cww9pdSJVHk0hUf0inxWyyjPogFIjmHKYgMKmtY= github.com/puzpuzpuz/xsync/v4 v4.5.0 h1:vOSWu6b57/emh+L/Cw0BeQfvxa/cogFywXHeGUxQxAg= github.com/puzpuzpuz/xsync/v4 v4.5.0/go.mod h1:VJDmTCJMBt8igNxnkQd86r+8KUeN1quSfNKu5bLYFQo= +github.com/redis/rueidis v1.0.78 h1:hJXpEgC9IYfdwY4hCdaGYsfK+oUaAqvhI/GMy5akVJI= +github.com/redis/rueidis v1.0.78/go.mod h1:L8mnCQJJaSNL6I4pIR6Rz732HTGS9vmuXm0yT9dRvjo= github.com/rivo/uniseg v0.1.0/go.mod h1:J6wj4VEh+S6ZtnVlnTBMWIodfgj8LQOQFoIToxlJtxc= github.com/rivo/uniseg v0.2.0/go.mod h1:J6wj4VEh+S6ZtnVlnTBMWIodfgj8LQOQFoIToxlJtxc= github.com/rivo/uniseg v0.4.7 h1:WUdvkW8uEhrYfLC4ZzdpI2ztxP1I582+49Oc5Mq64VQ= diff --git a/internal/cache/breaker.go b/internal/cache/breaker.go new file mode 100644 index 000000000..c90ab3a91 --- /dev/null +++ b/internal/cache/breaker.go @@ -0,0 +1,65 @@ +package cache + +import ( + "sync" + "time" +) + +// breaker stops a failing cache server from costing every request its full +// timeout: after threshold consecutive failures it opens, and while open +// callers skip the server entirely. Once openFor has passed, one caller is +// told to probe; the probe's outcome closes the breaker or reopens it. +type breaker struct { + threshold int + openFor time.Duration + now func() time.Time + + mu sync.Mutex + failures int + open bool + openedAt time.Time + probing bool +} + +func newBreaker(threshold int, openFor time.Duration, now func() time.Time) *breaker { + return &breaker{threshold: threshold, openFor: openFor, now: now} +} + +// allow reports whether a call may go to the server, and whether the caller +// should start the one probe that decides whether an open breaker closes. +func (b *breaker) allow() (ok, probe bool) { + b.mu.Lock() + defer b.mu.Unlock() + if !b.open { + return true, false + } + if b.probing || b.now().Sub(b.openedAt) < b.openFor { + return false, false + } + b.probing = true + return false, true +} + +// success records a call the server answered, and closes the breaker. +func (b *breaker) success() { + b.mu.Lock() + defer b.mu.Unlock() + b.failures, b.open, b.probing = 0, false, false +} + +// failure records a call the server did not answer in time, and opens the +// breaker at the threshold — or at once, for a failed probe. +func (b *breaker) failure() { + b.mu.Lock() + defer b.mu.Unlock() + b.failures++ + if b.probing || b.failures >= b.threshold { + b.open, b.openedAt, b.probing = true, b.now(), false + } +} + +func (b *breaker) isOpen() bool { + b.mu.Lock() + defer b.mu.Unlock() + return b.open +} diff --git a/internal/cache/breaker_test.go b/internal/cache/breaker_test.go new file mode 100644 index 000000000..e6dc2296c --- /dev/null +++ b/internal/cache/breaker_test.go @@ -0,0 +1,59 @@ +package cache + +import ( + "testing" + "time" + + "github.com/stretchr/testify/assert" +) + +type fakeClock struct{ t time.Time } + +func (c *fakeClock) now() time.Time { return c.t } + +func TestBreaker(t *testing.T) { + t.Parallel() + clock := &fakeClock{t: time.Unix(0, 0)} + b := newBreaker(3, 5*time.Second, clock.now) + allow := func() (bool, bool) { return b.allow() } + + ok, probe := allow() + assert.True(t, ok) + assert.False(t, probe) + + b.failure() + b.failure() + b.success() // a success resets the run + b.failure() + b.failure() + assert.False(t, b.isOpen(), "two in a row is below the threshold") + b.failure() + assert.True(t, b.isOpen()) + + ok, probe = allow() + assert.False(t, ok, "open: skip the server") + assert.False(t, probe, "not due for a probe yet") + + clock.t = clock.t.Add(5 * time.Second) + ok, probe = allow() + assert.False(t, ok) + assert.True(t, probe, "due: this caller probes") + ok, probe = allow() + assert.False(t, ok) + assert.False(t, probe, "one probe at a time") + + b.failure() // the probe failed: open for another period + assert.True(t, b.isOpen()) + clock.t = clock.t.Add(4 * time.Second) + _, probe = allow() + assert.False(t, probe) + clock.t = clock.t.Add(time.Second) + _, probe = allow() + assert.True(t, probe) + + b.success() + assert.False(t, b.isOpen()) + ok, probe = allow() + assert.True(t, ok) + assert.False(t, probe) +} diff --git a/internal/cache/cache.go b/internal/cache/cache.go index c9e66a96b..11fbc8512 100644 --- a/internal/cache/cache.go +++ b/internal/cache/cache.go @@ -19,7 +19,8 @@ type Entry struct { // the query runs orphans the fill rather than re-homing pre-write rows under // the post-bump versions (#382). The zero Snapshot makes Set a no-op. type Snapshot struct { - key string // the backend's key for the entry at the observed versions + key string // the backend's key for the entry at the observed versions + tokens []byte // RedisCache: the version tokens read, in tokenKeys order } // ErrForeignDependency is a Lookup whose dependencies name a tenant other diff --git a/internal/cache/export_test.go b/internal/cache/export_test.go new file mode 100644 index 000000000..7c392c13e --- /dev/null +++ b/internal/cache/export_test.go @@ -0,0 +1,12 @@ +package cache + +// Hooks for the integration tests in package cache_test. + +// Pending reports how many token bumps r still owes the server. +func Pending(r *RedisCache) int { return r.pending.len() } + +// Bypassed reports whether r is skipping the server. +func Bypassed(r *RedisCache) bool { return r.bypassed() } + +// DecodedFactor is how many times MaxValueBytes a value may decompress to. +const DecodedFactor = decodedFactor diff --git a/internal/cache/metrics.go b/internal/cache/metrics.go new file mode 100644 index 000000000..cf76eb7b3 --- /dev/null +++ b/internal/cache/metrics.go @@ -0,0 +1,106 @@ +package cache + +import ( + "context" + "errors" + "time" + + "go.opentelemetry.io/otel" + "go.opentelemetry.io/otel/attribute" + "go.opentelemetry.io/otel/metric" +) + +// Lookup outcomes for the "result" attribute of wavehouse_cache_lookups_total. +const ( + resultHit = "hit" + resultMiss = "miss" // nothing stored + resultStale = "stale" // stored under versions since bumped + resultBypass = "bypass" // server skipped: breaker open or not yet connected + resultError = "error" // server failed or timed out +) + +// metrics are a shared-cache backend's instruments. No tenant attribute: +// lookups are the hot path, and tenants are unbounded. +type metrics struct { + backend attribute.KeyValue + lookups metric.Int64Counter + duration metric.Float64Histogram + invalidation metric.Int64Counter + valueBytes metric.Int64Histogram + oversize metric.Int64Counter + setFailures metric.Int64Counter + registration metric.Registration +} + +// newMetrics builds the instruments on the global meter provider, with +// gauges read from breakerOpen and pending on every collection. Call it +// after observability.InitProvider, like every other instrument. +func newMetrics(backend string, breakerOpen func() bool, pending func() int) (*metrics, error) { + meter := otel.Meter("wavehouse-cache") + m := &metrics{backend: attribute.String("backend", backend)} + var errs [9]error + m.lookups, errs[0] = meter.Int64Counter("wavehouse_cache_lookups_total", + metric.WithDescription("Shared-cache lookups by result: hit, miss, stale (stored under since-bumped versions), bypass (server skipped), error")) + m.duration, errs[1] = meter.Float64Histogram("wavehouse_cache_op_duration_seconds", + metric.WithDescription("Shared-cache round-trip time by op: lookup, set, invalidate"), metric.WithUnit("s"), + metric.WithExplicitBucketBoundaries(.0001, .00025, .0005, .001, .0025, .005, .01, .025, .05, .1, .25)) + m.invalidation, errs[2] = meter.Int64Counter("wavehouse_cache_invalidations_total", + metric.WithDescription("Version-token bumps by result: ok, or deferred to the pending retry set")) + m.valueBytes, errs[3] = meter.Int64Histogram("wavehouse_cache_value_bytes", + metric.WithDescription("Size of each value written to the shared cache, after compression"), metric.WithUnit("By"), + metric.WithExplicitBucketBoundaries(256, 1<<10, 4<<10, 16<<10, 64<<10, 256<<10, 1<<20, 4<<20)) + m.oversize, errs[4] = meter.Int64Counter("wavehouse_cache_oversize_total", + metric.WithDescription("Results not cached because they exceed the value size limit")) + m.setFailures, errs[5] = meter.Int64Counter("wavehouse_cache_set_failures_total", + metric.WithDescription("Shared-cache writes that failed, by reason: oom, timeout, other")) + breakerGauge, err := meter.Int64ObservableGauge("wavehouse_cache_breaker_open", + metric.WithDescription("1 while the shared cache is being bypassed (circuit breaker open, or never connected), else 0")) + errs[6] = err + pendingGauge, err := meter.Int64ObservableGauge("wavehouse_cache_invalidations_pending", + metric.WithDescription("Version-token bumps not yet delivered to the shared cache; entries they would orphan may be served stale meanwhile")) + errs[7] = err + m.registration, errs[8] = meter.RegisterCallback(func(_ context.Context, o metric.Observer) error { + var open int64 + if breakerOpen() { + open = 1 + } + o.ObserveInt64(breakerGauge, open, metric.WithAttributes(m.backend)) + o.ObserveInt64(pendingGauge, int64(pending()), metric.WithAttributes(m.backend)) + return nil + }, breakerGauge, pendingGauge) + if err := errors.Join(errs[:]...); err != nil { + return nil, err + } + return m, nil +} + +func (m *metrics) lookup(result string) { + m.lookups.Add(context.Background(), 1, metric.WithAttributes(m.backend, attribute.String("result", result))) +} + +func (m *metrics) op(op string, start time.Time) { + m.duration.Record(context.Background(), time.Since(start).Seconds(), + metric.WithAttributes(m.backend, attribute.String("op", op))) +} + +func (m *metrics) invalidated(result string, n int) { + if n > 0 { + m.invalidation.Add(context.Background(), int64(n), metric.WithAttributes(m.backend, attribute.String("result", result))) + } +} + +func (m *metrics) stored(n int) { + m.valueBytes.Record(context.Background(), int64(n), metric.WithAttributes(m.backend)) +} + +func (m *metrics) tooLarge() { + m.oversize.Add(context.Background(), 1, metric.WithAttributes(m.backend)) +} + +func (m *metrics) setFailed(reason string) { + m.setFailures.Add(context.Background(), 1, metric.WithAttributes(m.backend, attribute.String("reason", reason))) +} + +func (m *metrics) close() { + _ = m.registration.Unregister() +} diff --git a/internal/cache/pending.go b/internal/cache/pending.go new file mode 100644 index 000000000..c19b21786 --- /dev/null +++ b/internal/cache/pending.go @@ -0,0 +1,78 @@ +package cache + +import ( + "sync" + + "github.com/Wave-RF/WaveHouse/internal/tenant" +) + +// pendingBumps holds the token bumps an invalidation could not deliver, until +// a retry lands them. Repeats of one key coalesce; past max keys the set +// collapses to one tenant bump per affected tenant — coarser, never less. +// Landing a bump late is still correct: a fresh token orphans the pre-write +// entries and any fill in between. +type pendingBumps struct { + prefix string + max int + + mu sync.Mutex + keys map[string]pendingKey + gen uint64 +} + +type pendingKey struct { + tenant tenant.ID + gen uint64 // when last added; take drops a key only if it is unchanged +} + +func newPendingBumps(prefix string, maxKeys int) *pendingBumps { + return &pendingBumps{prefix: prefix, max: maxKeys, keys: map[string]pendingKey{}} +} + +// add records keys of tenant id as owed a bump. +func (p *pendingBumps) add(id tenant.ID, keys ...string) { + p.mu.Lock() + defer p.mu.Unlock() + p.gen++ + for _, k := range keys { + p.keys[k] = pendingKey{tenant: id, gen: p.gen} + } + if len(p.keys) <= p.max { + return + } + collapsed := make(map[string]pendingKey, len(p.keys)) + for _, pk := range p.keys { + collapsed[tenantTokenKey(p.prefix, pk.tenant)] = pendingKey{tenant: pk.tenant, gen: p.gen} + } + p.keys = collapsed +} + +// snapshot returns the keys owed a bump with the generation each was added +// at, for a later done. +func (p *pendingBumps) snapshot() map[string]uint64 { + p.mu.Lock() + defer p.mu.Unlock() + out := make(map[string]uint64, len(p.keys)) + for k, pk := range p.keys { + out[k] = pk.gen + } + return out +} + +// done drops keys whose bump landed — unless one was added again since the +// snapshot, whose bump the landed one may have preceded. +func (p *pendingBumps) done(landed map[string]uint64) { + p.mu.Lock() + defer p.mu.Unlock() + for k, gen := range landed { + if pk, ok := p.keys[k]; ok && pk.gen == gen { + delete(p.keys, k) + } + } +} + +func (p *pendingBumps) len() int { + p.mu.Lock() + defer p.mu.Unlock() + return len(p.keys) +} diff --git a/internal/cache/pending_test.go b/internal/cache/pending_test.go new file mode 100644 index 000000000..cc8f331da --- /dev/null +++ b/internal/cache/pending_test.go @@ -0,0 +1,45 @@ +package cache + +import ( + "testing" + + "github.com/stretchr/testify/assert" +) + +func TestPendingBumps_Coalesce(t *testing.T) { + t.Parallel() + p := newPendingBumps("wh", 10) + p.add("acme", "k1", "k2") + p.add("acme", "k1") + assert.Equal(t, 2, p.len()) + + snap := p.snapshot() + p.done(snap) + assert.Zero(t, p.len()) +} + +// A key added again after the snapshot a drain worked from stays owed: the +// drain's bump may have landed before the write the new add is for. +func TestPendingBumps_ReAddedKeyStays(t *testing.T) { + t.Parallel() + p := newPendingBumps("wh", 10) + p.add("acme", "k1", "k2") + snap := p.snapshot() + p.add("acme", "k1") + p.done(snap) + assert.Equal(t, map[string]uint64{"k1": 2}, p.snapshot()) +} + +func TestPendingBumps_OverflowCollapsesToTenants(t *testing.T) { + t.Parallel() + p := newPendingBumps("wh", 3) + p.add("acme", "wh:{acme}:B:a", "wh:{acme}:B:b") + p.add("globex", "wh:{globex}:B:a") + assert.Equal(t, 3, p.len()) + + p.add("acme", "wh:{acme}:B:c") + got := p.snapshot() + assert.Len(t, got, 2) + assert.Contains(t, got, "wh:{acme}:T") + assert.Contains(t, got, "wh:{globex}:T") +} diff --git a/internal/cache/redis.go b/internal/cache/redis.go new file mode 100644 index 000000000..ede95a534 --- /dev/null +++ b/internal/cache/redis.go @@ -0,0 +1,644 @@ +package cache + +import ( + "bytes" + "context" + "crypto/tls" + "errors" + "fmt" + "log/slog" + "net" + "strings" + "sync" + "sync/atomic" + "time" + + "github.com/redis/rueidis" + + "github.com/Wave-RF/WaveHouse/internal/tenant" +) + +// Redis deployment modes for RedisConfig.Mode. +const ( + RedisStandalone = "standalone" + RedisCluster = "cluster" + RedisSentinel = "sentinel" +) + +// Defaults for RedisConfig's zero values. +const ( + DefaultRedisKeyPrefix = "wh" + DefaultRedisTimeout = 100 * time.Millisecond + DefaultRedisDialTimeout = time.Second + DefaultRedisMaxValueBytes = 1 << 20 + DefaultRedisCompressMinBytes = 1 << 10 + DefaultRedisVersionTTL = 7 * 24 * time.Hour + DefaultRedisPendingMax = 100_000 + defaultBreakerThreshold = 5 + defaultBreakerOpenFor = 5 * time.Second +) + +const ( + drainMinBackoff = 100 * time.Millisecond + drainMaxBackoff = 10 * time.Second + drainIdle = time.Second + drainBatch = 1000 + dialMaxBackoff = 30 * time.Second + closeDrainBudget = time.Second +) + +var errBypassed = errors.New("cache: redis unavailable") + +// RedisConfig configures a RedisCache. A zero field takes its Default* +// value, except CompressMinBytes, where 0 means never compress. +type RedisConfig struct { + Addrs []string // host:port; several are seeds (cluster) or sentinels + Mode string // RedisStandalone (""), RedisCluster or RedisSentinel + SentinelMaster string // the master set name, Mode RedisSentinel + Username string + Password string + DB int // standalone and sentinel only + TLS *tls.Config // nil: plaintext + + KeyPrefix string // leads every key; separates deployments sharing a server + Timeout time.Duration // per operation + DialTimeout time.Duration + MaxValueBytes int // largest value stored, after compression + CompressMinBytes int // zstd-compress values at least this large + VersionTTL time.Duration // idle lifetime of a version token, jittered ±10% + PendingMax int // undelivered bumps kept before collapsing to tenant bumps + + BreakerThreshold int // consecutive failures that open the breaker + BreakerOpenFor time.Duration // how long it stays open before a probe +} + +func (c RedisConfig) withDefaults() (RedisConfig, error) { + if len(c.Addrs) == 0 { + return c, errors.New("cache: redis needs at least one address") + } + for _, a := range c.Addrs { + if _, _, err := net.SplitHostPort(a); err != nil { + return c, fmt.Errorf("cache: redis address %q: %w", a, err) + } + } + switch c.Mode { + case "": + c.Mode = RedisStandalone + case RedisStandalone, RedisCluster, RedisSentinel: + default: + return c, fmt.Errorf("cache: redis mode %q: want %s, %s or %s", c.Mode, RedisStandalone, RedisCluster, RedisSentinel) + } + if c.Mode == RedisCluster && c.DB != 0 { + return c, errors.New("cache: redis cluster has only database 0") + } + if c.Mode == RedisSentinel && c.SentinelMaster == "" { + return c, errors.New("cache: redis sentinel mode needs the master set name") + } + if c.DB < 0 { + return c, fmt.Errorf("cache: redis db %d is negative", c.DB) + } + if strings.ContainsAny(c.KeyPrefix, "{}") { + return c, fmt.Errorf("cache: redis key prefix %q may not contain a hash tag brace", c.KeyPrefix) + } + for _, v := range []struct { + name string + n int64 + }{ + {"timeout", int64(c.Timeout)}, + {"dial timeout", int64(c.DialTimeout)}, + {"max value bytes", int64(c.MaxValueBytes)}, + {"compress min bytes", int64(c.CompressMinBytes)}, + {"version ttl", int64(c.VersionTTL)}, + {"pending max", int64(c.PendingMax)}, + {"breaker threshold", int64(c.BreakerThreshold)}, + {"breaker open for", int64(c.BreakerOpenFor)}, + } { + if v.n < 0 { + return c, fmt.Errorf("cache: redis %s is negative", v.name) + } + } + c.KeyPrefix = cmpOr(c.KeyPrefix, DefaultRedisKeyPrefix) + c.Timeout = cmpOr(c.Timeout, DefaultRedisTimeout) + c.DialTimeout = cmpOr(c.DialTimeout, DefaultRedisDialTimeout) + c.MaxValueBytes = cmpOr(c.MaxValueBytes, DefaultRedisMaxValueBytes) + c.VersionTTL = cmpOr(c.VersionTTL, DefaultRedisVersionTTL) + c.PendingMax = cmpOr(c.PendingMax, DefaultRedisPendingMax) + c.BreakerThreshold = cmpOr(c.BreakerThreshold, defaultBreakerThreshold) + c.BreakerOpenFor = cmpOr(c.BreakerOpenFor, defaultBreakerOpenFor) + return c, nil +} + +func cmpOr[T comparable](v, def T) T { + var zero T + if v == zero { + return def + } + return v +} + +func (c RedisConfig) clientOption() rueidis.ClientOption { + opt := rueidis.ClientOption{ + InitAddress: c.Addrs, + Username: c.Username, + Password: c.Password, + SelectDB: c.DB, + TLSConfig: c.TLS, + Dialer: net.Dialer{Timeout: c.DialTimeout}, + ClientName: "wavehouse", + DisableCache: true, // no client-side caching until the near-cache (E5) + ForceSingleClient: c.Mode == RedisStandalone, + } + if c.Mode == RedisSentinel { + opt.Sentinel = rueidis.SentinelOption{MasterSet: c.SentinelMaster, TLSConfig: c.TLS, Dialer: opt.Dialer} + } + return opt +} + +// RedisCache is a Cache shared by every process pointed at one Redis — +// or Valkey, Dragonfly, ElastiCache, MemoryDB: it uses only GET, SET and +// MGET, no scripts and no client tracking. +// +// Versions are random tokens, one per tenant, per table and per scope, +// under the tenant's hash tag; a bump sets a fresh one. A value carries the +// tokens it was computed under and is a hit only while they are all still +// current, so a lost token (eviction, expiry, a restart) can only cause +// misses. A lookup is one round trip. The server failing or timing out is +// a miss, a skipped fill and a deferred invalidation — never a failed query. +type RedisCache struct { + cfg RedisConfig + opt rueidis.ClientOption + maxDecoded int + + client atomic.Pointer[rueidis.Client] // nil until the first connection + codec *codec + breaker *breaker + pending *pendingBumps + metrics *metrics + + ctx context.Context // cancelled by Close + cancel context.CancelFunc + wg sync.WaitGroup + wake chan struct{} + closeOnce sync.Once +} + +var _ Cache = (*RedisCache)(nil) + +// NewRedis builds a RedisCache. A malformed cfg is an error; a server that +// cannot be reached within the dial timeout is not — the cache starts +// bypassed and keeps dialing in the background. +func NewRedis(cfg RedisConfig) (*RedisCache, error) { + cfg, err := cfg.withDefaults() + if err != nil { + return nil, err + } + maxDecoded := cfg.MaxValueBytes * decodedFactor + cd, err := newCodec(cfg.CompressMinBytes, maxDecoded) + if err != nil { + return nil, fmt.Errorf("cache: zstd: %w", err) + } + ctx, cancel := context.WithCancel(context.Background()) + r := &RedisCache{ + cfg: cfg, + opt: cfg.clientOption(), + maxDecoded: maxDecoded, + codec: cd, + breaker: newBreaker(cfg.BreakerThreshold, cfg.BreakerOpenFor, time.Now), + pending: newPendingBumps(cfg.KeyPrefix, cfg.PendingMax), + ctx: ctx, + cancel: cancel, + wake: make(chan struct{}, 1), + } + if r.metrics, err = newMetrics("redis", r.bypassed, r.pending.len); err != nil { + cancel() + cd.close() + return nil, fmt.Errorf("cache: metrics: %w", err) + } + r.wg.Add(1) + go r.drainLoop() + if err := r.dial(); err != nil { + r.wg.Add(1) + go r.dialLoop() + } + return r, nil +} + +func (r *RedisCache) dial() error { + c, err := rueidis.NewClient(r.opt) + if err != nil { + level := slog.LevelWarn + if isAuthError(err) { + level = slog.LevelError + } + slog.Log(r.ctx, level, "cache: redis unreachable; bypassing the cache and retrying", + "addrs", r.cfg.Addrs, "error", err) + return err + } + if r.ctx.Err() != nil { + c.Close() + return r.ctx.Err() + } + r.client.Store(&c) + r.nudge() + return nil +} + +func (r *RedisCache) dialLoop() { + defer r.wg.Done() + backoff := time.Second + for { + select { + case <-r.ctx.Done(): + return + case <-time.After(backoff): + } + if r.dial() == nil { + slog.InfoContext(r.ctx, "cache: redis connected", "addrs", r.cfg.Addrs) + return + } + backoff = min(backoff*2, dialMaxBackoff) + } +} + +func isAuthError(err error) bool { + msg := err.Error() + return strings.Contains(msg, "WRONGPASS") || strings.Contains(msg, "NOAUTH") || strings.Contains(msg, "NOPERM") +} + +// conn returns the client to send to, or nil when the server is to be +// skipped: not connected yet, or the breaker open. The caller that finds an +// open breaker due for a probe starts it. +func (r *RedisCache) conn() rueidis.Client { + cp := r.client.Load() + if cp == nil { + return nil + } + ok, probe := r.breaker.allow() + if probe { + go r.probe(*cp) + } + if !ok { + return nil + } + return *cp +} + +func (r *RedisCache) probe(c rueidis.Client) { + ctx, cancel := context.WithTimeout(r.ctx, r.cfg.Timeout) + defer cancel() + if err := c.Do(ctx, c.B().Ping().Build()).Error(); err != nil { + r.breaker.failure() + return + } + r.breaker.success() + slog.InfoContext(ctx, "cache: redis reachable again; cache back in use") + r.nudge() +} + +// bypassed reports whether operations are skipping the server. +func (r *RedisCache) bypassed() bool { + return r.client.Load() == nil || r.breaker.isOpen() +} + +// record feeds an operation's outcome to the breaker. A reply from the +// server, even an error reply, shows it is up; a caller that gave up first +// shows nothing about it. +func (r *RedisCache) record(parent context.Context, err error) { + if err == nil || rueidis.IsRedisNil(err) { + r.breaker.success() + return + } + if _, ok := rueidis.IsRedisErr(err); ok { + r.breaker.success() + return + } + if parent.Err() != nil { + return + } + r.breaker.failure() +} + +// Lookup reads the tokens deps fold and the entry for sha in one pipelined +// round trip. A token that does not exist yet is created, never read as a +// value, so its first use is a miss. +func (r *RedisCache) Lookup(ctx context.Context, id tenant.ID, sha string, deps []Namespace) (Entry, Snapshot, error) { + for _, d := range deps { + if d.Tenant != id { + return Entry{}, Snapshot{}, fmt.Errorf("%w: %q under %q", ErrForeignDependency, d.Tenant, id) + } + } + keys := tokenKeys(r.cfg.KeyPrefix, id, deps) + if len(keys) > maxTokenKeys { + return Entry{}, Snapshot{}, fmt.Errorf("cache: %d dependencies is more than a value can record", len(deps)) + } + c := r.conn() + if c == nil { + r.metrics.lookup(resultBypass) + return Entry{}, Snapshot{}, nil + } + defer r.metrics.op("lookup", time.Now()) + opCtx, cancel := context.WithTimeout(ctx, r.cfg.Timeout) + defer cancel() + + vkey := valueKey(r.cfg.KeyPrefix, id, sha, deps) + res := c.DoMulti(opCtx, c.B().Mget().Key(keys...).Build(), c.B().Get().Key(vkey).Build()) + tokens, missing, err := readTokens(res[0]) + if err != nil { + return r.lookupFailed(ctx, err) + } + val, err := res[1].AsBytes() + if err != nil && !rueidis.IsRedisNil(err) { + return r.lookupFailed(ctx, err) + } + r.record(ctx, nil) + + if len(missing) > 0 { + if tokens, err = r.createTokens(opCtx, c, keys, missing); err != nil { + return r.lookupFailed(ctx, err) + } + r.metrics.lookup(resultMiss) + return Entry{}, Snapshot{key: vkey, tokens: tokens}, nil + } + snap := Snapshot{key: vkey, tokens: tokens} + if val == nil { + r.metrics.lookup(resultMiss) + return Entry{}, snap, nil + } + stored, expiresAt, payload, err := r.codec.decode(val) + if err != nil { + slog.DebugContext(ctx, "cache: unreadable value; treating as a miss", "key", vkey, "error", err) + r.metrics.lookup(resultMiss) + return Entry{}, snap, nil + } + remaining := time.Until(expiresAt) + if !bytes.Equal(stored, tokens) || remaining <= 0 { + r.metrics.lookup(resultStale) + return Entry{}, snap, nil + } + r.metrics.lookup(resultHit) + return Entry{Value: payload, TTL: remaining}, snap, nil +} + +func (r *RedisCache) lookupFailed(ctx context.Context, err error) (Entry, Snapshot, error) { + r.record(ctx, err) + r.metrics.lookup(resultError) + return Entry{}, Snapshot{}, fmt.Errorf("cache: redis lookup: %w", err) +} + +// readTokens concatenates an MGET reply's tokens, listing the indexes of the +// keys that do not exist. +func readTokens(res rueidis.RedisResult) (tokens []byte, missing []int, err error) { + msgs, err := res.ToArray() + if err != nil { + return nil, nil, err + } + tokens = make([]byte, 0, len(msgs)*tokenLen) + for i := range msgs { + if msgs[i].IsNil() { + missing = append(missing, i) + tokens = append(tokens, make([]byte, tokenLen)...) + continue + } + s, err := msgs[i].ToString() + if err != nil { + return nil, nil, err + } + if len(s) != tokenLen { + return nil, nil, fmt.Errorf("version token is %d bytes, not %d: is another program writing under this key prefix?", len(s), tokenLen) + } + tokens = append(tokens, s...) + } + return tokens, missing, nil +} + +// createTokens sets each missing token — only if still missing, as another +// process may create it first — and reads them all back, in one round trip: +// the tokens share a slot, so the pipeline runs in order on one node. +func (r *RedisCache) createTokens(ctx context.Context, c rueidis.Client, keys []string, missing []int) ([]byte, error) { + cmds := make(rueidis.Commands, 0, len(missing)+1) + for _, i := range missing { + tok := newToken() + cmds = append(cmds, c.B().Set().Key(keys[i]).Value(rueidis.BinaryString(tok)).Nx().Ex(jitter(r.cfg.VersionTTL, tok)).Build()) + } + cmds = append(cmds, c.B().Mget().Key(keys...).Build()) + res := c.DoMulti(ctx, cmds...) + for _, rr := range res[:len(missing)] { + if err := rr.Error(); err != nil && !rueidis.IsRedisNil(err) { + return nil, err + } + } + tokens, still, err := readTokens(res[len(missing)]) + if err != nil { + return nil, err + } + if len(still) > 0 { + return nil, errors.New("version token vanished as it was created: is the server evicting everything?") + } + return tokens, nil +} + +// Set stores value with the tokens snap read, for ttl. A value over the size +// limit is not stored; neither is anything while the server is bypassed. +func (r *RedisCache) Set(ctx context.Context, snap Snapshot, value []byte, ttl time.Duration) error { + if snap.key == "" || ttl <= 0 { + return nil + } + if len(value) > r.maxDecoded { + r.metrics.tooLarge() + return nil + } + b := r.codec.encode(snap.tokens, time.Now().Add(ttl), value) + if len(b) > r.cfg.MaxValueBytes { + r.metrics.tooLarge() + return nil + } + c := r.conn() + if c == nil { + return nil + } + defer r.metrics.op("set", time.Now()) + opCtx, cancel := context.WithTimeout(ctx, r.cfg.Timeout) + defer cancel() + err := c.Do(opCtx, c.B().Set().Key(snap.key).Value(rueidis.BinaryString(b)).Px(max(ttl, time.Millisecond)).Build()).Error() + r.record(ctx, err) + if err != nil { + r.metrics.setFailed(setFailureReason(err)) + return fmt.Errorf("cache: redis set: %w", err) + } + r.metrics.stored(len(b)) + return nil +} + +func setFailureReason(err error) string { + if re, ok := rueidis.IsRedisErr(err); ok && strings.HasPrefix(re.Error(), "OOM") { + return "oom" + } + if errors.Is(err, context.DeadlineExceeded) { + return "timeout" + } + return "other" +} + +// Invalidate sets a fresh token for every token the namespaces' writes +// reach, in one pipelined round trip. Bumps the server does not take are +// kept and retried until it does, and reported as an error meanwhile. +func (r *RedisCache) Invalidate(ctx context.Context, namespaces []Namespace) (uint64, error) { + owner := map[string]tenant.ID{} + for _, ns := range namespaces { + for _, k := range bumpKeys(r.cfg.KeyPrefix, ns) { + owner[k] = ns.Tenant + } + } + return uint64(len(namespaces)), r.bump(ctx, owner) +} + +// InvalidateTenant sets a fresh tenant token, orphaning every entry of id. +func (r *RedisCache) InvalidateTenant(ctx context.Context, id tenant.ID) error { + return r.bump(ctx, map[string]tenant.ID{tenantTokenKey(r.cfg.KeyPrefix, id): id}) +} + +func (r *RedisCache) bump(ctx context.Context, owner map[string]tenant.ID) error { + if len(owner) == 0 { + return nil + } + c := r.conn() + if c == nil { + r.deferBumps(owner) + return fmt.Errorf("%w: %d invalidations deferred", errBypassed, len(owner)) + } + defer r.metrics.op("invalidate", time.Now()) + opCtx, cancel := context.WithTimeout(ctx, r.cfg.Timeout) + defer cancel() + keys := make([]string, 0, len(owner)) + cmds := make(rueidis.Commands, 0, len(owner)) + for k := range owner { + keys = append(keys, k) + cmds = append(cmds, r.bumpCmd(c, k)) + } + failed := map[string]tenant.ID{} + var firstErr error + for i, rr := range c.DoMulti(opCtx, cmds...) { + if err := rr.Error(); err != nil { + failed[keys[i]] = owner[keys[i]] + firstErr = cmpOr(firstErr, err) + } + } + r.record(ctx, firstErr) + r.metrics.invalidated("ok", len(owner)-len(failed)) + if len(failed) > 0 { + r.deferBumps(failed) + return fmt.Errorf("cache: redis invalidate (%d deferred): %w", len(failed), firstErr) + } + return nil +} + +func (r *RedisCache) bumpCmd(c rueidis.Client, key string) rueidis.Completed { + tok := newToken() + return c.B().Set().Key(key).Value(rueidis.BinaryString(tok)).Ex(jitter(r.cfg.VersionTTL, tok)).Build() +} + +func (r *RedisCache) deferBumps(owner map[string]tenant.ID) { + for k, id := range owner { + r.pending.add(id, k) + } + r.metrics.invalidated("deferred", len(owner)) +} + +// nudge wakes the drain loop now, rather than at its next tick. +func (r *RedisCache) nudge() { + select { + case r.wake <- struct{}{}: + default: + } +} + +func (r *RedisCache) drainLoop() { + defer r.wg.Done() + backoff := drainMinBackoff + timer := time.NewTimer(drainIdle) + defer timer.Stop() + for { + select { + case <-r.ctx.Done(): + return + case <-r.wake: + backoff = drainMinBackoff + case <-timer.C: + } + wait := drainIdle + if r.breaker.isOpen() { + r.conn() // starts the probe when due, so an idle process recovers too + } + if r.pending.len() > 0 { + if r.drain(r.ctx) { + backoff = drainMinBackoff + } else { + wait, backoff = backoff, min(backoff*2, drainMaxBackoff) + } + } + timer.Reset(wait) + } +} + +// drain delivers the pending bumps, reporting whether none remain. +func (r *RedisCache) drain(ctx context.Context) bool { + owed := r.pending.snapshot() + keys := make([]string, 0, len(owed)) + for k := range owed { + keys = append(keys, k) + } + for len(keys) > 0 { + batch := keys[:min(drainBatch, len(keys))] + keys = keys[len(batch):] + c := r.conn() + if c == nil { + return false + } + cmds := make(rueidis.Commands, 0, len(batch)) + for _, k := range batch { + cmds = append(cmds, r.bumpCmd(c, k)) + } + opCtx, cancel := context.WithTimeout(ctx, r.cfg.Timeout) + landed := map[string]uint64{} + var firstErr error + for i, rr := range c.DoMulti(opCtx, cmds...) { + if err := rr.Error(); err != nil { + firstErr = cmpOr(firstErr, err) + continue + } + landed[batch[i]] = owed[batch[i]] + } + cancel() + r.record(ctx, firstErr) + r.pending.done(landed) + r.metrics.invalidated("ok", len(landed)) + if firstErr != nil { + return false + } + } + return r.pending.len() == 0 +} + +// Close stops the background loops, makes one last attempt at the pending +// bumps, and closes the connection. Bumps still undelivered are lost: the +// entries they would orphan are served until their TTL. +func (r *RedisCache) Close() error { + r.closeOnce.Do(func() { + r.cancel() + r.wg.Wait() + if r.pending.len() > 0 { + ctx, cancel := context.WithTimeout(context.Background(), closeDrainBudget) + r.drain(ctx) + cancel() + if n := r.pending.len(); n > 0 { + slog.Warn("cache: closing with undelivered invalidations; entries they orphan stay cached until their TTL", "pending", n) + } + } + if cp := r.client.Load(); cp != nil { + (*cp).Close() + } + r.codec.close() + r.metrics.close() + }) + return nil +} diff --git a/internal/cache/redis_codec.go b/internal/cache/redis_codec.go new file mode 100644 index 000000000..cf6f4b5ec --- /dev/null +++ b/internal/cache/redis_codec.go @@ -0,0 +1,194 @@ +package cache + +import ( + "cmp" + "crypto/rand" + "crypto/sha256" + "encoding/binary" + "encoding/hex" + "errors" + "fmt" + "slices" + "time" + + "github.com/klauspost/compress/zstd" + + "github.com/Wave-RF/WaveHouse/internal/tenant" +) + +// tokenLen is the size of a version token: random, so a token key that is +// lost (evicted, expired, flushed) and recreated can never match a value +// stored under its predecessor, as a counter restarting at 0 would. +const tokenLen = 8 + +// decodedFactor bounds a value's decompressed size at this multiple of the +// stored-size limit, refusing a zip bomb planted in a shared server. +const decodedFactor = 8 + +// Value layout: format, flags, expires-at (unix ms), token count, tokens, +// payload. Big-endian. +const ( + valueFormat = 1 + flagZstd = 1 << 0 + headerLen = 1 + 1 + 8 + 2 + maxTokenKeys = 1<<16 - 1 +) + +var errCorruptValue = errors.New("cache: corrupt value") + +func newToken() []byte { + b := make([]byte, tokenLen) + _, _ = rand.Read(b) // never fails (crypto/rand, Go ≥ 1.24) + return b +} + +// tenantTokenKey is the key of tenant id's token. Every token key carries +// the tenant as a hash tag, so all of a tenant's tokens share one cluster +// slot and a lookup reads them with one MGET. +func tenantTokenKey(prefix string, id tenant.ID) string { + return prefix + ":{" + string(id) + "}:T" +} + +func tableTokenKey(prefix string, id tenant.ID, table string) string { + return prefix + ":{" + string(id) + "}:B:" + table +} + +func scopeTokenKey(prefix string, id tenant.ID, table, scope string) string { + return prefix + ":{" + string(id) + "}:S:" + table + ":" + scope +} + +// sortedDeps returns deps in canonical order without duplicates. +func sortedDeps(deps []Namespace) []Namespace { + out := slices.Clone(deps) + slices.SortFunc(out, func(a, b Namespace) int { + return cmp.Or(cmp.Compare(a.Table, b.Table), cmp.Compare(a.Scope, b.Scope)) + }) + return slices.Compact(out) +} + +// tokenKeys lists the tokens a result for deps of tenant id is filed under, +// in canonical order: the tenant's, then each dep's table and scope tokens. +// A dep with scope s folds B:table (bumped by a whole-table write) and +// S:table:s (bumped by a write to s, and — for s == "" — by any scoped write +// to the table), the same lattice LocalCache's version index encodes. +func tokenKeys(prefix string, id tenant.ID, deps []Namespace) []string { + keys := []string{tenantTokenKey(prefix, id)} + for _, d := range sortedDeps(deps) { + keys = append(keys, tableTokenKey(prefix, id, d.Table), scopeTokenKey(prefix, id, d.Table, d.Scope)) + } + slices.Sort(keys[1:]) + return append(keys[:1], slices.Compact(keys[1:])...) +} + +// bumpKeys lists the tokens an invalidation of ns replaces: a whole-table +// write the table's, a scoped write its scope's and the whole-table view's. +func bumpKeys(prefix string, ns Namespace) []string { + if ns.Scope == "" { + return []string{tableTokenKey(prefix, ns.Tenant, ns.Table)} + } + return []string{scopeTokenKey(prefix, ns.Tenant, ns.Table, ns.Scope), scopeTokenKey(prefix, ns.Tenant, ns.Table, "")} +} + +// valueKey names the entry for sha over deps. It carries no versions, so a +// refill overwrites in place, and no hash tag, so one tenant's values spread +// across a cluster's shards. +func valueKey(prefix string, id tenant.ID, sha string, deps []Namespace) string { + h := sha256.New() + h.Write([]byte(sha)) + h.Write([]byte{0}) + for _, d := range sortedDeps(deps) { + h.Write([]byte(d.Table)) + h.Write([]byte{0}) + h.Write([]byte(d.Scope)) + h.Write([]byte{0}) + } + return prefix + ":q:" + string(id) + ":" + hex.EncodeToString(h.Sum(nil)) +} + +// codec compresses and frames values. Its zstd encoder and decoder are safe +// for concurrent EncodeAll/DecodeAll. +type codec struct { + enc *zstd.Encoder + dec *zstd.Decoder + compressMin int + maxDecoded int +} + +func newCodec(compressMin, maxDecoded int) (*codec, error) { + enc, err := zstd.NewWriter(nil, zstd.WithEncoderLevel(zstd.SpeedFastest)) + if err != nil { + return nil, err + } + dec, err := zstd.NewReader(nil, zstd.WithDecoderMaxMemory(uint64(maxDecoded)), zstd.WithDecoderConcurrency(0)) //nolint:gosec // maxDecoded is a positive config-derived int + if err != nil { + return nil, err + } + return &codec{enc: enc, dec: dec, compressMin: compressMin, maxDecoded: maxDecoded}, nil +} + +func (c *codec) close() { + _ = c.enc.Close() + c.dec.Close() +} + +// encode frames payload with the tokens it was computed under, compressing +// it when that is enabled, the payload is large enough, and it helps. +func (c *codec) encode(tokens []byte, expiresAt time.Time, payload []byte) []byte { + var flags byte + body := payload + if c.compressMin > 0 && len(payload) >= c.compressMin { + if z := c.enc.EncodeAll(payload, nil); len(z) < len(payload) { + body, flags = z, flagZstd + } + } + out := make([]byte, headerLen, headerLen+len(tokens)+len(body)) + out[0] = valueFormat + out[1] = flags + binary.BigEndian.PutUint64(out[2:10], uint64(expiresAt.UnixMilli())) + binary.BigEndian.PutUint16(out[10:12], uint16(len(tokens)/tokenLen)) //nolint:gosec // bounded by maxTokenKeys at Lookup + out = append(out, tokens...) + return append(out, body...) +} + +// decode is encode's inverse. An unknown format — a newer process's value +// during a rolling upgrade — is an error, which the caller reads as a miss. +func (c *codec) decode(b []byte) (tokens []byte, expiresAt time.Time, payload []byte, err error) { + if len(b) < headerLen { + return nil, time.Time{}, nil, errCorruptValue + } + if b[0] != valueFormat { + return nil, time.Time{}, nil, fmt.Errorf("%w: format %d", errCorruptValue, b[0]) + } + flags := b[1] + expiresAt = time.UnixMilli(int64(binary.BigEndian.Uint64(b[2:10]))) //nolint:gosec // written by encode + n := int(binary.BigEndian.Uint16(b[10:12])) * tokenLen + if len(b) < headerLen+n { + return nil, time.Time{}, nil, errCorruptValue + } + tokens = b[headerLen : headerLen+n] + payload = b[headerLen+n:] + switch flags { + case 0: + case flagZstd: + if payload, err = c.dec.DecodeAll(payload, nil); err != nil { + return nil, time.Time{}, nil, fmt.Errorf("%w: %w", errCorruptValue, err) + } + if len(payload) > c.maxDecoded { + return nil, time.Time{}, nil, fmt.Errorf("%w: decodes to %d bytes", errCorruptValue, len(payload)) + } + default: + return nil, time.Time{}, nil, fmt.Errorf("%w: flags %#x", errCorruptValue, flags) + } + return tokens, expiresAt, payload, nil +} + +// jitter spreads d by ±10%, drawing on the token being written so a batch of +// bumps doesn't expire together. +func jitter(d time.Duration, token []byte) time.Duration { + span := int64(d) / 5 + if span <= 0 { + return d + } + r := int64(binary.BigEndian.Uint64(token) % uint64(span)) //nolint:gosec // r < span, an int64 + return d - time.Duration(span/2) + time.Duration(r) +} diff --git a/internal/cache/redis_codec_test.go b/internal/cache/redis_codec_test.go new file mode 100644 index 000000000..68c423d0f --- /dev/null +++ b/internal/cache/redis_codec_test.go @@ -0,0 +1,198 @@ +package cache + +import ( + "bytes" + "strings" + "testing" + "time" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" +) + +func TestTokenKeys(t *testing.T) { + t.Parallel() + tests := []struct { + name string + deps []Namespace + want []string + }{ + {"no deps is the tenant token alone", nil, []string{"wh:{acme}:T"}}, + { + "a dep folds its table and scope tokens", + []Namespace{{Tenant: "acme", Table: "events", Scope: "org_1"}}, + []string{"wh:{acme}:T", "wh:{acme}:B:events", "wh:{acme}:S:events:org_1"}, + }, + { + "a scopeless dep reads the whole-table view", + []Namespace{{Tenant: "acme", Table: "events"}}, + []string{"wh:{acme}:T", "wh:{acme}:B:events", "wh:{acme}:S:events:"}, + }, + { + "shared table tokens and duplicate deps appear once, sorted", + []Namespace{ + {Tenant: "acme", Table: "orders"}, + {Tenant: "acme", Table: "events", Scope: "b"}, + {Tenant: "acme", Table: "events", Scope: "a"}, + {Tenant: "acme", Table: "orders"}, + }, + []string{ + "wh:{acme}:T", "wh:{acme}:B:events", "wh:{acme}:B:orders", + "wh:{acme}:S:events:a", "wh:{acme}:S:events:b", "wh:{acme}:S:orders:", + }, + }, + } + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + t.Parallel() + assert.Equal(t, tt.want, tokenKeys("wh", "acme", tt.deps)) + }) + } +} + +func TestBumpKeys(t *testing.T) { + t.Parallel() + assert.Equal(t, []string{"p:{acme}:B:events"}, bumpKeys("p", Namespace{Tenant: "acme", Table: "events"})) + assert.Equal(t, []string{"p:{acme}:S:events:org_1", "p:{acme}:S:events:"}, + bumpKeys("p", Namespace{Tenant: "acme", Table: "events", Scope: "org_1"})) +} + +// Every key a bump writes is one some lookup reads: otherwise the bump +// orphans nothing. +func TestBumpKeysAreReadByLookups(t *testing.T) { + t.Parallel() + for _, ns := range []Namespace{{Tenant: "acme", Table: "events"}, {Tenant: "acme", Table: "events", Scope: "org_1"}} { + read := map[string]bool{} + for _, scope := range []string{"", ns.Scope} { + for _, k := range tokenKeys("wh", "acme", []Namespace{{Tenant: "acme", Table: ns.Table, Scope: scope}}) { + read[k] = true + } + } + for _, k := range bumpKeys("wh", ns) { + assert.True(t, read[k], k) + } + } +} + +func TestValueKey(t *testing.T) { + t.Parallel() + a, b := Namespace{Tenant: "acme", Table: "events"}, Namespace{Tenant: "acme", Table: "orders", Scope: "x"} + k := valueKey("wh", "acme", "acme:query:abc", []Namespace{a, b}) + assert.True(t, strings.HasPrefix(k, "wh:q:acme:"), k) + assert.NotContains(t, k, "{", "values carry no hash tag, so they spread across shards") + assert.Equal(t, k, valueKey("wh", "acme", "acme:query:abc", []Namespace{b, a, b}), "order and duplicates do not matter") + assert.NotEqual(t, k, valueKey("wh", "acme", "acme:query:abc", []Namespace{a})) + assert.NotEqual(t, k, valueKey("wh", "acme", "acme:query:abd", []Namespace{a, b})) + assert.NotEqual(t, + valueKey("wh", "acme", "q", []Namespace{{Tenant: "acme", Table: "ab", Scope: "c"}}), + valueKey("wh", "acme", "q", []Namespace{{Tenant: "acme", Table: "a", Scope: "bc"}})) +} + +func newTestCodec(t *testing.T, compressMin, maxDecoded int) *codec { + t.Helper() + c, err := newCodec(compressMin, maxDecoded) + require.NoError(t, err) + t.Cleanup(c.close) + return c +} + +func TestCodec_RoundTrip(t *testing.T) { + t.Parallel() + c := newTestCodec(t, 64, 1<<20) + tokens := append(newToken(), newToken()...) + exp := time.UnixMilli(time.Now().Add(time.Minute).UnixMilli()) + tests := []struct { + name string + payload []byte + compressed bool + }{ + {"empty", []byte{}, false}, + {"below the threshold", []byte(`[{"a":1}]`), false}, + {"compressible", bytes.Repeat([]byte(`{"user":"u1","n":42},`), 200), true}, + {"incompressible stays raw", randomBytes(4096), false}, + } + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + t.Parallel() + b := c.encode(tokens, exp, tt.payload) + assert.Equal(t, tt.compressed, b[1]&flagZstd != 0) + if tt.compressed { + assert.Less(t, len(b), len(tt.payload)) + } + gotTokens, gotExp, gotPayload, err := c.decode(b) + require.NoError(t, err) + assert.Equal(t, tokens, gotTokens) + assert.True(t, exp.Equal(gotExp)) + assert.Equal(t, tt.payload, gotPayload) + }) + } +} + +func TestCodec_CompressionOff(t *testing.T) { + t.Parallel() + c := newTestCodec(t, 0, 1<<20) + b := c.encode(newToken(), time.Now(), bytes.Repeat([]byte("a"), 4096)) + assert.Zero(t, b[1]) +} + +func TestCodec_RefusesBadValues(t *testing.T) { + t.Parallel() + c := newTestCodec(t, 1, 1<<10) + good := c.encode(newToken(), time.Now(), []byte("rows")) + bomb := newTestCodec(t, 1, 1<<30).encode(newToken(), time.Now(), make([]byte, 1<<20)) + require.Less(t, len(bomb), 1<<10, "the bomb is small when stored") + + withByte := func(i int, v byte) []byte { + b := bytes.Clone(good) + b[i] = v + return b + } + tests := []struct { + name string + b []byte + }{ + {"shorter than the header", good[:headerLen-1]}, + {"unknown format", withByte(0, 2)}, + {"unknown flags", withByte(1, 0x80)}, + {"fewer tokens than it declares", withByte(11, 200)}, + {"not zstd", append(withByte(1, flagZstd)[:headerLen+tokenLen], "not zstd"...)}, + {"decodes past the limit", bomb}, + } + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + t.Parallel() + _, _, _, err := c.decode(tt.b) + require.ErrorIs(t, err, errCorruptValue) + }) + } +} + +func TestJitter(t *testing.T) { + t.Parallel() + d := 100 * time.Second + for range 1000 { + j := jitter(d, newToken()) + assert.GreaterOrEqual(t, j, 90*time.Second) + assert.Less(t, j, 110*time.Second) + } + assert.Equal(t, time.Nanosecond, jitter(time.Nanosecond, newToken())) +} + +func TestNewToken(t *testing.T) { + t.Parallel() + seen := map[string]bool{} + for range 1000 { + tok := newToken() + require.Len(t, tok, tokenLen) + require.False(t, seen[string(tok)]) + seen[string(tok)] = true + } +} + +func randomBytes(n int) []byte { + b := make([]byte, 0, n) + for len(b) < n { + b = append(b, newToken()...) + } + return b[:n] +} diff --git a/internal/cache/redis_integration_test.go b/internal/cache/redis_integration_test.go new file mode 100644 index 000000000..fb902e735 --- /dev/null +++ b/internal/cache/redis_integration_test.go @@ -0,0 +1,418 @@ +//go:build integration + +package cache_test + +import ( + "bytes" + "context" + "fmt" + "net" + "strconv" + "strings" + "sync/atomic" + "testing" + "time" + + "github.com/moby/moby/api/types/container" + "github.com/moby/moby/api/types/network" + "github.com/moby/moby/client" + "github.com/redis/rueidis" + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + "github.com/testcontainers/testcontainers-go" + "github.com/testcontainers/testcontainers-go/wait" + + "github.com/Wave-RF/WaveHouse/internal/cache" + "github.com/Wave-RF/WaveHouse/internal/tenant" + "github.com/Wave-RF/WaveHouse/internal/testutil/cachetest" +) + +// Pinned: the servers the shared cache is documented to run on. +const ( + redisImage = "redis:8.10.2-alpine" + valkeyImage = "valkey/valkey:8.1.10-alpine" + dragonflyImage = "docker.dragonflydb.io/dragonflydb/dragonfly:v2.0.0" +) + +// maxValue is the stored-size limit the tests run with; small, so the +// oversize case stays cheap. +const maxValue = 64 << 10 + +type server struct { + ctr testcontainers.Container + addr string + mode string +} + +// noPersistence keeps the image's VOLUME /data off an anonymous volume. +func noPersistence(hc *container.HostConfig) { + hc.Tmpfs = map[string]string{"/data": ""} +} + +func startContainer(t *testing.T, req testcontainers.ContainerRequest, port string) (testcontainers.Container, string) { + t.Helper() + ctx := context.Background() + if req.HostConfigModifier == nil { + req.HostConfigModifier = noPersistence + } + req.WaitingFor = wait.ForListeningPort(port).WithStartupTimeout(90 * time.Second) + ctr, err := testcontainers.GenericContainer(ctx, testcontainers.GenericContainerRequest{ContainerRequest: req, Started: true}) + testcontainers.CleanupContainer(t, ctr) + require.NoError(t, err) + host, err := ctr.Host(ctx) + require.NoError(t, err) + mapped, err := ctr.MappedPort(ctx, port) + require.NoError(t, err) + return ctr, net.JoinHostPort(host, mapped.Port()) +} + +func startStandalone(t *testing.T, image string, cmd ...string) *server { + t.Helper() + ctr, addr := startContainer(t, testcontainers.ContainerRequest{ + Image: image, Cmd: cmd, ExposedPorts: []string{"6379/tcp"}, + }, "6379/tcp") + s := &server{ctr: ctr, addr: addr, mode: cache.RedisStandalone} + waitReady(t, s, nil) + return s +} + +func startRedis(t *testing.T) *server { + return startStandalone(t, redisImage, "redis-server", "--save", "", "--appendonly", "no") +} + +func startValkey(t *testing.T) *server { + return startStandalone(t, valkeyImage, "valkey-server", "--save", "", "--appendonly", "no") +} + +func startDragonfly(t *testing.T) *server { + return startStandalone(t, dragonflyImage, "--proactor_threads=2", "--maxmemory=512mb") +} + +// startCluster runs a one-node Redis Cluster owning every slot: enough for +// the server to enforce cluster semantics — CROSSSLOT on a multi-key +// command, MOVED routing through the client — which is what the key schema +// must survive. The node announces 127.0.0.1 on a host port bound to the +// same number, so the address the client learns from CLUSTER SLOTS is +// dialable from the test. +func startCluster(t *testing.T) *server { + t.Helper() + port := freePort(t) + p := network.MustParsePort(port + "/tcp") + ctr, _ := startContainer(t, testcontainers.ContainerRequest{ + Image: redisImage, + Cmd: []string{ + "redis-server", "--port", port, "--cluster-enabled", "yes", "--cluster-port", "16379", + "--cluster-announce-ip", "127.0.0.1", "--save", "", "--appendonly", "no", + }, + ExposedPorts: []string{port + "/tcp"}, + HostConfigModifier: func(hc *container.HostConfig) { + noPersistence(hc) + hc.PortBindings = network.PortMap{p: {{HostPort: port}}} + }, + }, port+"/tcp") + code, out, err := ctr.Exec(context.Background(), []string{"redis-cli", "-p", port, "cluster", "addslotsrange", "0", "16383"}) + require.NoError(t, err) + require.Zero(t, code, "%v", out) + s := &server{ctr: ctr, addr: net.JoinHostPort("127.0.0.1", port), mode: cache.RedisCluster} + waitReady(t, s, func(c rueidis.Client) error { + info, err := c.Do(context.Background(), c.B().ClusterInfo().Build()).ToString() + if err == nil && !strings.Contains(info, "cluster_state:ok") { + err = fmt.Errorf("cluster not ready: %q", info) + } + return err + }) + return s +} + +func freePort(t *testing.T) string { + t.Helper() + var lc net.ListenConfig + for { + ln, err := lc.Listen(context.Background(), "tcp", "127.0.0.1:0") + require.NoError(t, err) + port := strconv.Itoa(ln.Addr().(*net.TCPAddr).Port) + require.NoError(t, ln.Close()) + if port != "16379" { + return port + } + } +} + +// raw opens a plain client on s, for the test to reach under the cache. +func raw(t *testing.T, s *server) rueidis.Client { + t.Helper() + c, err := rueidis.NewClient(rueidis.ClientOption{ + InitAddress: []string{s.addr}, DisableCache: true, ForceSingleClient: s.mode == cache.RedisStandalone, + }) + require.NoError(t, err) + t.Cleanup(c.Close) + return c +} + +// waitReady blocks until s answers — a listening port is not yet a server +// that takes commands — and, given ready, until ready passes too. +func waitReady(t *testing.T, s *server, ready func(rueidis.Client) error) { + t.Helper() + var last error + require.Eventually(t, func() bool { + c, err := rueidis.NewClient(rueidis.ClientOption{ + InitAddress: []string{s.addr}, DisableCache: true, ForceSingleClient: true, + }) + if last = err; err != nil { + return false + } + defer c.Close() + if last = c.Do(context.Background(), c.B().Ping().Build()).Error(); last != nil { + return false + } + if ready != nil { + last = ready(c) + } + return last == nil + }, 60*time.Second, 100*time.Millisecond, "server %s not ready: %v", s.addr, last) +} + +var prefixes atomic.Uint64 + +func uniquePrefix() string { return fmt.Sprintf("t%d", prefixes.Add(1)) } + +// open builds a RedisCache on s. Its timeout is generous: the suite runs +// in parallel under -race, and a timed-out lookup is a miss the conformance +// cases would read as a wrong answer. +func open(t *testing.T, s *server, prefix string, tune ...func(*cache.RedisConfig)) *cache.RedisCache { + t.Helper() + cfg := cache.RedisConfig{ + Addrs: []string{s.addr}, Mode: s.mode, KeyPrefix: prefix, + Timeout: 5 * time.Second, MaxValueBytes: maxValue, CompressMinBytes: cache.DefaultRedisCompressMinBytes, + } + for _, f := range tune { + f(&cfg) + } + c, err := cache.NewRedis(cfg) + require.NoError(t, err) + t.Cleanup(func() { _ = c.Close() }) + return c +} + +func TestRedis_Conformance(t *testing.T) { + t.Parallel() + servers := []struct { + name string + start func(*testing.T) *server + }{ + {"redis", startRedis}, + {"valkey", startValkey}, + {"dragonfly", startDragonfly}, + {"redis cluster", startCluster}, + } + for _, sv := range servers { + t.Run(sv.name, func(t *testing.T) { + t.Parallel() + s := sv.start(t) + cachetest.Run(t, + func(t *testing.T) cache.Cache { return open(t, s, uniquePrefix()) }, + cachetest.Options{ + // The raw-size bound: a value past it is refused before + // compression, and the suite's oversize value is zeros, + // which would compress under the stored-size one. + MaxValueBytes: maxValue * cache.DecodedFactor, + NewPair: func(t *testing.T) (cache.Cache, cache.Cache) { + p := uniquePrefix() + return open(t, s, p), open(t, s, p) + }, + }) + t.Run("cross-tenant invalidation spans slots", func(t *testing.T) { + t.Parallel() + testCrossTenantInvalidate(t, open(t, s, uniquePrefix())) + }) + t.Run("compressed values round-trip", func(t *testing.T) { + t.Parallel() + testCompression(t, s) + }) + }) + } +} + +// The ingest worker's shared-tables fan-out bumps one table under several +// tenants in one call: tokens in as many slots, one pipeline. +func testCrossTenantInvalidate(t *testing.T, c *cache.RedisCache) { + ctx := context.Background() + var deps [][]cache.Namespace + for i := range 20 { + id := tenantID(i) + d := []cache.Namespace{{Tenant: id, Table: "events"}} + deps = append(deps, d) + _, snap, err := c.Lookup(ctx, id, "q", d) + require.NoError(t, err) + require.NoError(t, c.Set(ctx, snap, []byte("rows"), time.Minute)) + } + var all []cache.Namespace + for _, d := range deps { + all = append(all, d...) + } + n, err := c.Invalidate(ctx, all) + require.NoError(t, err) + assert.Equal(t, uint64(len(all)), n) + for i, d := range deps { + e, _, err := c.Lookup(ctx, tenantID(i), "q", d) + require.NoError(t, err) + assert.Nil(t, e.Value, tenantID(i)) + } +} + +func tenantID(i int) tenant.ID { return tenant.ID(fmt.Sprintf("tenant-%d", i)) } + +func testCompression(t *testing.T, s *server) { + ctx := context.Background() + prefix := uniquePrefix() + c := open(t, s, prefix) + rows := bytes.Repeat([]byte(`{"user_id":"u-1","event":"click","value":42.5},`), 10_000) + require.Greater(t, len(rows), maxValue, "stored only because it compresses under the limit") + deps := []cache.Namespace{{Tenant: "acme", Table: "events"}} + _, snap, err := c.Lookup(ctx, "acme", "big", deps) + require.NoError(t, err) + require.NoError(t, c.Set(ctx, snap, rows, time.Minute)) + e, _, err := c.Lookup(ctx, "acme", "big", deps) + require.NoError(t, err) + assert.Equal(t, rows, e.Value) + + r := raw(t, s) + keys, err := r.Do(ctx, r.B().Keys().Pattern(prefix+":q:*").Build()).AsStrSlice() + require.NoError(t, err) + require.Len(t, keys, 1) + stored, err := r.Do(ctx, r.B().Strlen().Key(keys[0]).Build()).AsInt64() + require.NoError(t, err) + assert.Less(t, stored, int64(len(rows)/10)) +} + +// A token that is lost — evicted, expired, flushed, a restart without +// persistence — is recreated fresh, so a value stored under its predecessor +// can only miss. A counter recreated at its initial value would serve the +// value filed at that value again: this test fails for one. +func TestRedis_LostTokensAreMisses(t *testing.T) { + t.Parallel() + ctx := context.Background() + s := startRedis(t) + r := raw(t, s) + prefix := uniquePrefix() + c := open(t, s, prefix) + deps := []cache.Namespace{{Tenant: "acme", Table: "events", Scope: "org_1"}} + fill := func(t *testing.T, sha string, deps []cache.Namespace) { + t.Helper() + _, snap, err := c.Lookup(ctx, "acme", sha, deps) + require.NoError(t, err) + require.NoError(t, c.Set(ctx, snap, []byte("rows"), time.Minute)) + e, _, err := c.Lookup(ctx, "acme", sha, deps) + require.NoError(t, err) + require.Equal(t, "rows", string(e.Value)) + } + // Twice: the first lookup recreates the lost tokens, and it is the + // second, reading them back, that a recreated counter would fool. + requireMiss := func(t *testing.T, sha string, deps []cache.Namespace) { + t.Helper() + for range 2 { + e, _, err := c.Lookup(ctx, "acme", sha, deps) + require.NoError(t, err) + require.Nil(t, e.Value) + } + } + keys := func(t *testing.T, pattern string) []string { + t.Helper() + keys, err := r.Do(ctx, r.B().Keys().Pattern(pattern).Build()).AsStrSlice() + require.NoError(t, err) + return keys + } + + for _, lost := range []string{"T", "B:events", "S:events:org_1", "*"} { + fill(t, "q", deps) + fill(t, "pipe", nil) + tokens := keys(t, prefix+":{acme}:"+lost) + require.NotEmpty(t, tokens, lost) + require.NoError(t, r.Do(ctx, r.B().Del().Key(tokens...).Build()).Error()) + require.Len(t, keys(t, prefix+":q:*"), 2, "lost %s: the values survive, only their tokens are gone", lost) + requireMiss(t, "q", deps) + if lost == "T" || lost == "*" { + requireMiss(t, "pipe", nil) + } + } + + fill(t, "q", deps) + require.NoError(t, r.Do(ctx, r.B().Flushall().Build()).Error()) + requireMiss(t, "q", deps) + fill(t, "q", deps) +} + +func dockerClient(t *testing.T) *testcontainers.DockerClient { + t.Helper() + d, err := testcontainers.NewDockerClientWithOpts(context.Background()) + require.NoError(t, err) + t.Cleanup(func() { _ = d.Close() }) + return d +} + +// A server that stops answering costs a request at most about the op +// timeout, then nothing: the breaker opens and the cache is bypassed. +// Invalidations made meanwhile are kept and land once it answers again, and +// a process that boots while it is down starts bypassed and connects later. +func TestRedis_ServerStopsAnswering(t *testing.T) { + t.Parallel() + ctx := context.Background() + s := startRedis(t) + d := dockerClient(t) + prefix := uniquePrefix() + const timeout = 100 * time.Millisecond + a := open(t, s, prefix, func(c *cache.RedisConfig) { + c.Timeout, c.BreakerThreshold, c.BreakerOpenFor = timeout, 3, 300*time.Millisecond + }) + deps := []cache.Namespace{{Tenant: "acme", Table: "events"}} + _, snap, err := a.Lookup(ctx, "acme", "q", deps) + require.NoError(t, err) + require.NoError(t, a.Set(ctx, snap, []byte("pre-write rows"), time.Minute)) + + _, err = d.ContainerPause(ctx, s.ctr.GetContainerID(), client.ContainerPauseOptions{}) + require.NoError(t, err) + paused := true + unpause := func() { + if paused { + paused = false + _, err := d.ContainerUnpause(ctx, s.ctr.GetContainerID(), client.ContainerUnpauseOptions{}) + require.NoError(t, err) + } + } + t.Cleanup(unpause) + + for i := range 3 { + start := time.Now() + e, snap, err := a.Lookup(ctx, "acme", "q", deps) + require.Error(t, err, "lookup %d", i) + assert.Nil(t, e.Value) + assert.Less(t, time.Since(start), 10*timeout, "a lookup costs at most about the timeout") + require.NoError(t, a.Set(ctx, snap, []byte("rows"), time.Minute), "the failed lookup's snapshot files nothing") + } + require.True(t, cache.Bypassed(a), "three timeouts open the breaker") + start := time.Now() + e, _, err := a.Lookup(ctx, "acme", "q", deps) + require.NoError(t, err, "bypassed is a miss, not a failure") + assert.Nil(t, e.Value) + assert.Less(t, time.Since(start), timeout/2, "bypassed costs no round trip") + + _, err = a.Invalidate(ctx, deps) + require.Error(t, err) + assert.Equal(t, 1, cache.Pending(a)) + + bootStart := time.Now() + late := open(t, s, prefix, func(c *cache.RedisConfig) { c.DialTimeout = 200 * time.Millisecond }) + assert.Less(t, time.Since(bootStart), 5*time.Second, "an unanswering server does not hold boot") + assert.True(t, cache.Bypassed(late)) + + unpause() + require.Eventually(t, func() bool { return cache.Pending(a) == 0 && !cache.Bypassed(a) }, 15*time.Second, 50*time.Millisecond, + "the deferred bump lands once the server answers") + require.Eventually(t, func() bool { return !cache.Bypassed(late) }, 15*time.Second, 50*time.Millisecond, + "the late process connects") + + b := open(t, s, prefix) + e, _, err = b.Lookup(ctx, "acme", "q", deps) + require.NoError(t, err) + assert.Nil(t, e.Value, "the fill from before the deferred bump is orphaned for every process") +} diff --git a/internal/cache/redis_test.go b/internal/cache/redis_test.go new file mode 100644 index 000000000..9feff6327 --- /dev/null +++ b/internal/cache/redis_test.go @@ -0,0 +1,231 @@ +package cache + +import ( + "context" + "errors" + "net" + "testing" + "time" + + "github.com/redis/rueidis" + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + "go.opentelemetry.io/otel" + "go.opentelemetry.io/otel/attribute" + sdkmetric "go.opentelemetry.io/otel/sdk/metric" + "go.opentelemetry.io/otel/sdk/metric/metricdata" +) + +func TestRedisConfig_Validation(t *testing.T) { + t.Parallel() + ok := RedisConfig{Addrs: []string{"redis:6379"}} + tests := []struct { + name string + mutate func(c *RedisConfig) + wantErr string + }{ + {"no address", func(c *RedisConfig) { c.Addrs = nil }, "at least one address"}, + {"address without a port", func(c *RedisConfig) { c.Addrs = []string{"redis"} }, `address "redis"`}, + {"unknown mode", func(c *RedisConfig) { c.Mode = "ring" }, `mode "ring"`}, + {"cluster with a db", func(c *RedisConfig) { c.Mode, c.DB = RedisCluster, 1 }, "only database 0"}, + {"sentinel without a master", func(c *RedisConfig) { c.Mode = RedisSentinel }, "master set name"}, + {"negative db", func(c *RedisConfig) { c.DB = -1 }, "negative"}, + {"hash tag in the prefix", func(c *RedisConfig) { c.KeyPrefix = "{wh}" }, "brace"}, + {"negative timeout", func(c *RedisConfig) { c.Timeout = -time.Second }, "timeout is negative"}, + {"negative size", func(c *RedisConfig) { c.MaxValueBytes = -1 }, "max value bytes is negative"}, + } + for _, tt := range tests { + t.Run(tt.name, func(t *testing.T) { + t.Parallel() + c := ok + tt.mutate(&c) + _, err := NewRedis(c) + require.ErrorContains(t, err, tt.wantErr) + }) + } +} + +func TestRedisConfig_Defaults(t *testing.T) { + t.Parallel() + c, err := RedisConfig{Addrs: []string{"redis:6379"}}.withDefaults() + require.NoError(t, err) + assert.Equal(t, RedisStandalone, c.Mode) + assert.Equal(t, DefaultRedisKeyPrefix, c.KeyPrefix) + assert.Equal(t, DefaultRedisTimeout, c.Timeout) + assert.Equal(t, DefaultRedisMaxValueBytes, c.MaxValueBytes) + assert.Equal(t, DefaultRedisVersionTTL, c.VersionTTL) + assert.Zero(t, c.CompressMinBytes, "0 means never compress, not the default") + assert.True(t, c.clientOption().ForceSingleClient) + + c.Mode, c.SentinelMaster = RedisSentinel, "mymaster" + assert.Equal(t, "mymaster", c.clientOption().Sentinel.MasterSet) + c.Mode = RedisCluster + assert.False(t, c.clientOption().ForceSingleClient) +} + +// closedAddr is an address nothing listens on: dials are refused at once. +func closedAddr(t *testing.T) string { + t.Helper() + var lc net.ListenConfig + ln, err := lc.Listen(context.Background(), "tcp", "127.0.0.1:0") + require.NoError(t, err) + addr := ln.Addr().String() + require.NoError(t, ln.Close()) + return addr +} + +// An unreachable server does not fail construction: the cache is bypassed +// — a miss that files nothing, a no-op fill, a deferred invalidation — and +// keeps dialing. +func TestRedis_UnreachableIsBypassed(t *testing.T) { + t.Parallel() + ctx := context.Background() + r, err := NewRedis(RedisConfig{Addrs: []string{closedAddr(t)}, DialTimeout: 100 * time.Millisecond, PendingMax: 2}) + require.NoError(t, err) + t.Cleanup(func() { _ = r.Close() }) + assert.True(t, r.bypassed()) + + deps := []Namespace{{Tenant: "acme", Table: "events"}} + e, snap, err := r.Lookup(ctx, "acme", "q", deps) + require.NoError(t, err) + assert.Nil(t, e.Value) + assert.Empty(t, snap.key, "nothing to file under") + + _, _, err = r.Lookup(ctx, "acme", "q", []Namespace{{Tenant: "globex", Table: "events"}}) + require.ErrorIs(t, err, ErrForeignDependency) + + require.NoError(t, r.Set(ctx, Snapshot{key: "k"}, []byte("rows"), time.Minute)) + + n, err := r.Invalidate(ctx, []Namespace{{Tenant: "acme", Table: "events", Scope: "org_1"}}) + require.ErrorIs(t, err, errBypassed) + assert.Equal(t, uint64(1), n) + assert.Equal(t, 2, r.pending.len()) + require.ErrorIs(t, r.InvalidateTenant(ctx, "globex"), errBypassed) + owed := r.pending.snapshot() + assert.Len(t, owed, 2, "past PendingMax: one bump per tenant") + assert.Contains(t, owed, "wh:{acme}:T") + assert.Contains(t, owed, "wh:{globex}:T") + + require.NoError(t, r.Close()) + require.NoError(t, r.Close(), "idempotent") +} + +func TestRedis_SetDeclinesWithoutTouchingTheServer(t *testing.T) { + t.Parallel() + r, err := NewRedis(RedisConfig{Addrs: []string{closedAddr(t)}, DialTimeout: 100 * time.Millisecond, MaxValueBytes: 64}) + require.NoError(t, err) + t.Cleanup(func() { _ = r.Close() }) + ctx := context.Background() + snap := Snapshot{key: "k", tokens: newToken()} + for _, tt := range []struct { + name string + snap Snapshot + value []byte + ttl time.Duration + }{ + {"zero snapshot", Snapshot{}, []byte("rows"), time.Minute}, + {"zero ttl", snap, []byte("rows"), 0}, + {"over the stored limit", snap, randomBytes(100), time.Minute}, + {"over the decoded limit", snap, make([]byte, 64*decodedFactor+1), time.Minute}, + } { + require.NoError(t, r.Set(ctx, tt.snap, tt.value, tt.ttl), tt.name) + } +} + +func TestRedis_Record(t *testing.T) { + t.Parallel() + r := &RedisCache{breaker: newBreaker(1, time.Hour, time.Now)} + live, cancelled := context.Background(), cancelledCtx() + + r.record(live, rueidis.Nil) + r.record(live, &rueidis.RedisError{}) + r.record(cancelled, context.Canceled) + assert.False(t, r.breaker.isOpen(), "a reply, even an error reply, or the caller giving up says nothing against the server") + + r.record(live, context.DeadlineExceeded) + assert.True(t, r.breaker.isOpen()) +} + +func cancelledCtx() context.Context { + ctx, cancel := context.WithCancel(context.Background()) + cancel() + return ctx +} + +func TestSetFailureReason(t *testing.T) { + t.Parallel() + assert.Equal(t, "timeout", setFailureReason(context.DeadlineExceeded)) + assert.Equal(t, "other", setFailureReason(errors.New("broken pipe"))) +} + +func TestIsAuthError(t *testing.T) { + t.Parallel() + assert.True(t, isAuthError(errors.New("WRONGPASS invalid username-password pair"))) + assert.True(t, isAuthError(errors.New("NOAUTH authentication required"))) + assert.False(t, isAuthError(errors.New("dial tcp: connection refused"))) +} + +func TestReadTokens_RefusesForeignValues(t *testing.T) { + t.Parallel() + _, _, err := readTokens(rueidis.RedisResult{}) + require.Error(t, err) +} + +func TestRedisMetrics(t *testing.T) { + // No t.Parallel(): swaps the global meter provider. + saved := otel.GetMeterProvider() + reader := sdkmetric.NewManualReader() + mp := sdkmetric.NewMeterProvider(sdkmetric.WithReader(reader)) + otel.SetMeterProvider(mp) + t.Cleanup(func() { + _ = mp.Shutdown(context.Background()) + otel.SetMeterProvider(saved) + }) + + r, err := NewRedis(RedisConfig{Addrs: []string{closedAddr(t)}, DialTimeout: 100 * time.Millisecond}) + require.NoError(t, err) + t.Cleanup(func() { _ = r.Close() }) + ctx := context.Background() + _, _, _ = r.Lookup(ctx, "acme", "q", nil) + _, _ = r.Invalidate(ctx, []Namespace{{Tenant: "acme", Table: "events"}}) + r.metrics.op("set", time.Now()) + r.metrics.stored(10) + r.metrics.tooLarge() + r.metrics.setFailed("oom") + + var rm metricdata.ResourceMetrics + require.NoError(t, reader.Collect(ctx, &rm)) + got := map[string]metricdata.Aggregation{} + for _, sm := range rm.ScopeMetrics { + for _, m := range sm.Metrics { + got[m.Name] = m.Data + } + } + sumOf := func(name, key, value string) int64 { + t.Helper() + var n int64 + switch d := got[name].(type) { + case metricdata.Sum[int64]: + for _, dp := range d.DataPoints { + if v, ok := dp.Attributes.Value(attribute.Key(key)); key == "" || ok && v.AsString() == value { + n += dp.Value + } + } + case metricdata.Gauge[int64]: + for _, dp := range d.DataPoints { + n += dp.Value + } + default: + t.Fatalf("%s: %T", name, got[name]) + } + return n + } + assert.Equal(t, int64(1), sumOf("wavehouse_cache_lookups_total", "result", resultBypass)) + assert.Equal(t, int64(1), sumOf("wavehouse_cache_invalidations_total", "result", "deferred")) + assert.Equal(t, int64(1), sumOf("wavehouse_cache_invalidations_pending", "", "")) + assert.Equal(t, int64(1), sumOf("wavehouse_cache_breaker_open", "", "")) + assert.Equal(t, int64(1), sumOf("wavehouse_cache_oversize_total", "", "")) + assert.Equal(t, int64(1), sumOf("wavehouse_cache_set_failures_total", "reason", "oom")) + assert.Contains(t, got, "wavehouse_cache_op_duration_seconds") + assert.Contains(t, got, "wavehouse_cache_value_bytes") +} From e346426faccb0b05eeb9738b4cafb5abfeb11703 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 00:20:07 -0400 Subject: [PATCH 09/79] fix(cache): count only transport failures against the breaker Review round: a reply the backend cannot use (a foreign value under a token key, an unreadable MGET) no longer opens the breaker, and such a token is replaced; the probe follows the same rule. Close drains the pending bumps past an open breaker, and starts no probe once closing. VersionTTL under 2s is refused (it would round to EX 0). Docs: PING in the command list, the breaker rule, compression wording. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 6 +- internal/cache/redis.go | 100 +++++++++++++++-------- internal/cache/redis_integration_test.go | 63 ++++++++++++++ internal/cache/redis_test.go | 9 +- 5 files changed, 137 insertions(+), 43 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 6a39bc8df..08960f289 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A Redis-compatible shared cache backend, built but not yet selectable** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values over 1 KiB are zstd-compressed and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row open a circuit breaker that skips the server until a probe succeeds, and deferred invalidations are retried until they land, collapsing to one tenant-wide bump per tenant past 100,000 keys. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending`, `wavehouse_cache_value_bytes`, `wavehouse_cache_oversize_total`, `wavehouse_cache_set_failures_total{backend,reason}`. Nothing selects it yet: the `cache.backend` switch and the `cache.redis.*` settings arrive with the wiring (E4), so every deployment still runs `LocalCache`. Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression and server-stops-answering cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. +- **A Redis-compatible shared cache backend, built but not yet selectable** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row open a circuit breaker that skips the server until a probe succeeds, and deferred invalidations are retried until they land, collapsing to one tenant-wide bump per tenant past 100,000 keys. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending`, `wavehouse_cache_value_bytes`, `wavehouse_cache_oversize_total`, `wavehouse_cache_set_failures_total{backend,reason}`. Nothing selects it yet: the `cache.backend` switch and the `cache.redis.*` settings arrive with the wiring (E4), so every deployment still runs `LocalCache`. Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression and server-stops-answering cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 53850f49f..32f743875 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -112,10 +112,10 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on, each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope: one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)). `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. - **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: a namespace key is `..
.
.`, and a query key is folded with the tenant's version and each dependency's namespace key and namespace version, so bumping a table (a scopeless write) or one scope — scope is reserved and empty today, so every write is the whole-table bump — orphans every dependent entry without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results, and the flat directory's single tenant simply carries the `0` prefix. `BumpTenant` (behind `InvalidateTenant`) advances the tenant version that leads every namespace key of one tenant, orphaning its every namespace and every cached query in one step (the tenant version is folded into every query key, so a pipe result with no dependencies is orphaned too) — a table no bump ever keyed included, which an enumeration of the index would miss — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. -- **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET` and `MGET`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, so a token lost to eviction, expiry or a restart is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing. Nothing selects this backend yet — the boot config and wiring arrive with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s E4. +- **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET`, `MGET` and `PING`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token is replaced, so a token lost to eviction, expiry or a restart is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing. Nothing selects this backend yet — the boot config and wiring arrive with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s E4. - **redis_codec.go** — the key schema (`tokenKeys`, `bumpKeys`, `valueKey`) and the value frame: format, flags, expiry, the token list, the payload, zstd-compressed from `CompressMinBytes` when that is smaller. A stored value is capped at `MaxValueBytes` and a decoded one at eight times that, which refuses a zip bomb planted in a shared server; an unknown format, as a newer process writes during a rolling upgrade, is a miss. -- **breaker.go** — the circuit breaker: `BreakerThreshold` (5) consecutive failures open it, and while open every operation skips the server at once; after `BreakerOpenFor` (5 s) one background `PING` decides whether it closes. A reply from the server, error replies included, counts as a success; a caller that gave up first counts as nothing. -- **pending.go** — invalidations the server did not take, kept per token key (repeats coalesce) and retried by a background loop with backoff from 100 ms to 10 s until they land. Landing late is still correct: a fresh token orphans the pre-write entries and any fill made meanwhile. Past `PendingMax` (100,000) keys the set collapses to one tenant bump per affected tenant. Bumps still owed when a process stops are lost after one last attempt, and the entries they would orphan are served until their TTL — the same failure as a worker stopping between an insert and its invalidation. +- **breaker.go** — the circuit breaker: `BreakerThreshold` (5) consecutive failures open it, and while open every operation skips the server at once; after `BreakerOpenFor` (5 s) one background `PING` decides whether it closes. Only a transport failure or a timeout counts against the server: any reply, an error reply or one the backend cannot use included, counts as a success, for operations and the probe alike, and a caller that gave up first counts as nothing. +- **pending.go** — invalidations the server did not take, kept per token key (repeats coalesce) and retried by a background loop with backoff from 100 ms to 10 s until they land. Landing late is still correct: a fresh token orphans the pre-write entries and any fill made meanwhile. Past `PendingMax` (100,000) keys the set collapses to one tenant bump per affected tenant. `Close` makes one last attempt at the bumps still owed, past the breaker (an open one is why they are owed); what that attempt cannot deliver is lost, and the entries they would orphan are served until their TTL — the same failure as a worker stopping between an insert and its invalidation. - **metrics.go** — the shared backend's instruments (meter `wavehouse-cache`, `backend="redis"`, no tenant attribute): `wavehouse_cache_lookups_total` by `result` (`hit`, `miss`, `stale` — filed under since-bumped tokens, `bypass` — server skipped, `error`), `wavehouse_cache_op_duration_seconds` by `op` (`lookup`, `set`, `invalidate`), `wavehouse_cache_breaker_open` (1 while bypassed, including before the first connection), `wavehouse_cache_invalidations_total` by `result` (`ok`, `deferred`), `wavehouse_cache_invalidations_pending` (bumps owed: entries they would orphan may be served stale meanwhile), `wavehouse_cache_value_bytes` (stored size), `wavehouse_cache_oversize_total` and `wavehouse_cache_set_failures_total` by `reason` (`oom`, `timeout`, `other`). ### `config/` — Configuration diff --git a/internal/cache/redis.go b/internal/cache/redis.go index ede95a534..30ed35ba6 100644 --- a/internal/cache/redis.go +++ b/internal/cache/redis.go @@ -47,7 +47,12 @@ const ( closeDrainBudget = time.Second ) -var errBypassed = errors.New("cache: redis unavailable") +var ( + errBypassed = errors.New("cache: redis unavailable") + // errMalformedReply is a reply the server sent that this code cannot + // use: the server is up, so it never counts against the breaker. + errMalformedReply = errors.New("cache: malformed reply") +) // RedisConfig configures a RedisCache. A zero field takes its Default* // value, except CompressMinBytes, where 0 means never compress. @@ -117,6 +122,9 @@ func (c RedisConfig) withDefaults() (RedisConfig, error) { return c, fmt.Errorf("cache: redis %s is negative", v.name) } } + if c.VersionTTL != 0 && c.VersionTTL < 2*time.Second { + return c, fmt.Errorf("cache: redis version ttl %s is under 2s: jittered, it would round to EX 0", c.VersionTTL) + } c.KeyPrefix = cmpOr(c.KeyPrefix, DefaultRedisKeyPrefix) c.Timeout = cmpOr(c.Timeout, DefaultRedisTimeout) c.DialTimeout = cmpOr(c.DialTimeout, DefaultRedisDialTimeout) @@ -155,8 +163,8 @@ func (c RedisConfig) clientOption() rueidis.ClientOption { } // RedisCache is a Cache shared by every process pointed at one Redis — -// or Valkey, Dragonfly, ElastiCache, MemoryDB: it uses only GET, SET and -// MGET, no scripts and no client tracking. +// or Valkey, Dragonfly, ElastiCache, MemoryDB: it uses only GET, SET, MGET +// and PING, no scripts and no client tracking. // // Versions are random tokens, one per tenant, per table and per scope, // under the tenant's hash tag; a bump sets a fresh one. A value carries the @@ -274,7 +282,7 @@ func (r *RedisCache) conn() rueidis.Client { return nil } ok, probe := r.breaker.allow() - if probe { + if probe && r.ctx.Err() == nil { go r.probe(*cp) } if !ok { @@ -286,11 +294,10 @@ func (r *RedisCache) conn() rueidis.Client { func (r *RedisCache) probe(c rueidis.Client) { ctx, cancel := context.WithTimeout(r.ctx, r.cfg.Timeout) defer cancel() - if err := c.Do(ctx, c.B().Ping().Build()).Error(); err != nil { - r.breaker.failure() + r.record(r.ctx, c.Do(ctx, c.B().Ping().Build()).Error()) + if r.breaker.isOpen() { return } - r.breaker.success() slog.InfoContext(ctx, "cache: redis reachable again; cache back in use") r.nudge() } @@ -301,14 +308,14 @@ func (r *RedisCache) bypassed() bool { } // record feeds an operation's outcome to the breaker. A reply from the -// server, even an error reply, shows it is up; a caller that gave up first +// server, even an error reply or one this code cannot use, shows it is up; a caller that gave up first // shows nothing about it. func (r *RedisCache) record(parent context.Context, err error) { if err == nil || rueidis.IsRedisNil(err) { r.breaker.success() return } - if _, ok := rueidis.IsRedisErr(err); ok { + if _, ok := rueidis.IsRedisErr(err); ok || errors.Is(err, errMalformedReply) { r.breaker.success() return } @@ -342,18 +349,23 @@ func (r *RedisCache) Lookup(ctx context.Context, id tenant.ID, sha string, deps vkey := valueKey(r.cfg.KeyPrefix, id, sha, deps) res := c.DoMulti(opCtx, c.B().Mget().Key(keys...).Build(), c.B().Get().Key(vkey).Build()) - tokens, missing, err := readTokens(res[0]) + for _, rr := range res { + if err := rr.Error(); err != nil && !rueidis.IsRedisNil(err) { + return r.lookupFailed(ctx, err) + } + } + r.record(ctx, nil) + tokens, missing, foreign, err := readTokens(res[0]) if err != nil { return r.lookupFailed(ctx, err) } val, err := res[1].AsBytes() if err != nil && !rueidis.IsRedisNil(err) { - return r.lookupFailed(ctx, err) + return r.lookupFailed(ctx, fmt.Errorf("%w: %w", errMalformedReply, err)) } - r.record(ctx, nil) - if len(missing) > 0 { - if tokens, err = r.createTokens(opCtx, c, keys, missing); err != nil { + if len(missing) > 0 || len(foreign) > 0 { + if tokens, err = r.createTokens(opCtx, c, keys, missing, foreign); err != nil { return r.lookupFailed(ctx, err) } r.metrics.lookup(resultMiss) @@ -386,11 +398,11 @@ func (r *RedisCache) lookupFailed(ctx context.Context, err error) (Entry, Snapsh } // readTokens concatenates an MGET reply's tokens, listing the indexes of the -// keys that do not exist. -func readTokens(res rueidis.RedisResult) (tokens []byte, missing []int, err error) { +// keys that do not exist and of those holding something that is not a token. +func readTokens(res rueidis.RedisResult) (tokens []byte, missing, foreign []int, err error) { msgs, err := res.ToArray() if err != nil { - return nil, nil, err + return nil, nil, nil, fmt.Errorf("%w: %w", errMalformedReply, err) } tokens = make([]byte, 0, len(msgs)*tokenLen) for i := range msgs { @@ -400,39 +412,47 @@ func readTokens(res rueidis.RedisResult) (tokens []byte, missing []int, err erro continue } s, err := msgs[i].ToString() - if err != nil { - return nil, nil, err - } - if len(s) != tokenLen { - return nil, nil, fmt.Errorf("version token is %d bytes, not %d: is another program writing under this key prefix?", len(s), tokenLen) + if err != nil || len(s) != tokenLen { + foreign = append(foreign, i) + tokens = append(tokens, make([]byte, tokenLen)...) + continue } tokens = append(tokens, s...) } - return tokens, missing, nil + return tokens, missing, foreign, nil } // createTokens sets each missing token — only if still missing, as another -// process may create it first — and reads them all back, in one round trip: -// the tokens share a slot, so the pipeline runs in order on one node. -func (r *RedisCache) createTokens(ctx context.Context, c rueidis.Client, keys []string, missing []int) ([]byte, error) { - cmds := make(rueidis.Commands, 0, len(missing)+1) +// process may create it first — replaces each foreign one, and reads them +// all back, in one round trip: the tokens share a slot, so the pipeline runs +// in order on one node. A fresh token can only cause misses, so replacing +// whatever held a token key is safe. +func (r *RedisCache) createTokens(ctx context.Context, c rueidis.Client, keys []string, missing, foreign []int) ([]byte, error) { + if len(foreign) > 0 { + slog.WarnContext(ctx, "cache: replacing values that are not version tokens; is another program writing under this key prefix?", + "keys", len(foreign), "prefix", r.cfg.KeyPrefix) + } + cmds := make(rueidis.Commands, 0, len(missing)+len(foreign)+1) for _, i := range missing { tok := newToken() cmds = append(cmds, c.B().Set().Key(keys[i]).Value(rueidis.BinaryString(tok)).Nx().Ex(jitter(r.cfg.VersionTTL, tok)).Build()) } + for _, i := range foreign { + cmds = append(cmds, r.bumpCmd(c, keys[i])) + } cmds = append(cmds, c.B().Mget().Key(keys...).Build()) res := c.DoMulti(ctx, cmds...) - for _, rr := range res[:len(missing)] { + for _, rr := range res { if err := rr.Error(); err != nil && !rueidis.IsRedisNil(err) { return nil, err } } - tokens, still, err := readTokens(res[len(missing)]) + tokens, still, bad, err := readTokens(res[len(res)-1]) if err != nil { return nil, err } - if len(still) > 0 { - return nil, errors.New("version token vanished as it was created: is the server evicting everything?") + if len(still) > 0 || len(bad) > 0 { + return nil, fmt.Errorf("%w: version token gone or replaced as it was written", errMalformedReply) } return tokens, nil } @@ -570,7 +590,7 @@ func (r *RedisCache) drainLoop() { r.conn() // starts the probe when due, so an idle process recovers too } if r.pending.len() > 0 { - if r.drain(r.ctx) { + if r.drain(r.ctx, r.conn) { backoff = drainMinBackoff } else { wait, backoff = backoff, min(backoff*2, drainMaxBackoff) @@ -580,8 +600,9 @@ func (r *RedisCache) drainLoop() { } } -// drain delivers the pending bumps, reporting whether none remain. -func (r *RedisCache) drain(ctx context.Context) bool { +// drain delivers the pending bumps through the client conn returns, +// reporting whether none remain. +func (r *RedisCache) drain(ctx context.Context, conn func() rueidis.Client) bool { owed := r.pending.snapshot() keys := make([]string, 0, len(owed)) for k := range owed { @@ -590,7 +611,7 @@ func (r *RedisCache) drain(ctx context.Context) bool { for len(keys) > 0 { batch := keys[:min(drainBatch, len(keys))] keys = keys[len(batch):] - c := r.conn() + c := conn() if c == nil { return false } @@ -628,7 +649,14 @@ func (r *RedisCache) Close() error { r.wg.Wait() if r.pending.len() > 0 { ctx, cancel := context.WithTimeout(context.Background(), closeDrainBudget) - r.drain(ctx) + // Past the breaker: an open one is why bumps are pending, and this + // is the last chance to deliver them. + r.drain(ctx, func() rueidis.Client { + if cp := r.client.Load(); cp != nil { + return *cp + } + return nil + }) cancel() if n := r.pending.len(); n > 0 { slog.Warn("cache: closing with undelivered invalidations; entries they orphan stay cached until their TTL", "pending", n) diff --git a/internal/cache/redis_integration_test.go b/internal/cache/redis_integration_test.go index fb902e735..155caca92 100644 --- a/internal/cache/redis_integration_test.go +++ b/internal/cache/redis_integration_test.go @@ -342,6 +342,33 @@ func TestRedis_LostTokensAreMisses(t *testing.T) { fill(t, "q", deps) } +// A token key holding something that is not a token — another program +// under the prefix, a different token size mid-upgrade — is a reply, not a +// failure: it is replaced, which can only cause misses, and the breaker +// stays closed. +func TestRedis_ForeignTokenIsReplaced(t *testing.T) { + t.Parallel() + ctx := context.Background() + s := startRedis(t) + r := raw(t, s) + prefix := uniquePrefix() + c := open(t, s, prefix, func(c *cache.RedisConfig) { c.BreakerThreshold = 1 }) + deps := []cache.Namespace{{Tenant: "acme", Table: "events"}} + _, snap, err := c.Lookup(ctx, "acme", "q", deps) + require.NoError(t, err) + require.NoError(t, c.Set(ctx, snap, []byte("rows"), time.Minute)) + require.NoError(t, r.Do(ctx, r.B().Set().Key(prefix+":{acme}:B:events").Value("abc").Build()).Error()) + + e, snap, err := c.Lookup(ctx, "acme", "q", deps) + require.NoError(t, err) + assert.Nil(t, e.Value) + assert.False(t, cache.Bypassed(c)) + require.NoError(t, c.Set(ctx, snap, []byte("new rows"), time.Minute)) + e, _, err = c.Lookup(ctx, "acme", "q", deps) + require.NoError(t, err) + assert.Equal(t, "new rows", string(e.Value)) +} + func dockerClient(t *testing.T) *testcontainers.DockerClient { t.Helper() d, err := testcontainers.NewDockerClientWithOpts(context.Background()) @@ -416,3 +443,39 @@ func TestRedis_ServerStopsAnswering(t *testing.T) { require.NoError(t, err) assert.Nil(t, e.Value, "the fill from before the deferred bump is orphaned for every process") } + +// Close makes its last attempt at the pending bumps past the breaker: an +// open one is why they are pending, and the server may be back by now. +func TestRedis_CloseDeliversPastAnOpenBreaker(t *testing.T) { + t.Parallel() + ctx := context.Background() + s := startRedis(t) + d := dockerClient(t) + prefix := uniquePrefix() + a := open(t, s, prefix, func(c *cache.RedisConfig) { + c.Timeout, c.BreakerThreshold, c.BreakerOpenFor = 100*time.Millisecond, 1, time.Hour + }) + deps := []cache.Namespace{{Tenant: "acme", Table: "events"}} + _, snap, err := a.Lookup(ctx, "acme", "q", deps) + require.NoError(t, err) + require.NoError(t, a.Set(ctx, snap, []byte("pre-write rows"), time.Minute)) + + _, err = d.ContainerPause(ctx, s.ctr.GetContainerID(), client.ContainerPauseOptions{}) + require.NoError(t, err) + _, _, err = a.Lookup(ctx, "acme", "q", deps) + require.Error(t, err) + require.True(t, cache.Bypassed(a)) + _, err = a.Invalidate(ctx, deps) + require.Error(t, err) + _, err = d.ContainerUnpause(ctx, s.ctr.GetContainerID(), client.ContainerUnpauseOptions{}) + require.NoError(t, err) + + require.True(t, cache.Bypassed(a), "the breaker stays open for its hour") + require.Equal(t, 1, cache.Pending(a)) + require.NoError(t, a.Close()) + assert.Zero(t, cache.Pending(a)) + + e, _, err := open(t, s, prefix).Lookup(ctx, "acme", "q", deps) + require.NoError(t, err) + assert.Nil(t, e.Value, "the bump Close delivered orphans the fill") +} diff --git a/internal/cache/redis_test.go b/internal/cache/redis_test.go index 9feff6327..9aafb4f12 100644 --- a/internal/cache/redis_test.go +++ b/internal/cache/redis_test.go @@ -3,6 +3,7 @@ package cache import ( "context" "errors" + "fmt" "net" "testing" "time" @@ -33,6 +34,7 @@ func TestRedisConfig_Validation(t *testing.T) { {"hash tag in the prefix", func(c *RedisConfig) { c.KeyPrefix = "{wh}" }, "brace"}, {"negative timeout", func(c *RedisConfig) { c.Timeout = -time.Second }, "timeout is negative"}, {"negative size", func(c *RedisConfig) { c.MaxValueBytes = -1 }, "max value bytes is negative"}, + {"version ttl under EX's resolution", func(c *RedisConfig) { c.VersionTTL = 1500 * time.Millisecond }, "under 2s"}, } for _, tt := range tests { t.Run(tt.name, func(t *testing.T) { @@ -140,6 +142,7 @@ func TestRedis_Record(t *testing.T) { r.record(live, rueidis.Nil) r.record(live, &rueidis.RedisError{}) r.record(cancelled, context.Canceled) + r.record(live, fmt.Errorf("%w: token is 3 bytes", errMalformedReply)) assert.False(t, r.breaker.isOpen(), "a reply, even an error reply, or the caller giving up says nothing against the server") r.record(live, context.DeadlineExceeded) @@ -165,10 +168,10 @@ func TestIsAuthError(t *testing.T) { assert.False(t, isAuthError(errors.New("dial tcp: connection refused"))) } -func TestReadTokens_RefusesForeignValues(t *testing.T) { +func TestReadTokens_NotAnArrayIsMalformed(t *testing.T) { t.Parallel() - _, _, err := readTokens(rueidis.RedisResult{}) - require.Error(t, err) + _, _, _, err := readTokens(rueidis.RedisResult{}) + require.ErrorIs(t, err, errMalformedReply) } func TestRedisMetrics(t *testing.T) { From d2a87ab7ef14abef09e16c5d4abe12c82aa34716 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 00:25:55 -0400 Subject: [PATCH 10/79] fix(cache): replace token and value keys of the wrong type MGET reads a non-string token key as nil and SET NX will not overwrite it, so such a key disabled caching behind it; it is now replaced on a second round trip. A value key of the wrong type is a miss the fill's SET replaces, not a lookup failure. Docs: VersionTTL is a lifetime from the last bump, and each metric's labels are listed. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- internal/cache/redis.go | 39 ++++++++++++++-------- internal/cache/redis_integration_test.go | 42 ++++++++++++++++-------- 4 files changed, 57 insertions(+), 28 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 08960f289..78a4e3ca0 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A Redis-compatible shared cache backend, built but not yet selectable** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row open a circuit breaker that skips the server until a probe succeeds, and deferred invalidations are retried until they land, collapsing to one tenant-wide bump per tenant past 100,000 keys. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending`, `wavehouse_cache_value_bytes`, `wavehouse_cache_oversize_total`, `wavehouse_cache_set_failures_total{backend,reason}`. Nothing selects it yet: the `cache.backend` switch and the `cache.redis.*` settings arrive with the wiring (E4), so every deployment still runs `LocalCache`. Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression and server-stops-answering cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. +- **A Redis-compatible shared cache backend, built but not yet selectable** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row open a circuit breaker that skips the server until a probe succeeds, and deferred invalidations are retried until they land, collapsing to one tenant-wide bump per tenant past 100,000 keys. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. Nothing selects it yet: the `cache.backend` switch and the `cache.redis.*` settings arrive with the wiring (E4), so every deployment still runs `LocalCache`. Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression and server-stops-answering cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 32f743875..0d617e972 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -112,7 +112,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on, each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope: one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)). `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. - **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: a namespace key is `..
.
.`, and a query key is folded with the tenant's version and each dependency's namespace key and namespace version, so bumping a table (a scopeless write) or one scope — scope is reserved and empty today, so every write is the whole-table bump — orphans every dependent entry without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results, and the flat directory's single tenant simply carries the `0` prefix. `BumpTenant` (behind `InvalidateTenant`) advances the tenant version that leads every namespace key of one tenant, orphaning its every namespace and every cached query in one step (the tenant version is folded into every query key, so a pipe result with no dependencies is orphaned too) — a table no bump ever keyed included, which an enumeration of the index would miss — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. -- **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET`, `MGET` and `PING`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token is replaced, so a token lost to eviction, expiry or a restart is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing. Nothing selects this backend yet — the boot config and wiring arrive with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s E4. +- **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET`, `MGET` and `PING`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token (a string of another length, a list, a hash) is replaced, so a token lost to eviction, expiry or a restart is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing. Nothing selects this backend yet — the boot config and wiring arrive with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s E4. - **redis_codec.go** — the key schema (`tokenKeys`, `bumpKeys`, `valueKey`) and the value frame: format, flags, expiry, the token list, the payload, zstd-compressed from `CompressMinBytes` when that is smaller. A stored value is capped at `MaxValueBytes` and a decoded one at eight times that, which refuses a zip bomb planted in a shared server; an unknown format, as a newer process writes during a rolling upgrade, is a miss. - **breaker.go** — the circuit breaker: `BreakerThreshold` (5) consecutive failures open it, and while open every operation skips the server at once; after `BreakerOpenFor` (5 s) one background `PING` decides whether it closes. Only a transport failure or a timeout counts against the server: any reply, an error reply or one the backend cannot use included, counts as a success, for operations and the probe alike, and a caller that gave up first counts as nothing. - **pending.go** — invalidations the server did not take, kept per token key (repeats coalesce) and retried by a background loop with backoff from 100 ms to 10 s until they land. Landing late is still correct: a fresh token orphans the pre-write entries and any fill made meanwhile. Past `PendingMax` (100,000) keys the set collapses to one tenant bump per affected tenant. `Close` makes one last attempt at the bumps still owed, past the breaker (an open one is why they are owed); what that attempt cannot deliver is lost, and the entries they would orphan are served until their TTL — the same failure as a worker stopping between an insert and its invalidation. diff --git a/internal/cache/redis.go b/internal/cache/redis.go index 30ed35ba6..f9ce1a0fe 100644 --- a/internal/cache/redis.go +++ b/internal/cache/redis.go @@ -70,7 +70,7 @@ type RedisConfig struct { DialTimeout time.Duration MaxValueBytes int // largest value stored, after compression CompressMinBytes int // zstd-compress values at least this large - VersionTTL time.Duration // idle lifetime of a version token, jittered ±10% + VersionTTL time.Duration // a token's lifetime from its last bump, jittered ±10%; reads do not extend it PendingMax int // undelivered bumps kept before collapsing to tenant bumps BreakerThreshold int // consecutive failures that open the breaker @@ -349,19 +349,26 @@ func (r *RedisCache) Lookup(ctx context.Context, id tenant.ID, sha string, deps vkey := valueKey(r.cfg.KeyPrefix, id, sha, deps) res := c.DoMulti(opCtx, c.B().Mget().Key(keys...).Build(), c.B().Get().Key(vkey).Build()) - for _, rr := range res { - if err := rr.Error(); err != nil && !rueidis.IsRedisNil(err) { - return r.lookupFailed(ctx, err) - } + if err := res[0].Error(); err != nil { + return r.lookupFailed(ctx, err) + } + // An error reply such as WRONGTYPE: the value key holds something else, + // which the fill's plain SET replaces, so it is a miss to fill. + valErr := res[1].Error() + _, valReplied := rueidis.IsRedisErr(valErr) + if valErr != nil && !rueidis.IsRedisNil(valErr) && !valReplied { + return r.lookupFailed(ctx, valErr) } r.record(ctx, nil) tokens, missing, foreign, err := readTokens(res[0]) if err != nil { return r.lookupFailed(ctx, err) } - val, err := res[1].AsBytes() - if err != nil && !rueidis.IsRedisNil(err) { - return r.lookupFailed(ctx, fmt.Errorf("%w: %w", errMalformedReply, err)) + var val []byte + if valErr == nil { + if val, err = res[1].AsBytes(); err != nil { + return r.lookupFailed(ctx, fmt.Errorf("%w: %w", errMalformedReply, err)) + } } if len(missing) > 0 || len(foreign) > 0 { @@ -423,8 +430,9 @@ func readTokens(res rueidis.RedisResult) (tokens []byte, missing, foreign []int, } // createTokens sets each missing token — only if still missing, as another -// process may create it first — replaces each foreign one, and reads them -// all back, in one round trip: the tokens share a slot, so the pipeline runs +// process may create it first — replaces each foreign one (a string that is +// not a token, or, on a second round trip, a key of another type), and reads +// them all back, in one round trip: the tokens share a slot, so the pipeline runs // in order on one node. A fresh token can only cause misses, so replacing // whatever held a token key is safe. func (r *RedisCache) createTokens(ctx context.Context, c rueidis.Client, keys []string, missing, foreign []int) ([]byte, error) { @@ -451,10 +459,15 @@ func (r *RedisCache) createTokens(ctx context.Context, c rueidis.Client, keys [] if err != nil { return nil, err } - if len(still) > 0 || len(bad) > 0 { - return nil, fmt.Errorf("%w: version token gone or replaced as it was written", errMalformedReply) + if len(still) == 0 && len(bad) == 0 { + return tokens, nil + } + // Still nil after SET NX: the key holds a list, hash or other non-string, + // which MGET reads as nil and NX will not overwrite. Replace it too. + if len(foreign) == 0 && len(still) > 0 { + return r.createTokens(ctx, c, keys, nil, still) } - return tokens, nil + return nil, fmt.Errorf("%w: version token gone or replaced as it was written", errMalformedReply) } // Set stores value with the tokens snap read, for ttl. A value over the size diff --git a/internal/cache/redis_integration_test.go b/internal/cache/redis_integration_test.go index 155caca92..422e372e7 100644 --- a/internal/cache/redis_integration_test.go +++ b/internal/cache/redis_integration_test.go @@ -342,10 +342,9 @@ func TestRedis_LostTokensAreMisses(t *testing.T) { fill(t, "q", deps) } -// A token key holding something that is not a token — another program -// under the prefix, a different token size mid-upgrade — is a reply, not a -// failure: it is replaced, which can only cause misses, and the breaker -// stays closed. +// A token or value key holding something else — another program under the +// prefix, a different token size mid-upgrade — is a reply, not a failure: +// it is replaced, which can only cause misses, and the breaker stays closed. func TestRedis_ForeignTokenIsReplaced(t *testing.T) { t.Parallel() ctx := context.Background() @@ -357,16 +356,33 @@ func TestRedis_ForeignTokenIsReplaced(t *testing.T) { _, snap, err := c.Lookup(ctx, "acme", "q", deps) require.NoError(t, err) require.NoError(t, c.Set(ctx, snap, []byte("rows"), time.Minute)) - require.NoError(t, r.Do(ctx, r.B().Set().Key(prefix+":{acme}:B:events").Value("abc").Build()).Error()) - - e, snap, err := c.Lookup(ctx, "acme", "q", deps) - require.NoError(t, err) - assert.Nil(t, e.Value) - assert.False(t, cache.Bypassed(c)) - require.NoError(t, c.Set(ctx, snap, []byte("new rows"), time.Minute)) - e, _, err = c.Lookup(ctx, "acme", "q", deps) + valueKeys, err := r.Do(ctx, r.B().Keys().Pattern(prefix+":q:*").Build()).AsStrSlice() require.NoError(t, err) - assert.Equal(t, "new rows", string(e.Value)) + require.Len(t, valueKeys, 1) + + for _, plant := range []struct { + name string + cmd rueidis.Completed + }{ + {"a short string under the table token", r.B().Set().Key(prefix + ":{acme}:B:events").Value("abc").Build()}, + {"", r.B().Del().Key(prefix + ":{acme}:T").Build()}, + {"a list under the tenant token", r.B().Rpush().Key(prefix + ":{acme}:T").Element("x").Build()}, + {"", r.B().Del().Key(valueKeys[0]).Build()}, + {"a hash under the value key", r.B().Hset().Key(valueKeys[0]).FieldValue().FieldValue("f", "v").Build()}, + } { + require.NoError(t, r.Do(ctx, plant.cmd).Error()) + if plant.name == "" { // the first half of a two-step plant + continue + } + e, snap, err := c.Lookup(ctx, "acme", "q", deps) + require.NoError(t, err, plant.name) + assert.Nil(t, e.Value, plant.name) + assert.False(t, cache.Bypassed(c), plant.name) + require.NoError(t, c.Set(ctx, snap, []byte("new rows"), time.Minute), plant.name) + e, _, err = c.Lookup(ctx, "acme", "q", deps) + require.NoError(t, err, plant.name) + assert.Equal(t, "new rows", string(e.Value), plant.name) + } } func dockerClient(t *testing.T) *testcontainers.DockerClient { From d67a46af2e63800bd8cb0721ff49e138e094787f Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 00:54:26 -0400 Subject: [PATCH 11/79] test(cov): keep the unreachable Redis backend out of the e2e gate The e2e stack runs LocalCache and no config selects the Redis backend yet, so its files pulled the e2e suite to 56.4% (floor 60). The integration suite covers them against real servers, and per-suite excludes leave the merged total unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- .testcoverage.yml | 6 ++++++ 1 file changed, 6 insertions(+) diff --git a/.testcoverage.yml b/.testcoverage.yml index aff1a694d..def851f08 100644 --- a/.testcoverage.yml +++ b/.testcoverage.yml @@ -73,3 +73,9 @@ exclude: - ^internal/settings/ - ^cmd/wavehouse/validate\.go$ - ^cmd/wavehouse/bootstrap\.go$ + # The Redis-compatible cache backend: the e2e stack runs LocalCache + # (no Redis server, and no config selects the backend yet — #613 E4), + # so the binary carries these files but e2e can never reach them; they + # pulled the e2e gate to 56.4%. The integration suite runs them against + # real servers, and the merged total still counts them. + - ^internal/cache/(redis|redis_codec|breaker|pending|metrics)\.go$ From c06808400e5a5eb24fdf65b26ac3531d31ed9884 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 02:53:43 -0400 Subject: [PATCH 12/79] feat(config): cache.backend=redis selects the shared cache The boot config's cache.backend now takes redis, configured by a new cache.redis block (WH_CACHE_REDIS_*), and wireCache builds a RedisCache from it, reading the TLS files. A malformed block refuses boot; an unreachable server or a rejected password boots bypassed and keeps reconnecting. An integration test boots two instances over one Redis and one ClickHouse: an ingest through one is served fresh by the other inside the stale entry's TTL, and a paused Redis leaves queries succeeding. The e2e suite now runs against Redis, so its coverage exclude is gone. Part of #613 (E4). Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- .testcoverage.yml | 6 - AGENTS.md | 4 +- CHANGELOG.md | 3 +- README.md | 2 +- config.yaml | 13 +- deployments/compose/dependencies.yaml | 13 + docs/src/content/docs/api.md | 2 +- docs/src/content/docs/architecture.md | 13 +- docs/src/content/docs/configuration.mdx | 78 ++++- docs/src/content/docs/deployment.md | 20 ++ docs/src/content/docs/development.md | 4 +- docs/src/content/docs/getting-started.md | 2 +- docs/src/content/docs/pipes.mdx | 2 +- docs/src/content/docs/settings-directory.mdx | 2 +- internal/app/app_test.go | 71 +++++ internal/app/wire.go | 41 ++- internal/config/backends.go | 40 ++- internal/config/backends_test.go | 2 +- internal/config/cache_redis.go | 153 ++++++++++ internal/config/cache_redis_test.go | 298 +++++++++++++++++++ scripts/orchestrator/main.go | 51 +++- tests/e2e/fixtures/config.yaml | 16 +- tests/integration/shared_cache_test.go | 271 +++++++++++++++++ 23 files changed, 1056 insertions(+), 51 deletions(-) create mode 100644 internal/config/cache_redis.go create mode 100644 internal/config/cache_redis_test.go create mode 100644 tests/integration/shared_cache_test.go diff --git a/.testcoverage.yml b/.testcoverage.yml index def851f08..aff1a694d 100644 --- a/.testcoverage.yml +++ b/.testcoverage.yml @@ -73,9 +73,3 @@ exclude: - ^internal/settings/ - ^cmd/wavehouse/validate\.go$ - ^cmd/wavehouse/bootstrap\.go$ - # The Redis-compatible cache backend: the e2e stack runs LocalCache - # (no Redis server, and no config selects the backend yet — #613 E4), - # so the binary carries these files but e2e can never reach them; they - # pulled the e2e gate to 56.4%. The integration suite runs them against - # real servers, and the merged total still counts them. - - ^internal/cache/(redis|redis_codec|breaker|pending|metrics)\.go$ diff --git a/AGENTS.md b/AGENTS.md index ebdd877c9..68c8f8df9 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -31,10 +31,10 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers - **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` and the `mq.max_bytes_gb` reconcile handing the MQ each served tenant's own gap window and byte budget, and `defaultSetting`/`onDefaultAdopt` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, the same hook's `Hub.Prune` ends the open streams of a tenant no longer served, and `wireCache`'s hook drops, through `LocalCache.Prune`, the cache version index of a tenant no longer served ([#262](https://github.com/Wave-RF/WaveHouse/issues/262))), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it - **`auth/`** — JWT auth middleware: HMAC **or** JWKS verification with `alg` pinned to the active verifier, role extraction from a configurable claim path; always runs, never rejects (bad token → empty role + stashed reason). One verifier per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 9): `Authenticator` keys them by `tenant.ID` — the request store's `settings.Store.Tenant()`, through an injected `TenantSource`; `tenant.Default` on the tenant-exempt routes — built from each tenant's `auth` block by `Reconfigure`, dropped by `Prune` once the tenant stops being served (rejected or removed), released by `Close`; the secrets (`Config`) are boot-level and shared. A JWKS key set is fetched off the boot and reload paths: until one has been stored the verifier is pending and a token-bearing request gets `503` + `Retry-After` from `api.refuseUnverifiable` (`auth.ErrVerifierPending`), never a `default_role` evaluation; refresh is library-managed (Eric, 2026-09-22), response capped at 1 MiB; the operator key's admin role is the request tenant's -- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index), and `RedisCache`, the Redis-compatible shared backend (random version tokens under the tenant's hash tag, one-round-trip lookups, bypass on failure behind a circuit breaker, deferred invalidations retried; built and tested, not yet selectable by config — [#613](https://github.com/Wave-RF/WaveHouse/issues/613) E4). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, a version per tenant, per (tenant, table) and per (tenant, table, scope), keyed by name and bumped in place (one entry per live namespace however often it is bumped, [#262](https://github.com/Wave-RF/WaveHouse/issues/262)) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` drops the tenant's index so its next key gets a process-unique generation, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read and `Set` files the fill under it, so a write landing mid-query orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) +- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index), and `RedisCache`, the Redis-compatible shared backend (random version tokens under the tenant's hash tag, one-round-trip lookups, bypass on failure behind a circuit breaker, deferred invalidations retried; selected by `cache.backend: redis`, configured by the boot config's `cache.redis` block — [#613](https://github.com/Wave-RF/WaveHouse/issues/613)). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, a version per tenant, per (tenant, table) and per (tenant, table, scope), keyed by name and bumped in place (one entry per live namespace however often it is bumped, [#262](https://github.com/Wave-RF/WaveHouse/issues/262)) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` drops the tenant's index so its next key gets a process-unique generation, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read and `Set` files the fill under it, so a write landing mid-query orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) -- **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability when a selected backend keeps state there (`NeedsDataDir`); `backends.go` holds each layer's `.backend` (only the in-process value today; `coord.backend` reserved) — boot is the validator, there is no dry run +- **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability when a selected backend keeps state there (`NeedsDataDir`); `backends.go` holds each layer's `.backend` (the in-process value by default; `cache.backend` also takes `redis`, whose sub-block is `cache_redis.go`; `coord.backend` reserved) — boot is the validator, there is no dry run - **`dedupe/`** — `Deduplicator` interface → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) diff --git a/CHANGELOG.md b/CHANGELOG.md index 5bbe30535..ab628fdff 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,8 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A Redis-compatible shared cache backend, built but not yet selectable** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row open a circuit breaker that skips the server until a probe succeeds, and deferred invalidations are retried until they land, collapsing to one tenant-wide bump per tenant past 100,000 keys. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. Nothing selects it yet: the `cache.backend` switch and the `cache.redis.*` settings arrive with the wiring (E4), so every deployment still runs `LocalCache`. Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression and server-stops-answering cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. +- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone`, `cluster`, `sentinel`), `sentinel_master`, `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `-1` never compresses, since the loader reads a `0` in the file as unset) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. A malformed block — no address, an address without a port, an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. +- **A Redis-compatible shared cache backend** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row open a circuit breaker that skips the server until a probe succeeds, and deferred invalidations are retried until they land, collapsing to one tenant-wide bump per tenant past 100,000 keys. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. `cache.backend: redis` selects it (the entry above). Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression and server-stops-answering cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block; no backend has settings yet, so every such sub-block is an unknown key for now and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name the `mq`, `cache` and `dedupe` backends: the zero value is not the default, and `app.New` refuses it. `coord.backend` is reserved: nothing reads it until the lease layer lands, and the sweeper still runs in every process. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. diff --git a/README.md b/README.md index fce01a0a5..65f6da141 100644 --- a/README.md +++ b/README.md @@ -72,7 +72,7 @@ ClickHouse is a phenomenal OLAP database, but pointing a frontend right at it le If you're building user-facing analytics, WaveHouse is like **Supabase for ClickHouse**. Or an **open-source Tinybird** that pushes data to the frontend in real time over SSE, not just pull-based REST. - **Ingest** — async durable WAL (embedded NATS JetStream), `200 OK` instantly, background batch-flush; schema-validated against `system.columns`; optional ID-based dedup (idempotent ingest); dead-letter queue for failed inserts. -- **Query** — in-process Ristretto cache + `singleflight` coalescing; type-safe structured query AST; Tinybird-style named pipes (parameterized SQL endpoints). +- **Query** — result cache (in-process Ristretto, or a Redis shared by every instance) + `singleflight` coalescing; type-safe structured query AST; Tinybird-style named pipes (parameterized SQL endpoints). - **Real-time** — native SSE push, broadcast *before* the ClickHouse flush, with JetStream gap-fill for late/reconnecting clients. - **Security** — Hasura-style per-table, per-role column + row policies with JWT claim templating, defined in the hot-reloadable settings directory. - **Client** — `@wavehouse/sdk`: TypeScript client with query builder, live queries, streaming, and schema codegen; one runtime dependency (an SSE frame parser, ~1.4 KB gzipped). diff --git a/config.yaml b/config.yaml index 5519b78cd..aa9171e42 100644 --- a/config.yaml +++ b/config.yaml @@ -43,8 +43,8 @@ clickhouse: password: "" max_total_conns: 0 # ceiling on open native connections across pools; 0 = none -# Each layer's implementation, chosen at boot. Only the in-process backend -# exists for each today, and it is the default. +# Each layer's implementation, chosen at boot. The in-process backend is the +# default for each, and the only one for these three. mq: backend: embedded # NATS JetStream under /nats dedupe: @@ -52,11 +52,18 @@ dedupe: coord: backend: local # reserved: nothing is elected yet -# In-process L1 cache size. The query time-bucket +# The query-result cache: local (in-process, sized by l1_max_cost) or redis +# (one Redis-compatible server shared by every instance; see the redis block +# below and the Configuration page for every key). The query time-bucket # (query.timestamp_bucket_seconds) is a settings key. cache: backend: local l1_max_cost: 67108864 + # redis: # read only with backend: redis + # addrs: ["localhost:6379"] # `docker compose -f deployments/compose/dependencies.yaml --profile redis up -d` + # key_prefix: wh + # timeout: 100ms # per operation; slower is a miss, never a failed query + # The password is a secret: WH_CACHE_REDIS_PASSWORD, not this file. # Auth has no on/off switch — the JWT middleware always runs. A request with no # token, or an invalid/expired one, falls back to the policy default_role; diff --git a/deployments/compose/dependencies.yaml b/deployments/compose/dependencies.yaml index de147c726..4cc4f2fd6 100644 --- a/deployments/compose/dependencies.yaml +++ b/deployments/compose/dependencies.yaml @@ -33,5 +33,18 @@ services: timeout: 2s retries: 15 + # Optional shared cache for trying cache.backend=redis locally, e.g. two + # host-side instances on different ports: `--profile redis`. No + # persistence, and /data on tmpfs, so it leaves no volume behind. + redis: + profiles: [redis] + # Pinned to match internal/cache's integration suite. + image: redis:8.10.2-alpine + command: ["redis-server", "--save", "", "--appendonly", "no", "--maxmemory", "256mb", "--maxmemory-policy", "allkeys-lru"] + ports: + - "6379:6379" + tmpfs: + - /data + volumes: clickhouse-data: diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index 1634aaabc..c82739c07 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -560,7 +560,7 @@ The inbound request body is capped at 1 MiB; a body over the cap is rejected wit ### `GET/POST /v1/pipes/{name}` — Execute Named Pipe -Executes a pre-defined named query (pipe) with parameter binding. Parameters can be supplied via query string and/or JSON body. Results are cached in the shared L1 (Ristretto) with singleflight coalescing — same machinery as the structured query endpoint, keyed by [tenant](/deployment#multi-tenant-deployments) like it, and again, unlike `/v1/ops/query`. +Executes a pre-defined named query (pipe) with parameter binding. Parameters can be supplied via query string and/or JSON body. Results are cached in the query cache ([`cache.backend`](/configuration#backends): in-process, or a Redis shared by every instance) with singleflight coalescing — same machinery as the structured query endpoint, keyed by [tenant](/deployment#multi-tenant-deployments) like it, and again, unlike `/v1/ops/query`. **Query Parameters:** Any key matching a pipe parameter name. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 2a06d9fff..6c59e5f16 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -28,7 +28,7 @@ flowchart TD MQ --> BC["Buffer Consumer
(batch flush)"] BC -.->|failed inserts| DLQ["DLQ"]:::fail - QH["Query Handler"] --> Cache["Cache
(Ristretto + singleflight)"] + QH["Query Handler"] --> Cache["Cache
(local or Redis + singleflight)"] SSH["SSE Handler"] --> Hub["Stream Hub
(project once per role)"] @@ -90,7 +90,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability (after which each `Config.Warnings` line is logged at `WARN`), ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the whole cache — structured-query and pipe results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth, dedupe and cache hooks use too), ending the open streams of a tenant no longer served, and `wireCache`'s hook prunes the cache's version index the same way (`LocalCache.Prune`, [#262](https://github.com/Wave-RF/WaveHouse/issues/262)), so a tenant no longer served stops holding it. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. `wireCache` has two: `local`, the in-process `LocalCache`, and `redis`, the shared `RedisCache` built from the `cache.redis` block, which boots bypassed rather than failing when its server is unreachable. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the whole cache — structured-query and pipe results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth, dedupe and cache hooks use too), ending the open streams of a tenant no longer served, and `wireCache`'s hook prunes the cache's version index the same way (`LocalCache.Prune`, [#262](https://github.com/Wave-RF/WaveHouse/issues/262)), so a tenant no longer served stops holding it. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out @@ -112,7 +112,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on, each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope: one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)). `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. - **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: one version per tenant, per (tenant, table) and per (tenant, table, scope), each keyed by its name alone and bumped in place, so the index holds one entry per live tenant, table and scope however often each is bumped ([#262](https://github.com/Wave-RF/WaveHouse/issues/262)). A query key folds the tenant's version and, for each dependency, its tenant's, table's and scope's, so bumping a table (a scopeless write) orphans every scope of it, and bumping one scope orphans that scope and the whole-table view — scope is reserved and empty today, so every write is the whole-table bump — all without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results. A tenant's version is a *generation*, unique within the process and handed out by the first key built for the tenant; `BumpTenant` (behind `InvalidateTenant`) drops the tenant's whole index, so the next key gets a fresh generation no cached entry folds, orphaning every cached result of the tenant in one step — a pipe result with no dependencies, and a table no bump ever keyed, included — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). `LocalCache.Prune` does the same for every tenant no longer served, which `internal/app` runs after each settings reload, so a tenant removed or rejected stops holding its index. A table bump drops the table's scope versions with it, since every key they were folded into also folds the old table version; and a bump of a tenant with no index is a no-op, since no key folds its next generation yet. The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. -- **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET`, `MGET` and `PING`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token (a string of another length, a list, a hash) is replaced, so a token lost to eviction, expiry or a restart is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing. Nothing selects this backend yet — the boot config and wiring arrive with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s E4. +- **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET`, `MGET` and `PING`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token (a string of another length, a list, a hash) is replaced, so a token lost to eviction, expiry or a restart is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing. `cache.backend: redis` selects it: `internal/app`'s `wireCache` maps the boot config's `cache.redis` block onto `RedisConfig`, reading the TLS files, and releases it with the other components. - **redis_codec.go** — the key schema (`tokenKeys`, `bumpKeys`, `valueKey`) and the value frame: format, flags, expiry, the token list, the payload, zstd-compressed from `CompressMinBytes` when that is smaller. A stored value is capped at `MaxValueBytes` and a decoded one at eight times that, which refuses a zip bomb planted in a shared server; an unknown format, as a newer process writes during a rolling upgrade, is a miss. - **breaker.go** — the circuit breaker: `BreakerThreshold` (5) consecutive failures open it, and while open every operation skips the server at once; after `BreakerOpenFor` (5 s) one background `PING` decides whether it closes. Only a transport failure or a timeout counts against the server: any reply, an error reply or one the backend cannot use included, counts as a success, for operations and the probe alike, and a caller that gave up first counts as nothing. - **pending.go** — invalidations the server did not take, kept per token key (repeats coalesce) and retried by a background loop with backoff from 100 ms to 10 s until they land. Landing late is still correct: a fresh token orphans the pre-write entries and any fill made meanwhile. Past `PendingMax` (100,000) keys the set collapses to one tenant bump per affected tenant. `Close` makes one last attempt at the bumps still owed, past the breaker (an open one is why they are owed); what that attempt cannot deliver is lost, and the entries they would orphan are served until their TTL — the same failure as a worker stopping between an insert and its invalidation. @@ -120,9 +120,10 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ ### `config/` — Configuration -- **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — the implementation each layer runs on, resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). +- **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — the implementation each layer runs on, resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `cache.redis.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). - **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load` when `Config.NeedsDataDir` says a selected backend keeps state there, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. -- **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` at the end of `Validate`), which refuses a value not on the list and names the ones that are. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns the valid combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), which `app.New` logs at `WARN`. +- **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` at the end of `Validate`), which refuses a value not on the list and names the ones that are. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns what a valid configuration is still likely to get wrong — the combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), a `cache.redis` block that is not read, certificate verification turned off — which `app.New` logs at `WARN`. +- **cache_redis.go** — `CacheRedisConfig`, the `cache.redis` sub-block, and its checks: an address, each `host:port`, a known mode, a sentinel's master name, `db` 0 in cluster mode, positive timeouts and sizes, a `version_ttl` of at least 2 s, and `compress_min_bytes` positive or `-1` (never: cleanenv reads a `0` in the file as unset and applies the default, so `0` cannot mean off). `CacheRedisTLS.Config` builds the `tls.Config`, reading the files; `Validate` calls it so an unreadable file refuses boot, and `internal/app` calls it again to build the connection. A TLS key set while `tls.enabled` is off is an error rather than a plaintext connection. - **strict.go** — `rejectUnknownKeys`, the YAML half: re-reads the file as a generic tree and walks it against the struct's `yaml` tags, listing every key the struct doesn't declare. cleanenv itself is lenient by design, which is exactly wrong for boot config once keys have moved to the settings directory. - **persistence.go** — `WarnIfFreshDataDir` logs the startup `WARN` when `data_dir` is missing or empty (on a redeploy, the sign that the volume didn't persist); `LogStorageInitError` attaches the UID-65532 `permissionHint` to a NATS or Pebble open failure that looks like a permission denial — the same hint string `CheckDataDir` uses. @@ -366,7 +367,7 @@ Client GET /v1/stream | Analytics DB | ClickHouse | Primary data store + schema source of truth | | Message Queue | NATS + JetStream | Durable event streaming | | L1 Cache | Ristretto v2 | In-process memory cache | -| Shared cache | [rueidis](https://github.com/redis/rueidis) | Redis-compatible client for the shared backend (not yet selectable) | +| Shared cache | [rueidis](https://github.com/redis/rueidis) | Redis-compatible client for the shared backend (`cache.backend: redis`) | | Embedded KV | Pebble | Optional deduplication | | Config | cleanenv | YAML + env var config loading | | Release | GoReleaser | Cross-platform binary builds | diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 193a6c21b..82744a359 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -39,16 +39,16 @@ This page is boot config only — what the platform operator owns (wiring, lifec ### Backends -Each layer's implementation is chosen once, at boot. Today every layer has one backend, the in-process one, and it is the default, so a config that sets none of these keys runs as it always has. A value this build has no backend for refuses boot and names the valid ones. +Each layer's implementation is chosen once, at boot. Every layer defaults to its in-process backend, so a config that sets none of these keys runs as it always has; the cache also has a shared one, `redis`. A value this build has no backend for refuses boot and names the valid ones. | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | | `mq.backend` | `WH_MQ_BACKEND` | `embedded` | The message queue. `embedded`: NATS JetStream inside this process, under `/nats`. It listens on no port, so no other process can reach its queue. | -| `cache.backend` | `WH_CACHE_BACKEND` | `local` | The query-result cache. `local`: in this process, sized by `cache.l1_max_cost`. | +| `cache.backend` | `WH_CACHE_BACKEND` | `local` | The query-result cache. `local`: in this process, sized by `cache.l1_max_cost`. `redis`: one Redis-compatible server shared by every process, configured by [`cache.redis`](#cache), so an insert one process makes invalidates what every process has cached. | | `dedupe.backend` | `WH_DEDUPE_BACKEND` | `pebble` | Where ingest dedupe keeps the event ids it has seen. `pebble`: in this process, under `/pebble`, open while any tenant has dedupe on. | | `coord.backend` | `WH_COORD_BACKEND` | `local` | Reserved for the leases that will elect work only one process may do at a time, such as the sweeper. Nothing is elected yet: every process runs its own sweeper, and `local`, the only value, changes nothing. | -Settings for one backend will go in a sub-block named after it, `.`, read only when that backend is selected. No backend has settings yet, so today any such sub-block, `mq.embedded` included, is an unknown key and refuses boot. `mq` and `dedupe` also appear in the settings directory's `config.json`, with different keys (`mq.max_bytes_gb`, `dedupe.enabled`, …): those are per-tenant tunables and stay there, and one written in `config.yaml` refuses boot as an unknown key. +Settings for one backend go in a sub-block named after it, `.`, read only when that backend is selected. `cache.redis` is the only one so far; any other, `mq.embedded` included, is an unknown key and refuses boot. A `cache.redis.addrs` set while `cache.backend` is `local` is logged at `WARN` at boot, since the block is not read. `mq` and `dedupe` also appear in the settings directory's `config.json`, with different keys (`mq.max_bytes_gb`, `dedupe.enabled`, …): those are per-tenant tunables and stay there, and one written in `config.yaml` refuses boot as an unknown key. ### Server @@ -120,7 +120,34 @@ Each tenant's queue has its own disk budget, `mq.max_bytes_gb`, a hot-reloadable | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | -| `cache.l1_max_cost` | `WH_CACHE_L1_MAX_COST` | `67108864` | Maximum L1 cache size in bytes (~64 MB). The time-range bucket structured queries normalize to is `query.timestamp_bucket_seconds` in the [Settings Directory](/settings-directory#configjson-keys). | +| `cache.l1_max_cost` | `WH_CACHE_L1_MAX_COST` | `67108864` | Maximum size in bytes (~64 MB) of the `local` backend's in-process cache. The time-range bucket structured queries normalize to is `query.timestamp_bucket_seconds` in the [Settings Directory](/settings-directory#configjson-keys). | + +The `redis` backend's settings, read only when `cache.backend` is `redis`. It runs on Redis, Valkey, Dragonfly, ElastiCache (including Serverless) and MemoryDB: it sends only `GET`, `SET`, `MGET` and `PING`. [Deployment](/deployment#multiple-instances-and-the-shared-cache) covers sizing, `maxmemory-policy` and what a reader on another instance can see. + +| YAML Key | Env Var | Default | Description | +| --- | --- | ------- | ----------- | +| `cache.redis.addrs` | `WH_CACHE_REDIS_ADDRS` | — | **Required** with `backend: redis.` `host:port` of the server; several are a cluster's seed nodes or the sentinels. Comma-separated in the env var. | +| `cache.redis.mode` | `WH_CACHE_REDIS_MODE` | `standalone` | `standalone`, `cluster` or `sentinel`. | +| `cache.redis.sentinel_master` | `WH_CACHE_REDIS_SENTINEL_MASTER` | — | The master set name. Required with `mode: sentinel`. | +| `cache.redis.username` | `WH_CACHE_REDIS_USERNAME` | — | ACL user. Empty uses the server's `default` user. | +| `cache.redis.password` | `WH_CACHE_REDIS_PASSWORD` | — | A secret: set it through the environment (or your secret store's env injection), not in a tracked `config.yaml`. | +| `cache.redis.db` | `WH_CACHE_REDIS_DB` | `0` | Database number (`SELECT`). Standalone and sentinel only: a cluster has only database `0`, and any other value refuses boot. | +| `cache.redis.tls.enabled` | `WH_CACHE_REDIS_TLS_ENABLED` | `false` | Connect over TLS, verifying the server against the system roots or `ca_file`. Any other `tls` key set while this is off refuses boot, rather than connecting in plaintext. | +| `cache.redis.tls.ca_file` | `WH_CACHE_REDIS_TLS_CA_FILE` | — | PEM file of the authorities to trust instead of the system roots. | +| `cache.redis.tls.cert_file` | `WH_CACHE_REDIS_TLS_CERT_FILE` | — | Client certificate (PEM) for mutual TLS. Set together with `key_file`. | +| `cache.redis.tls.key_file` | `WH_CACHE_REDIS_TLS_KEY_FILE` | — | The client certificate's private key (PEM). | +| `cache.redis.tls.server_name` | `WH_CACHE_REDIS_TLS_SERVER_NAME` | — | Name to verify the server's certificate against, when it differs from the address. | +| `cache.redis.tls.insecure_skip_verify` | `WH_CACHE_REDIS_TLS_INSECURE_SKIP_VERIFY` | `false` | Accept any server certificate. Logged at `WARN` at boot: whoever can intercept the connection can read and replace cached results. | +| `cache.redis.key_prefix` | `WH_CACHE_REDIS_KEY_PREFIX` | `wh` | Leads every key, so several deployments can share one server. No `{` or `}`. | +| `cache.redis.timeout` | `WH_CACHE_REDIS_TIMEOUT` | `100ms` | Per operation. A lookup or fill that takes longer is a miss or a skipped fill, never a failed query. | +| `cache.redis.dial_timeout` | `WH_CACHE_REDIS_DIAL_TIMEOUT` | `1s` | Per connection attempt. | +| `cache.redis.max_value_bytes` | `WH_CACHE_REDIS_MAX_VALUE_BYTES` | `1048576` | Largest result stored, after compression (1 MiB). A larger one is returned to the caller but not cached. | +| `cache.redis.compress_min_bytes` | `WH_CACHE_REDIS_COMPRESS_MIN_BYTES` | `1024` | Results at least this large are zstd-compressed when that makes them smaller. `-1` never compresses. `0` is not "off": in the YAML file it reads as unset and takes the default, and in the env var it refuses boot. | +| `cache.redis.version_ttl` | `WH_CACHE_REDIS_VERSION_TTL` | `168h` | How long a table's or tenant's version token outlives its last write, so the tokens of dropped tables and removed tenants eventually expire. At least `2s`. An expired token only causes misses. | + +**When the server is unreachable or misbehaves, the cache is bypassed; queries are not.** A failure or a timeout makes the lookup a miss and the fill a no-op, and five in a row open a circuit breaker that skips the server entirely until a probe, every 5 s, gets an answer. Queries then go straight to ClickHouse, still coalesced per instance by `singleflight`. An invalidation the server did not take is kept and retried until it lands. `/readyz` does not depend on the cache. + +**Boot does not wait for the server.** A malformed block (an address without a port, `mode: cluster` with `db` other than `0`, an unreadable or unparsable TLS file) refuses boot. A server that cannot be reached, or that refuses the credentials, does not: the process boots with the cache bypassed and keeps reconnecting, with backoff up to 30 s. A rejected password (`WRONGPASS`, `NOAUTH`) is logged at `ERROR` on every attempt; any other failure at `WARN`. This is deliberate: a rotated Redis password must not crash-loop every instance at once. Watch `wavehouse_cache_breaker_open`, which reads `1` while the cache is bypassed, including before the first connection. ### Authentication @@ -210,8 +237,28 @@ mq: backend: embedded # in-process NATS JetStream under /nats cache: - backend: local - l1_max_cost: 67108864 + backend: local # local | redis + l1_max_cost: 67108864 # the local backend's size + redis: # read only with backend: redis + addrs: [] # required with backend: redis, e.g. ["redis:6379"] + mode: standalone # standalone | cluster | sentinel + sentinel_master: "" + username: "" + password: "" # a secret: set WH_CACHE_REDIS_PASSWORD instead + db: 0 + tls: + enabled: false + ca_file: "" + cert_file: "" + key_file: "" + server_name: "" + insecure_skip_verify: false + key_prefix: wh + timeout: 100ms + dial_timeout: 1s + max_value_bytes: 1048576 + compress_min_bytes: 1024 # -1 = never + version_ttl: 168h dedupe: backend: pebble # in-process Pebble under /pebble @@ -267,6 +314,25 @@ WH_CH_MAX_TOTAL_CONNS=0 WH_MQ_BACKEND=embedded WH_CACHE_BACKEND=local WH_CACHE_L1_MAX_COST=67108864 +# Read only with WH_CACHE_BACKEND=redis; WH_CACHE_REDIS_ADDRS is then required. +WH_CACHE_REDIS_ADDRS= +WH_CACHE_REDIS_MODE=standalone +WH_CACHE_REDIS_SENTINEL_MASTER= +WH_CACHE_REDIS_USERNAME= +WH_CACHE_REDIS_PASSWORD= +WH_CACHE_REDIS_DB=0 +WH_CACHE_REDIS_TLS_ENABLED=false +WH_CACHE_REDIS_TLS_CA_FILE= +WH_CACHE_REDIS_TLS_CERT_FILE= +WH_CACHE_REDIS_TLS_KEY_FILE= +WH_CACHE_REDIS_TLS_SERVER_NAME= +WH_CACHE_REDIS_TLS_INSECURE_SKIP_VERIFY=false +WH_CACHE_REDIS_KEY_PREFIX=wh +WH_CACHE_REDIS_TIMEOUT=100ms +WH_CACHE_REDIS_DIAL_TIMEOUT=1s +WH_CACHE_REDIS_MAX_VALUE_BYTES=1048576 +WH_CACHE_REDIS_COMPRESS_MIN_BYTES=1024 +WH_CACHE_REDIS_VERSION_TTL=168h WH_DEDUPE_BACKEND=pebble WH_COORD_BACKEND=local diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 79dae7d6a..63f1b1a9c 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -397,6 +397,26 @@ The folder name is the tenant id, and each folder is a complete settings directo `X-Tenant-ID` is a generic name, and some gateways and service meshes stamp one on every request. WaveHouse used to ignore it; now, over a settings directory that holds the four files, any value other than `0` names an unknown tenant, so **every `/v1` route outside `/v1/ops/*` answers `404 unknown tenant: `** (a `400` when the value is not a tenant id at all, a dotted hostname, say) — the SDK's `/v1/health` reachability ping included, while the bare probes and the admin surface stay green. Strip the inbound header at the edge ([header forwarding](/reverse-proxy#header-and-auth-forwarding)) unless you are using it deliberately. +## Multiple instances and the shared cache + +Several WaveHouse instances can serve one ClickHouse behind a load balancer, but most of what each one holds is its own. The message queue is embedded, so an event is inserted by the instance that took its `POST /v1/ingest`, and reaches only that instance's SSE subscribers. The dedupe store is per instance too, so an id one instance has seen is new to another. + +The query-result cache is the layer that can be shared today. With the default `cache.backend: local`, each instance caches in its own memory, and an insert invalidates only the cache of the instance that made it. Every other instance keeps serving its cached results for the rows before the insert until each entry's TTL runs out, between 10 s and 1 h depending on how long the query took. With [`cache.backend: redis`](/configuration#cache), every instance reads and fills one Redis-compatible server, and an insert on any instance invalidates the cached results of every instance. + +**What another instance can see.** Ingest is already asynchronous: `/v1/ingest` answers before the batch is inserted. Once the inserting instance's worker has written the batch to ClickHouse, it replaces the table's version token in Redis, and from then on a lookup on any instance misses and reads the new rows. The cache adds no delay of its own beyond that single write. The exceptions: + +- **The server is unreachable from the inserting instance.** The invalidation is kept and retried until it lands (`wavehouse_cache_invalidations_pending` counts what is owed). Meanwhile other instances that can still reach the server keep serving the older results, for as long as the outage lasts and at most until each entry's TTL. An instance that stops while invalidations are still owed loses them, with the same bound. The same thing happens today when a process stops between an insert and its invalidation. +- **A failover to a replica that had not yet received the latest token writes** can bring back entries filed under the older tokens, bounded by the replication lag at the moment of failover and those entries' TTL. WaveHouse never reads from replicas. +- **Admin writes through `POST /v1/ops/query`** do not invalidate the cache ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). With a shared cache, the stale results they leave are served by every instance, not only one. + +**Sizing the server.** Every key WaveHouse writes has a TTL, and a version token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses, never bring back an entry it had invalidated. So set `maxmemory` and let the server evict: `maxmemory-policy allkeys-lru` (or `allkeys-lfu`, `volatile-lru`, `volatile-lfu`). Under `noeviction`, a full server refuses the writes. Fills then fail (counted by `wavehouse_cache_set_failures_total{reason="oom"}`), lookups keep working, and invalidations are kept and retried. A stored result is capped at `cache.redis.max_value_bytes` (1 MiB compressed). A tenant's version tokens share one hash tag, so each lookup reads them in one `MGET` in cluster mode as well. The results themselves carry no hash tag and spread across shards. Persistence is not needed: an empty server after a restart is a cold cache, not a wrong one. + +**Coalescing stays per instance.** `singleflight` collapses identical concurrent queries within each instance, so a cold hot query costs at most one ClickHouse query per instance, not one per request. + +**Metrics** (meter `wavehouse-cache`, every series labeled `backend="redis"`, no tenant label): `wavehouse_cache_lookups_total{result}` (`hit`, `miss`, `stale`, `bypass`, `error`), `wavehouse_cache_op_duration_seconds{op}` (`lookup`, `set`, `invalidate`), `wavehouse_cache_breaker_open` (1 while the cache is bypassed), `wavehouse_cache_invalidations_total{result}` (`ok`, `deferred`), `wavehouse_cache_invalidations_pending`, `wavehouse_cache_value_bytes`, `wavehouse_cache_oversize_total` and `wavehouse_cache_set_failures_total{reason}` (`oom`, `timeout`, `other`). Two signals are worth alerting on: `wavehouse_cache_breaker_open` at 1, or `wavehouse_cache_invalidations_pending` above 0, for more than a few minutes. + +For local development, `docker compose -f deployments/compose/dependencies.yaml --profile redis up -d` starts a Redis on `localhost:6379` with no persistence. + ## ClickHouse Schema WaveHouse uses a **Bring Your Own Schema** model. You create your tables in ClickHouse with whatever columns and engines you need. WaveHouse discovers the schemas automatically via `system.columns` and validates ingest data against them — see [Schema Validation](/api#post-v1ingesttabletable--ingest-data) for the rules a record must satisfy. diff --git a/docs/src/content/docs/development.md b/docs/src/content/docs/development.md index 8e6a85036..2a9ea7d5a 100644 --- a/docs/src/content/docs/development.md +++ b/docs/src/content/docs/development.md @@ -82,7 +82,7 @@ make dev WaveHouse is now running at `http://localhost:8080` in standalone mode with: - **Embedded NATS** (JetStream) — no external MQ needed -- **L1 cache only** (Ristretto) — no external cache needed +- **In-process cache** (Ristretto, `cache.backend: local`) — no external cache needed; to try the shared one, start Redis with `docker compose -f deployments/compose/dependencies.yaml --profile redis up -d` and set `WH_CACHE_BACKEND=redis WH_CACHE_REDIS_ADDRS=localhost:6379` - **Trial policy** — the dev settings directory `./settings` is seeded on first run with the compose stack's permissive `public` policy, so tokenless requests to the demo tables work (see [Test the API](#test-the-api)) - **Dedup disabled** by default — no Pebble needed - **Schema discovery** — automatically finds your ClickHouse tables @@ -454,7 +454,7 @@ WaveHouse/ │ ├── api/ # HTTP handlers, router, middleware │ ├── app/ # Process wiring (build every component, run under one errgroup, release in reverse) │ ├── auth/ # JWT/JWKS authentication middleware -│ ├── cache/ # L1 (Ristretto) + L2 caching +│ ├── cache/ # Query cache: in-process (Ristretto) or shared (Redis-compatible) │ ├── chconn/ # ClickHouse pools, one per connection tuple (reconciled on settings reload) │ ├── chsql/ # Shared ClickHouse SQL helpers (quoting + bind-safety) │ ├── config/ # YAML + env var configuration diff --git a/docs/src/content/docs/getting-started.md b/docs/src/content/docs/getting-started.md index f24c3ee6f..667137d1a 100644 --- a/docs/src/content/docs/getting-started.md +++ b/docs/src/content/docs/getting-started.md @@ -75,7 +75,7 @@ curl -s -X POST "http://localhost:8080/v1/query?table=clicks" \ -d '{"columns": ["page", "button", "score"], "limit": 10}' ``` -`POST /v1/query?table={table}` and `GET/POST /v1/pipes/{name}` are cached in-process (L1 Ristretto) with singleflight coalescing — duplicate concurrent queries hit ClickHouse once. For raw SQL there's `POST /v1/ops/query` (an admin escape hatch that never caches, emitting `Cache-Control: no-store`), but it's **admin-only** — the trial `public` role can't reach it. To use it, swap the public default for real auth: configure a JWT secret and present a token whose role is the policy [`admin_role`](/access-control#admin_role--the-privileged-role). +`POST /v1/query?table={table}` and `GET/POST /v1/pipes/{name}` are cached — in-process by default, or in a Redis shared by every instance with [`cache.backend: redis`](/configuration#cache) — with singleflight coalescing, so duplicate concurrent queries hit ClickHouse once. For raw SQL there's `POST /v1/ops/query` (an admin escape hatch that never caches, emitting `Cache-Control: no-store`), but it's **admin-only** — the trial `public` role can't reach it. To use it, swap the public default for real auth: configure a JWT secret and present a token whose role is the policy [`admin_role`](/access-control#admin_role--the-privileged-role). :::tip[Prefer a type-safe client?] The [TypeScript SDK](/sdk) wraps this endpoint in a chainable query builder with autocomplete on your table names and row types — plus live queries and streaming. The raw shapes are in the [structured query reference](/api#post-v1querytabletable--structured-query). diff --git a/docs/src/content/docs/pipes.mdx b/docs/src/content/docs/pipes.mdx index b8c7efa04..b64129225 100644 --- a/docs/src/content/docs/pipes.mdx +++ b/docs/src/content/docs/pipes.mdx @@ -169,7 +169,7 @@ curl -X POST http://localhost:8080/v1/pipes/top_pages \ -d '{"start_date": "2024-01-01", "limit": 20}' ``` -The response is a JSON array of rows. Results flow through the shared in-process L1 cache (Ristretto) with singleflight coalescing, so concurrent identical calls hit ClickHouse once; an `X-Cache: HIT` or `X-Cache: MISS` header tells you which path served the response. +The response is a JSON array of rows. Results flow through the query cache (in-process by default, or a Redis shared by every instance with [`cache.backend: redis`](/configuration#cache)) with singleflight coalescing, so concurrent identical calls hit ClickHouse once; an `X-Cache: HIT` or `X-Cache: MISS` header tells you which path served the response. | Status | Body | Cause | | ------ | ---- | ----- | diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index 561354627..922f3a3f2 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -179,7 +179,7 @@ The tenant tunables. Every key is required (a missing one is a validation error) } ``` -What stays in boot config is only what cannot change under a running process — the implementation each layer runs on (`mq.backend`, `cache.backend`, `dedupe.backend`, `coord.backend`), resource sizing (`data_dir`, `cache.l1_max_cost`, `clickhouse.max_total_conns`), the listeners, the observability exporters — and the **secrets**: `clickhouse.password`, `auth.jwt_secret`, `auth.operator_key`. Secrets never belong in a tracked JSON file, so they stay in the environment and are combined with the wiring here on every (re)connect; rotating one is a restart. See [Configuration](/configuration). Everything else lives here and reloads. +What stays in boot config is only what cannot change under a running process — the implementation each layer runs on (`mq.backend`, `cache.backend`, `dedupe.backend`, `coord.backend`) and a shared backend's connection (`cache.redis`), resource sizing (`data_dir`, `cache.l1_max_cost`, `clickhouse.max_total_conns`), the listeners, the observability exporters — and the **secrets**: `clickhouse.password`, `cache.redis.password`, `auth.jwt_secret`, `auth.operator_key`. Secrets never belong in a tracked JSON file, so they stay in the environment and are combined with the wiring here on every (re)connect; rotating one is a restart. See [Configuration](/configuration). Everything else lives here and reloads. ## Deduplication diff --git a/internal/app/app_test.go b/internal/app/app_test.go index cf727b6d4..2ce9ca5bb 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -702,6 +702,77 @@ func TestReload_PrunesCacheIndexToServedTenants(t *testing.T) { assert.Equal(t, map[tenant.ID]bool{"acme": false, "globex": true}, rec.last(), "removed; the repaired one served again") } +// redisTestConfig is testConfig with cache.backend=redis at addr, carrying +// the defaults Load would apply. +func redisTestConfig(t *testing.T, settingsDir, addr string) *config.Config { + t.Helper() + cfg := testConfig(t, settingsDir) + cfg.Cache = config.Cache{Backend: config.CacheRedis, Redis: config.CacheRedisConfig{ + Addrs: []string{addr}, Mode: config.RedisStandalone, KeyPrefix: "wh", + Timeout: 100 * time.Millisecond, DialTimeout: 200 * time.Millisecond, + MaxValueBytes: 1 << 20, CompressMinBytes: 1 << 10, VersionTTL: time.Hour, + }} + require.NoError(t, cfg.Validate()) + return cfg +} + +// cache.backend=redis wires the shared backend. A server that cannot be +// reached does not refuse boot: the cache starts bypassed, and the reload +// hook that prunes an in-process index leaves it alone. +func TestNew_RedisCacheBootsBypassedWhenUnreachable(t *testing.T) { + root := writeNestedSettings(t, map[string]map[string]any{"acme": nil, "globex": nil}) + a := newApp(t, redisTestConfig(t, root, closedAddr(t)), Options{}) + _, ok := a.cache.(*cache.RedisCache) + require.True(t, ok, "cache is %T", a.cache) + assert.Contains(t, componentNames(a), "cache") + + require.NoError(t, os.RemoveAll(filepath.Join(root, "acme"))) + a.tenants.Reload("test") // the prune hook must not trip on a non-pruner + + entry, snap, err := a.cache.Lookup(t.Context(), "globex", "sha", nil) + require.NoError(t, err) + assert.Nil(t, entry.Value, "bypassed: a miss") + assert.NoError(t, a.cache.Set(t.Context(), snap, []byte("v"), time.Minute), "and the fill a no-op") +} + +// A TLS file that went missing between validation and wiring refuses boot, +// naming the key. +func TestNew_RedisCacheRefusesAnUnreadableTLSFile(t *testing.T) { + guardGlobals(t) + cfg := redisTestConfig(t, writeSettings(t, nil), closedAddr(t)) + cfg.Cache.Redis.TLS = config.CacheRedisTLS{Enabled: true, CAFile: filepath.Join(t.TempDir(), "gone.pem")} + _, err := New(t.Context(), Options{Config: cfg}) + require.ErrorContains(t, err, "cache init: cache.redis.tls.ca_file") +} + +// The boot config's defaults are the backend's, and -1 is the backend's +// "never compress". Driven from Load, not a literal, so a default changed on +// one side only fails here. +func TestRedisConfig_FromLoadedDefaults(t *testing.T) { + t.Setenv("WH_SETTINGS_DIR", t.TempDir()) + t.Setenv("WH_CACHE_BACKEND", "redis") + t.Setenv("WH_CACHE_REDIS_ADDRS", "a:6379,b:6379") + t.Setenv("WH_CACHE_REDIS_PASSWORD", "pw") + loaded, err := config.Load(filepath.Join(t.TempDir(), "none.yaml")) + require.NoError(t, err) + got, err := redisConfig(loaded.Cache.Redis) + require.NoError(t, err) + assert.Equal(t, cache.RedisConfig{ + Addrs: []string{"a:6379", "b:6379"}, Mode: cache.RedisStandalone, Password: "pw", + KeyPrefix: cache.DefaultRedisKeyPrefix, Timeout: cache.DefaultRedisTimeout, + DialTimeout: cache.DefaultRedisDialTimeout, MaxValueBytes: cache.DefaultRedisMaxValueBytes, + CompressMinBytes: cache.DefaultRedisCompressMinBytes, VersionTTL: cache.DefaultRedisVersionTTL, + }, got) + + loaded.Cache.Redis.CompressMinBytes = -1 + loaded.Cache.Redis.Mode = config.RedisCluster + got, err = redisConfig(loaded.Cache.Redis) + require.NoError(t, err) + assert.Zero(t, got.CompressMinBytes) + assert.Equal(t, cache.RedisCluster, got.Mode) + assert.Equal(t, cache.RedisSentinel, config.RedisSentinel) +} + // keepalive is a config.json patch setting the stream block's keepalive pair. func keepalive(interval, buckets int) map[string]any { return map[string]any{"stream": map[string]any{"keepalive_interval": interval, "keepalive_buckets": buckets, "gap_window_minutes": 15}} diff --git a/internal/app/wire.go b/internal/app/wire.go index 3cbe71a82..e1fe828d8 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -627,7 +627,7 @@ var _ pruner = (*cache.LocalCache)(nil) // is chosen. After every reload a tenant no longer served, removed or // rejected alike, has its in-process version index dropped (#262); its cache // is orphaned with it, as it would be anyway when it came back -// (wireClickHouse). +// (wireClickHouse). A shared backend keeps no such index and is skipped. func (a *App) wireCache() error { var c cache.Cache switch b := a.cfg.Cache.Backend; b { @@ -637,6 +637,16 @@ func (a *App) wireCache() error { return fmt.Errorf("cache init: %w", err) } c = l1 + case config.CacheRedis: + rc, err := redisConfig(a.cfg.Cache.Redis) + if err != nil { + return fmt.Errorf("cache init: %w", err) + } + r, err := cache.NewRedis(rc) + if err != nil { + return fmt.Errorf("cache init: %w", err) + } + c = r default: return unreachableBackend("cache.backend", b) } @@ -650,6 +660,35 @@ func (a *App) wireCache() error { return nil } +// redisConfig maps the boot config's cache.redis block onto the backend's +// config. Load has applied every default and validated the block; the TLS +// files are read again here, so the connection uses what is on disk now. +func redisConfig(r config.CacheRedisConfig) (cache.RedisConfig, error) { + t, err := r.TLS.Config() + if err != nil { + return cache.RedisConfig{}, err + } + compressMin := r.CompressMinBytes + if compressMin < 0 { + compressMin = 0 // the backend's "never" + } + return cache.RedisConfig{ + Addrs: r.Addrs, + Mode: r.Mode, + SentinelMaster: r.SentinelMaster, + Username: r.Username, + Password: r.Password, + DB: r.DB, + TLS: t, + KeyPrefix: r.KeyPrefix, + Timeout: r.Timeout, + DialTimeout: r.DialTimeout, + MaxValueBytes: r.MaxValueBytes, + CompressMinBytes: compressMin, + VersionTTL: r.VersionTTL, + }, nil +} + // unreachableBackend is each layer switch's default case. config.Validate // refuses a backend with no case, so reaching it means a Config built by hand // without one (the zero value is not the default), or a case missing here. diff --git a/internal/config/backends.go b/internal/config/backends.go index f2ab93308..893d2304b 100644 --- a/internal/config/backends.go +++ b/internal/config/backends.go @@ -35,21 +35,34 @@ func (m MQ) validate() error { // CacheBackend names the query-result cache implementation. type CacheBackend string -// CacheLocal is the in-process Ristretto cache, sized by cache.l1_max_cost. -const CacheLocal CacheBackend = "local" +const ( + // CacheLocal is the in-process Ristretto cache, sized by + // cache.l1_max_cost. + CacheLocal CacheBackend = "local" + // CacheRedis is one Redis-compatible server shared by every process, + // configured by cache.redis. + CacheRedis CacheBackend = "redis" +) -var cacheBackends = []CacheBackend{CacheLocal} +var cacheBackends = []CacheBackend{CacheLocal, CacheRedis} // Cache selects and sizes the query-result cache. The time-range bucket // structured queries normalize to is a settings-directory key // (query.timestamp_bucket_seconds) — query shaping, not process memory. type Cache struct { - Backend CacheBackend `yaml:"backend" env:"WH_CACHE_BACKEND" env-default:"local"` - L1MaxCost int64 `yaml:"l1_max_cost" env:"WH_CACHE_L1_MAX_COST" env-default:"67108864"` + Backend CacheBackend `yaml:"backend" env:"WH_CACHE_BACKEND" env-default:"local"` + L1MaxCost int64 `yaml:"l1_max_cost" env:"WH_CACHE_L1_MAX_COST" env-default:"67108864"` + Redis CacheRedisConfig `yaml:"redis"` } func (c Cache) validate() error { - return checkBackend("cache.backend", "WH_CACHE_BACKEND", c.Backend, cacheBackends) + if err := checkBackend("cache.backend", "WH_CACHE_BACKEND", c.Backend, cacheBackends); err != nil { + return err + } + if c.Backend == CacheRedis { + return c.Redis.validate() + } + return nil } // DedupeBackend names where ingest dedupe keeps the ids it has seen. @@ -126,13 +139,20 @@ func (c *Config) NeedsDataDir() bool { } // Warnings returns what a valid configuration is still likely to get wrong, -// one line each, for boot to log at WARN. They are not errors because each is -// correct for a single replica, and one process cannot count its replicas. +// one line each, for boot to log at WARN. The shared-queue ones are not +// errors because each is correct for a single replica, and one process +// cannot count its replicas. func (c *Config) Warnings() []string { + var out []string + if c.Cache.Backend == CacheRedis && c.Cache.Redis.TLS.InsecureSkipVerify { + out = append(out, "cache.redis.tls.insecure_skip_verify is on: the cache accepts any certificate, so whoever can intercept the connection can read and replace cached query results") + } + if c.Cache.Backend != CacheRedis && c.Cache.Redis.hasAddrs() { + out = append(out, "cache.redis.addrs is set but cache.backend is "+string(c.Cache.Backend)+": the redis block is not read; set cache.backend=redis to share the cache") + } if !c.Distributed() { - return nil + return out } - var out []string if c.Cache.Backend == CacheLocal { out = append(out, "cache.backend=local with a shared mq.backend is correct for one replica only: an event ingested on another replica never invalidates this one's cache, so its reads stay stale until the cached entry expires") } diff --git a/internal/config/backends_test.go b/internal/config/backends_test.go index 0e70dc7a3..778336971 100644 --- a/internal/config/backends_test.go +++ b/internal/config/backends_test.go @@ -111,7 +111,7 @@ func TestValidate_UnknownBackend(t *testing.T) { want string }{ {"mq", func(c *Config) { c.MQ.Backend = "kafka" }, `mq.backend (WH_MQ_BACKEND) "kafka" is not a backend this build has; valid: embedded`}, - {"cache", func(c *Config) { c.Cache.Backend = "redis" }, `cache.backend (WH_CACHE_BACKEND) "redis" is not a backend this build has; valid: local`}, + {"cache", func(c *Config) { c.Cache.Backend = "memcached" }, `cache.backend (WH_CACHE_BACKEND) "memcached" is not a backend this build has; valid: local, redis`}, {"dedupe", func(c *Config) { c.Dedupe.Backend = "dynamodb" }, `dedupe.backend (WH_DEDUPE_BACKEND) "dynamodb" is not a backend this build has; valid: pebble`}, {"coord", func(c *Config) { c.Coord.Backend = "nats" }, `coord.backend (WH_COORD_BACKEND) "nats" is not a backend this build has; valid: local`}, // The zero value, which a Config built without Load carries. diff --git a/internal/config/cache_redis.go b/internal/config/cache_redis.go new file mode 100644 index 000000000..f0af67137 --- /dev/null +++ b/internal/config/cache_redis.go @@ -0,0 +1,153 @@ +package config + +import ( + "crypto/tls" + "crypto/x509" + "errors" + "fmt" + "net" + "os" + "strings" + "time" +) + +// Redis deployment modes for cache.redis.mode. +const ( + RedisStandalone = "standalone" + RedisCluster = "cluster" + RedisSentinel = "sentinel" +) + +// CacheRedisConfig configures cache.backend=redis: one Redis-compatible server +// (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB) shared by every process. +// Read only when that backend is selected. +type CacheRedisConfig struct { + // Addrs are host:port pairs: the server, or seeds for a cluster, or the + // sentinels. + Addrs []string `yaml:"addrs" env:"WH_CACHE_REDIS_ADDRS"` + Mode string `yaml:"mode" env:"WH_CACHE_REDIS_MODE" env-default:"standalone"` + SentinelMaster string `yaml:"sentinel_master" env:"WH_CACHE_REDIS_SENTINEL_MASTER"` + Username string `yaml:"username" env:"WH_CACHE_REDIS_USERNAME"` + Password string `yaml:"password" env:"WH_CACHE_REDIS_PASSWORD"` + DB int `yaml:"db" env:"WH_CACHE_REDIS_DB"` + TLS CacheRedisTLS `yaml:"tls"` + // KeyPrefix leads every key, so deployments can share one server. + KeyPrefix string `yaml:"key_prefix" env:"WH_CACHE_REDIS_KEY_PREFIX" env-default:"wh"` + Timeout time.Duration `yaml:"timeout" env:"WH_CACHE_REDIS_TIMEOUT" env-default:"100ms"` + DialTimeout time.Duration `yaml:"dial_timeout" env:"WH_CACHE_REDIS_DIAL_TIMEOUT" env-default:"1s"` + // MaxValueBytes is the largest value stored, after compression. + MaxValueBytes int `yaml:"max_value_bytes" env:"WH_CACHE_REDIS_MAX_VALUE_BYTES" env-default:"1048576"` + // CompressMinBytes is the smallest value zstd-compressed; -1 never + // compresses. Not 0: the loader reads a 0 in the file as unset and + // applies the default, so 0 cannot mean off. + CompressMinBytes int `yaml:"compress_min_bytes" env:"WH_CACHE_REDIS_COMPRESS_MIN_BYTES" env-default:"1024"` + // VersionTTL is how long a version token outlives its last bump. + VersionTTL time.Duration `yaml:"version_ttl" env:"WH_CACHE_REDIS_VERSION_TTL" env-default:"168h"` +} + +// CacheRedisTLS is cache.redis.tls. The files are paths, read at boot. +type CacheRedisTLS struct { + Enabled bool `yaml:"enabled" env:"WH_CACHE_REDIS_TLS_ENABLED"` + CAFile string `yaml:"ca_file" env:"WH_CACHE_REDIS_TLS_CA_FILE"` + CertFile string `yaml:"cert_file" env:"WH_CACHE_REDIS_TLS_CERT_FILE"` + KeyFile string `yaml:"key_file" env:"WH_CACHE_REDIS_TLS_KEY_FILE"` + ServerName string `yaml:"server_name" env:"WH_CACHE_REDIS_TLS_SERVER_NAME"` + InsecureSkipVerify bool `yaml:"insecure_skip_verify" env:"WH_CACHE_REDIS_TLS_INSECURE_SKIP_VERIFY"` +} + +// hasAddrs reports whether any address is set. An env file's blank +// `WH_CACHE_REDIS_ADDRS=` loads as one empty address, which is none. +func (r CacheRedisConfig) hasAddrs() bool { + return len(r.Addrs) > 1 || len(r.Addrs) == 1 && r.Addrs[0] != "" +} + +func (r CacheRedisConfig) validate() error { + if !r.hasAddrs() { + return errors.New("cache.backend=redis needs cache.redis.addrs (WH_CACHE_REDIS_ADDRS): the server's host:port, or a cluster's seeds, or the sentinels") + } + for _, a := range r.Addrs { + if _, _, err := net.SplitHostPort(a); err != nil { + return fmt.Errorf("cache.redis.addrs (WH_CACHE_REDIS_ADDRS) %q: want host:port: %w", a, err) + } + } + switch r.Mode { + case RedisStandalone, RedisCluster, RedisSentinel: + default: + return fmt.Errorf("cache.redis.mode (WH_CACHE_REDIS_MODE) %q: valid: %s, %s, %s", r.Mode, RedisStandalone, RedisCluster, RedisSentinel) + } + if r.Mode == RedisSentinel && r.SentinelMaster == "" { + return errors.New("cache.redis.mode=sentinel needs cache.redis.sentinel_master (WH_CACHE_REDIS_SENTINEL_MASTER), the master set name") + } + if r.DB < 0 { + return fmt.Errorf("cache.redis.db (WH_CACHE_REDIS_DB) %d is negative", r.DB) + } + if r.Mode == RedisCluster && r.DB != 0 { + return fmt.Errorf("cache.redis.db (WH_CACHE_REDIS_DB) %d: a Redis cluster has only database 0", r.DB) + } + if r.KeyPrefix == "" || strings.ContainsAny(r.KeyPrefix, "{}") { + return fmt.Errorf("cache.redis.key_prefix (WH_CACHE_REDIS_KEY_PREFIX) %q: want a non-empty prefix without a hash-tag brace", r.KeyPrefix) + } + for _, d := range []struct { + key string + v time.Duration + }{ + {"cache.redis.timeout (WH_CACHE_REDIS_TIMEOUT)", r.Timeout}, + {"cache.redis.dial_timeout (WH_CACHE_REDIS_DIAL_TIMEOUT)", r.DialTimeout}, + } { + if d.v <= 0 { + return fmt.Errorf("%s %s must be positive", d.key, d.v) + } + } + if r.VersionTTL < 2*time.Second { + return fmt.Errorf("cache.redis.version_ttl (WH_CACHE_REDIS_VERSION_TTL) %s is under 2s", r.VersionTTL) + } + if r.MaxValueBytes <= 0 { + return fmt.Errorf("cache.redis.max_value_bytes (WH_CACHE_REDIS_MAX_VALUE_BYTES) %d must be positive", r.MaxValueBytes) + } + if r.CompressMinBytes == 0 || r.CompressMinBytes < -1 { + return fmt.Errorf("cache.redis.compress_min_bytes (WH_CACHE_REDIS_COMPRESS_MIN_BYTES) %d: want a positive size, or -1 to never compress", r.CompressMinBytes) + } + if _, err := r.TLS.Config(); err != nil { + return err + } + return nil +} + +// Config builds the tls.Config the block describes, reading its files, or +// nil when TLS is off. A file set while TLS is off is an error rather than +// a silently plaintext connection. +func (t CacheRedisTLS) Config() (*tls.Config, error) { + if !t.Enabled { + if t != (CacheRedisTLS{}) { + return nil, errors.New("cache.redis.tls: files, server_name or insecure_skip_verify are set but cache.redis.tls.enabled (WH_CACHE_REDIS_TLS_ENABLED) is off") + } + return nil, nil + } + if (t.CertFile == "") != (t.KeyFile == "") { + return nil, errors.New("cache.redis.tls: cert_file and key_file must be set together") + } + cfg := &tls.Config{ + MinVersion: tls.VersionTLS12, + ServerName: t.ServerName, + InsecureSkipVerify: t.InsecureSkipVerify, //nolint:gosec // G402: the operator's cache.redis.tls.insecure_skip_verify, warned about at boot + } + if t.CAFile != "" { + pemBytes, err := os.ReadFile(t.CAFile) + if err != nil { + return nil, fmt.Errorf("cache.redis.tls.ca_file: %w", err) + } + pool := x509.NewCertPool() + if !pool.AppendCertsFromPEM(pemBytes) { + return nil, fmt.Errorf("cache.redis.tls.ca_file: no certificates in %s", t.CAFile) + } + cfg.RootCAs = pool + } + if t.CertFile != "" { + cert, err := tls.LoadX509KeyPair(t.CertFile, t.KeyFile) + if err != nil { + return nil, fmt.Errorf("cache.redis.tls.cert_file: %w", err) + } + cfg.Certificates = []tls.Certificate{cert} + } + return cfg, nil +} diff --git a/internal/config/cache_redis_test.go b/internal/config/cache_redis_test.go new file mode 100644 index 000000000..6d989c737 --- /dev/null +++ b/internal/config/cache_redis_test.go @@ -0,0 +1,298 @@ +package config + +import ( + "crypto/ecdsa" + "crypto/elliptic" + "crypto/rand" + "crypto/x509" + "crypto/x509/pkix" + "encoding/pem" + "math/big" + "os" + "path/filepath" + "testing" + "time" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" +) + +// redisBackend is what Load produces for cache.backend=redis with only the +// address set. +func redisBackend() Config { + c := defaultBackends() + c.Cache.Backend = CacheRedis + c.Cache.Redis = CacheRedisConfig{ + Addrs: []string{"redis:6379"}, Mode: RedisStandalone, KeyPrefix: "wh", + Timeout: 100 * time.Millisecond, DialTimeout: time.Second, + MaxValueBytes: 1 << 20, CompressMinBytes: 1 << 10, VersionTTL: 168 * time.Hour, + } + return c +} + +func TestLoad_CacheRedisDefaults(t *testing.T) { + t.Setenv("WH_CACHE_BACKEND", "redis") + t.Setenv("WH_CACHE_REDIS_ADDRS", "redis:6379") + cfg, err := Load("nonexistent.yaml") + require.NoError(t, err) + want := redisBackend() + assert.Equal(t, want.Cache.Redis, cfg.Cache.Redis) + assert.Empty(t, cfg.Warnings()) +} + +func TestLoad_CacheRedisFromEnv(t *testing.T) { + dir := t.TempDir() + caFile, certFile, keyFile := writeTestPKI(t, dir) + for k, v := range map[string]string{ + "WH_CACHE_BACKEND": "redis", + "WH_CACHE_REDIS_ADDRS": "s1:26379,s2:26379", + "WH_CACHE_REDIS_MODE": "sentinel", + "WH_CACHE_REDIS_SENTINEL_MASTER": "mymaster", + "WH_CACHE_REDIS_USERNAME": "wavehouse", + "WH_CACHE_REDIS_PASSWORD": "s3cret", + "WH_CACHE_REDIS_DB": "2", + "WH_CACHE_REDIS_TLS_ENABLED": "true", + "WH_CACHE_REDIS_TLS_CA_FILE": caFile, + "WH_CACHE_REDIS_TLS_CERT_FILE": certFile, + "WH_CACHE_REDIS_TLS_KEY_FILE": keyFile, + "WH_CACHE_REDIS_TLS_SERVER_NAME": "redis.internal", + "WH_CACHE_REDIS_KEY_PREFIX": "staging", + "WH_CACHE_REDIS_TIMEOUT": "250ms", + "WH_CACHE_REDIS_DIAL_TIMEOUT": "3s", + "WH_CACHE_REDIS_MAX_VALUE_BYTES": "2048", + "WH_CACHE_REDIS_COMPRESS_MIN_BYTES": "-1", + "WH_CACHE_REDIS_VERSION_TTL": "24h", + "WH_CACHE_REDIS_TLS_INSECURE_SKIP_VERIFY": "false", + } { + t.Setenv(k, v) + } + cfg, err := Load("nonexistent.yaml") + require.NoError(t, err) + assert.Equal(t, CacheRedisConfig{ + Addrs: []string{"s1:26379", "s2:26379"}, Mode: RedisSentinel, SentinelMaster: "mymaster", + Username: "wavehouse", Password: "s3cret", DB: 2, + TLS: CacheRedisTLS{ + Enabled: true, CAFile: caFile, CertFile: certFile, KeyFile: keyFile, ServerName: "redis.internal", + }, + KeyPrefix: "staging", Timeout: 250 * time.Millisecond, DialTimeout: 3 * time.Second, + MaxValueBytes: 2048, CompressMinBytes: -1, VersionTTL: 24 * time.Hour, + }, cfg.Cache.Redis) + tc, err := cfg.Cache.Redis.TLS.Config() + require.NoError(t, err) + assert.Equal(t, "redis.internal", tc.ServerName) + assert.NotNil(t, tc.RootCAs) + assert.Len(t, tc.Certificates, 1) +} + +func TestLoad_CacheRedisFromYAML(t *testing.T) { + t.Parallel() + path := filepath.Join(t.TempDir(), "config.yaml") + require.NoError(t, os.WriteFile(path, []byte(` +settings: + dir: ./settings +cache: + backend: redis + redis: + addrs: ["n1:6379", "n2:6379"] + mode: cluster + key_prefix: prod + timeout: 50ms + version_ttl: 72h +`), 0o600)) + cfg, err := Load(path) + require.NoError(t, err) + r := cfg.Cache.Redis + assert.Equal(t, CacheRedis, cfg.Cache.Backend) + assert.Equal(t, []string{"n1:6379", "n2:6379"}, r.Addrs) + assert.Equal(t, RedisCluster, r.Mode) + assert.Equal(t, "prod", r.KeyPrefix) + assert.Equal(t, 50*time.Millisecond, r.Timeout) + assert.Equal(t, 72*time.Hour, r.VersionTTL) + assert.Equal(t, time.Second, r.DialTimeout, "an unset key takes its default") + assert.Equal(t, 1024, r.CompressMinBytes) +} + +// A 0 in the file is read as unset, so it takes the default instead of +// switching compression off: why "never" is -1. Pinned so a loader that +// starts honoring the 0 is noticed. +func TestLoad_CacheRedisCompressZeroInYAMLIsTheDefault(t *testing.T) { + t.Parallel() + path := filepath.Join(t.TempDir(), "config.yaml") + require.NoError(t, os.WriteFile(path, []byte(` +settings: + dir: ./settings +cache: + backend: redis + redis: + addrs: ["r:6379"] + compress_min_bytes: 0 +`), 0o600)) + cfg, err := Load(path) + require.NoError(t, err) + assert.Equal(t, 1024, cfg.Cache.Redis.CompressMinBytes) +} + +func TestLoad_CacheRedisRefusesUnknownKeys(t *testing.T) { + t.Parallel() + path := filepath.Join(t.TempDir(), "config.yaml") + require.NoError(t, os.WriteFile(path, []byte(` +settings: + dir: ./settings +cache: + backend: redis + redis: + addr: r:6379 + near_cache: + max_cost: 1 + tls: + ca: /x + memcached: + addrs: ["m:11211"] +`), 0o600)) + _, err := Load(path) + require.Error(t, err) + assert.Contains(t, err.Error(), "cache.memcached, cache.redis.addr, cache.redis.near_cache, cache.redis.tls.ca") +} + +// The documented env file lists WH_CACHE_REDIS_ADDRS blank: that is no +// address, not an unread block to warn about, and not a valid redis one. +func TestLoad_CacheRedisBlankAddrs(t *testing.T) { + t.Setenv("WH_CACHE_REDIS_ADDRS", "") + cfg, err := Load("nonexistent.yaml") + require.NoError(t, err) + assert.Empty(t, cfg.Warnings()) + + t.Setenv("WH_CACHE_BACKEND", "redis") + _, err = Load("nonexistent.yaml") + require.ErrorContains(t, err, "cache.backend=redis needs cache.redis.addrs") +} + +func TestUnboundEnv_KnowsTheCacheRedisVariables(t *testing.T) { + t.Parallel() + assert.Empty(t, unboundEnv([]string{ + "WH_CACHE_REDIS_ADDRS=r:6379", "WH_CACHE_REDIS_PASSWORD=x", "WH_CACHE_REDIS_TLS_CA_FILE=/ca.pem", + "WH_CACHE_REDIS_VERSION_TTL=1h", "WH_CACHE_REDIS_COMPRESS_MIN_BYTES=-1", + })) + assert.Equal(t, []string{"WH_CACHE_REDIS_ADDR"}, unboundEnv([]string{"WH_CACHE_REDIS_ADDR=r:6379"})) +} + +func TestValidate_CacheRedis(t *testing.T) { + t.Parallel() + dir := t.TempDir() + caFile, certFile, keyFile := writeTestPKI(t, dir) + notPEM := filepath.Join(dir, "not.pem") + require.NoError(t, os.WriteFile(notPEM, []byte("hello"), 0o600)) + cases := []struct { + name string + set func(*CacheRedisConfig) + want string // "" = valid + }{ + {"defaults", func(*CacheRedisConfig) {}, ""}, + {"no addrs", func(r *CacheRedisConfig) { r.Addrs = nil }, "cache.backend=redis needs cache.redis.addrs (WH_CACHE_REDIS_ADDRS)"}, + {"addr without port", func(r *CacheRedisConfig) { r.Addrs = []string{"redis"} }, `cache.redis.addrs (WH_CACHE_REDIS_ADDRS) "redis": want host:port`}, + {"mode", func(r *CacheRedisConfig) { r.Mode = "replica" }, `cache.redis.mode (WH_CACHE_REDIS_MODE) "replica": valid: standalone, cluster, sentinel`}, + {"sentinel without master", func(r *CacheRedisConfig) { r.Mode = RedisSentinel }, "needs cache.redis.sentinel_master"}, + {"sentinel", func(r *CacheRedisConfig) { r.Mode, r.SentinelMaster = RedisSentinel, "m" }, ""}, + {"cluster db", func(r *CacheRedisConfig) { r.Mode, r.DB = RedisCluster, 1 }, "a Redis cluster has only database 0"}, + {"standalone db", func(r *CacheRedisConfig) { r.DB = 3 }, ""}, + {"negative db", func(r *CacheRedisConfig) { r.DB = -1 }, "is negative"}, + {"empty prefix", func(r *CacheRedisConfig) { r.KeyPrefix = "" }, "cache.redis.key_prefix"}, + {"brace prefix", func(r *CacheRedisConfig) { r.KeyPrefix = "a{b}" }, "hash-tag brace"}, + {"zero timeout", func(r *CacheRedisConfig) { r.Timeout = 0 }, "cache.redis.timeout (WH_CACHE_REDIS_TIMEOUT) 0s must be positive"}, + {"negative dial timeout", func(r *CacheRedisConfig) { r.DialTimeout = -time.Second }, "cache.redis.dial_timeout"}, + {"short version ttl", func(r *CacheRedisConfig) { r.VersionTTL = time.Second }, "cache.redis.version_ttl (WH_CACHE_REDIS_VERSION_TTL) 1s is under 2s"}, + {"zero max value", func(r *CacheRedisConfig) { r.MaxValueBytes = 0 }, "cache.redis.max_value_bytes"}, + {"compress 0", func(r *CacheRedisConfig) { r.CompressMinBytes = 0 }, "or -1 to never compress"}, + {"compress -2", func(r *CacheRedisConfig) { r.CompressMinBytes = -2 }, "or -1 to never compress"}, + {"compress never", func(r *CacheRedisConfig) { r.CompressMinBytes = -1 }, ""}, + {"tls files while off", func(r *CacheRedisConfig) { r.TLS.CAFile = caFile }, "cache.redis.tls.enabled (WH_CACHE_REDIS_TLS_ENABLED) is off"}, + {"tls system roots", func(r *CacheRedisConfig) { r.TLS.Enabled = true }, ""}, + {"tls full", func(r *CacheRedisConfig) { + r.TLS = CacheRedisTLS{Enabled: true, CAFile: caFile, CertFile: certFile, KeyFile: keyFile} + }, ""}, + {"tls cert without key", func(r *CacheRedisConfig) { r.TLS = CacheRedisTLS{Enabled: true, CertFile: certFile} }, "cert_file and key_file must be set together"}, + {"tls missing ca", func(r *CacheRedisConfig) { + r.TLS = CacheRedisTLS{Enabled: true, CAFile: filepath.Join(dir, "missing.pem")} + }, "cache.redis.tls.ca_file"}, + {"tls ca not pem", func(r *CacheRedisConfig) { r.TLS = CacheRedisTLS{Enabled: true, CAFile: notPEM} }, "no certificates in"}, + {"tls bad pair", func(r *CacheRedisConfig) { + r.TLS = CacheRedisTLS{Enabled: true, CertFile: certFile, KeyFile: notPEM} + }, "cache.redis.tls.cert_file"}, + } + for _, tc := range cases { + t.Run(tc.name, func(t *testing.T) { + t.Parallel() + cfg := redisBackend() + tc.set(&cfg.Cache.Redis) + err := cfg.Validate() + if tc.want == "" { + require.NoError(t, err) + return + } + require.Error(t, err) + assert.Contains(t, err.Error(), tc.want) + }) + } +} + +// The redis block is read only when selected: an invalid one under +// backend=local does not refuse boot, it warns that it is ignored. +func TestValidate_CacheRedisIgnoredUnlessSelected(t *testing.T) { + t.Parallel() + cfg := defaultBackends() + cfg.Cache.Redis.Addrs = []string{"no-port"} + require.NoError(t, cfg.Validate()) + got := cfg.Warnings() + require.Len(t, got, 1) + assert.Contains(t, got[0], "cache.redis.addrs is set but cache.backend is local") +} + +func TestWarnings_CacheRedis(t *testing.T) { + t.Parallel() + cfg := redisBackend() + assert.Empty(t, cfg.Warnings()) + cfg.Cache.Redis.TLS = CacheRedisTLS{Enabled: true, InsecureSkipVerify: true} + require.NoError(t, cfg.Validate()) + got := cfg.Warnings() + require.Len(t, got, 1) + assert.Contains(t, got[0], "cache.redis.tls.insecure_skip_verify is on") + + // A shared cache clears the shared-queue warning about a local one. + cfg = redisBackend() + cfg.MQ.Backend = "shared" + got = cfg.Warnings() + require.Len(t, got, 1) + assert.Contains(t, got[0], "dedupe.backend=pebble") +} + +// writeTestPKI writes a self-signed authority and a client certificate it +// signed, returning the three paths cache.redis.tls names. +func writeTestPKI(t *testing.T, dir string) (caFile, certFile, keyFile string) { + t.Helper() + caKey, err := ecdsa.GenerateKey(elliptic.P256(), rand.Reader) + require.NoError(t, err) + ca := &x509.Certificate{ + SerialNumber: big.NewInt(1), Subject: pkix.Name{CommonName: "test ca"}, + NotBefore: time.Now().Add(-time.Hour), NotAfter: time.Now().Add(time.Hour), + IsCA: true, BasicConstraintsValid: true, KeyUsage: x509.KeyUsageCertSign, + } + caDER, err := x509.CreateCertificate(rand.Reader, ca, ca, &caKey.PublicKey, caKey) + require.NoError(t, err) + leafKey, err := ecdsa.GenerateKey(elliptic.P256(), rand.Reader) + require.NoError(t, err) + leaf := &x509.Certificate{ + SerialNumber: big.NewInt(2), Subject: pkix.Name{CommonName: "wavehouse"}, + NotBefore: time.Now().Add(-time.Hour), NotAfter: time.Now().Add(time.Hour), + KeyUsage: x509.KeyUsageDigitalSignature, ExtKeyUsage: []x509.ExtKeyUsage{x509.ExtKeyUsageClientAuth}, + } + leafDER, err := x509.CreateCertificate(rand.Reader, leaf, ca, &leafKey.PublicKey, caKey) + require.NoError(t, err) + keyDER, err := x509.MarshalECPrivateKey(leafKey) + require.NoError(t, err) + write := func(name, typ string, der []byte) string { + path := filepath.Join(dir, name) + require.NoError(t, os.WriteFile(path, pem.EncodeToMemory(&pem.Block{Type: typ, Bytes: der}), 0o600)) + return path + } + return write("ca.pem", "CERTIFICATE", caDER), write("client.pem", "CERTIFICATE", leafDER), write("client.key", "EC PRIVATE KEY", keyDER) +} diff --git a/scripts/orchestrator/main.go b/scripts/orchestrator/main.go index a128d402e..10eacba89 100644 --- a/scripts/orchestrator/main.go +++ b/scripts/orchestrator/main.go @@ -1,10 +1,12 @@ // E2E orchestrator — drives a clean, isolated E2E test session against -// one ClickHouse + one WaveHouse, then runs the vitest suite once. +// one ClickHouse + one Redis + one WaveHouse, then runs the vitest suite +// once. // // Lifecycle: // -// 1. Start ClickHouse via testcontainers-go (random host ports — no -// conflict with `make dev` or other compose stacks). +// 1. Start ClickHouse and Redis (the fixture's shared cache) via +// testcontainers-go (random host ports — no conflict with `make dev` or +// other compose stacks). // 2. Pick a random free TCP port on 127.0.0.1 and start bin/wavehouse-cov // bound to it (WH_SERVER_PORT) with auth enabled. Random port avoids // conflicts with `make dev`, dev servers, and previous runs that may @@ -44,6 +46,7 @@ import ( "syscall" "time" + "github.com/moby/moby/api/types/container" "github.com/testcontainers/testcontainers-go" "github.com/testcontainers/testcontainers-go/wait" ) @@ -181,6 +184,43 @@ func run() error { return fmt.Errorf("settings dir: %w", err) } + // The shared cache the fixture's cache.backend=redis names: the suite + // runs the backend a multi-instance deployment runs. No persistence, and + // /data on tmpfs so the image's VOLUME leaves no anonymous volume behind. + log.Println("→ starting Redis testcontainer...") + redis, err := testcontainers.GenericContainer(ctx, testcontainers.GenericContainerRequest{ + ContainerRequest: testcontainers.ContainerRequest{ + // Pinned to match internal/cache's integration suite. + Image: "redis:8.10.2-alpine", + Cmd: []string{"redis-server", "--save", "", "--appendonly", "no"}, + ExposedPorts: []string{"6379/tcp"}, + HostConfigModifier: func(hc *container.HostConfig) { + hc.Tmpfs = map[string]string{"/data": ""} + }, + WaitingFor: wait.ForLog("Ready to accept connections").WithStartupTimeout(60 * time.Second), + }, + Started: true, + }) + if err != nil { + return fmt.Errorf("redis start: %w", err) + } + defer func() { + log.Println("→ terminating Redis testcontainer...") + if err := redis.Terminate(context.Background()); err != nil { + log.Printf(" redis terminate: %v", err) + } + }() + redisHost, err := redis.Host(ctx) + if err != nil { + return fmt.Errorf("redis host: %w", err) + } + redisPort, err := redis.MappedPort(ctx, "6379") + if err != nil { + return fmt.Errorf("redis port: %w", err) + } + redisAddr := net.JoinHostPort(redisHost, redisPort.Port()) + log.Printf("✓ Redis ready: %s", redisAddr) + whPort, err := pickFreePort(ctx) if err != nil { return fmt.Errorf("pick free port: %w", err) @@ -198,14 +238,15 @@ func run() error { // tests/e2e/fixtures/config.yaml and the tunables, policy, roles, and // pipes in tests/e2e/fixtures/settings — edit them there, not here. The // vars below are the per-run dynamic overrides (port, scratch paths, the - // patched settings copy) plus GOCOVERDIR and WH_CONFIG, which can't live - // in YAML. + // patched settings copy, the Redis address) plus GOCOVERDIR and + // WH_CONFIG, which can't live in YAML. whCmd.Env = append(os.Environ(), "GOCOVERDIR="+coverDir, "WH_CONFIG="+filepath.Join(repoRoot, "tests", "e2e", "fixtures", "config.yaml"), "WH_SERVER_PORT="+strconv.Itoa(whPort), "WH_SETTINGS_DIR="+settingsDir, "WH_DATA_DIR="+dataDir, + "WH_CACHE_REDIS_ADDRS="+redisAddr, ) if os.Getenv("OTEL_EXPORTER_OTLP_ENDPOINT") == "" { diff --git a/tests/e2e/fixtures/config.yaml b/tests/e2e/fixtures/config.yaml index 4ce30280b..69e8cb33e 100644 --- a/tests/e2e/fixtures/config.yaml +++ b/tests/e2e/fixtures/config.yaml @@ -1,8 +1,9 @@ # WaveHouse config for the e2e harness (scripts/orchestrator + tests/e2e). # -# Dynamic values — ClickHouse addr/HTTP port (testcontainer), WaveHouse -# server port (free port), WH_DATA_DIR (per-run scratch) — are injected by -# the orchestrator via env vars. Everything below is pinned here so the +# Dynamic values — ClickHouse addr/HTTP port (testcontainer), the Redis +# address (testcontainer, WH_CACHE_REDIS_ADDRS), WaveHouse server port (free +# port), WH_DATA_DIR (per-run scratch) — are injected by the orchestrator via +# env vars. Everything below is pinned here so the # rig config is visible and editable without recompiling Go. # JWT validation. The suite signs test tokens with this fixed dev secret; @@ -13,6 +14,15 @@ auth: # operator_key: "" # non-JWT full-access operator credential (Authorization: Operator, or X-Operator-Key); unset in the rig +# The shared cache, as a multi-instance deployment runs it; the in-process +# backend is covered by the unit and integration suites. The timeout is ten +# times the default so a loaded runner's slow round trip is not a bypassed +# lookup that a HIT assertion reads as a failure. +cache: + backend: redis + redis: + timeout: 1s + # Tenant tunables (schema refresh_interval 5s so schema-discovery tests don't # wait a minute; dedupe enabled + id_field; DLQ on; CORS "*"; a 1 GiB NATS # stream budget so the testcontainer stays tiny), the policy, and the pipes diff --git a/tests/integration/shared_cache_test.go b/tests/integration/shared_cache_test.go new file mode 100644 index 000000000..b665ccb67 --- /dev/null +++ b/tests/integration/shared_cache_test.go @@ -0,0 +1,271 @@ +//go:build integration + +package tests + +import ( + "context" + "fmt" + "io" + "net" + "net/http" + "net/url" + "strings" + "sync/atomic" + "testing" + "time" + + "github.com/moby/moby/api/types/container" + "github.com/moby/moby/client" + "github.com/redis/rueidis" + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + "github.com/testcontainers/testcontainers-go" + "github.com/testcontainers/testcontainers-go/wait" + + "github.com/Wave-RF/WaveHouse/internal/app" + "github.com/Wave-RF/WaveHouse/internal/config" +) + +// Pinned, as internal/cache's integration suite pins it. +const redisImage = "redis:8.10.2-alpine" + +// minCacheTTL is cache.QueryTimeToTTL's floor: a fill made less than this +// long ago cannot have expired, so a miss inside it is an invalidation. +const minCacheTTL = 10 * time.Second + +var cachePrefixes atomic.Uint64 + +// startRedis runs a throwaway Redis with no persistence, its /data on tmpfs +// so the image's VOLUME leaves no anonymous volume behind. +func startRedis(t *testing.T) (testcontainers.Container, string) { + t.Helper() + ctx := context.Background() + ctr, err := testcontainers.GenericContainer(ctx, testcontainers.GenericContainerRequest{ + ContainerRequest: testcontainers.ContainerRequest{ + Image: redisImage, + Cmd: []string{"redis-server", "--save", "", "--appendonly", "no"}, + ExposedPorts: []string{"6379/tcp"}, + HostConfigModifier: func(hc *container.HostConfig) { + hc.Tmpfs = map[string]string{"/data": ""} + }, + WaitingFor: wait.ForLog("Ready to accept connections").WithStartupTimeout(90 * time.Second), + }, + Started: true, + }) + testcontainers.CleanupContainer(t, ctr) + require.NoError(t, err) + host, err := ctr.Host(ctx) + require.NoError(t, err) + port, err := ctr.MappedPort(ctx, "6379/tcp") + require.NoError(t, err) + return ctr, net.JoinHostPort(host, port.Port()) +} + +// bootRedisApp runs a second, independent WaveHouse — its own embedded NATS, +// ingest worker and data_dir — against the suite's ClickHouse, with +// cache.backend=redis. It returns the instance's base URL. +func bootRedisApp(t *testing.T, redisAddr, prefix string, timeout time.Duration) string { + t.Helper() + e := env(t) + ctx := context.Background() + settingsDir, err := writeTestSettings(e.ch) + require.NoError(t, err) + var lc net.ListenConfig + ln, err := lc.Listen(ctx, "tcp", "127.0.0.1:0") + require.NoError(t, err) + cfg := &config.Config{ + DataDir: t.TempDir(), + Server: config.Server{ShutdownTimeout: 10}, + ClickHouse: config.ClickHouse{Password: testCHPassword}, + MQ: config.MQ{Backend: config.MQEmbedded}, + Cache: config.Cache{Backend: config.CacheRedis, Redis: config.CacheRedisConfig{ + Addrs: []string{redisAddr}, Mode: config.RedisStandalone, KeyPrefix: prefix, + Timeout: timeout, DialTimeout: 2 * time.Second, + MaxValueBytes: 1 << 20, CompressMinBytes: 1 << 10, VersionTTL: time.Hour, + }}, + Dedupe: config.Dedupe{Backend: config.DedupePebble}, + Coord: config.Coord{Backend: config.CoordLocal}, + Settings: config.Settings{Dir: settingsDir}, + } + a, err := app.New(ctx, app.Options{Config: cfg, Listener: ln}) + require.NoError(t, err) + runCtx, stop := context.WithCancel(ctx) + runDone := make(chan error, 1) + go func() { runDone <- a.Run(runCtx) }() + t.Cleanup(func() { + stop() + assert.NoError(t, <-runDone) + closeCtx, cancel := context.WithTimeout(context.Background(), 10*time.Second) + defer cancel() + assert.NoError(t, a.Close(closeCtx)) + }) + baseURL := "http://" + ln.Addr().String() + require.NoError(t, waitForLive(ctx, baseURL, 30*time.Second)) + return baseURL +} + +// structuredQuery posts a select-all structured query and returns the +// status, the X-Cache header and the body. +func structuredQuery(t *testing.T, baseURL, table string) (int, string, string) { + t.Helper() + status, xc, body, err := tryStructuredQuery(baseURL, table) + require.NoError(t, err) + return status, xc, body +} + +// tryStructuredQuery is structuredQuery for an Eventually condition, which +// runs off the test goroutine and so must not call require. +func tryStructuredQuery(baseURL, table string) (int, string, string, error) { + req, err := http.NewRequestWithContext(context.Background(), http.MethodPost, + baseURL+"/v1/query?table="+url.QueryEscape(table), strings.NewReader(`{"select_all":true}`)) + if err != nil { + return 0, "", "", err + } + req.Header.Set("Content-Type", "application/json") + resp, err := http.DefaultClient.Do(req) + if err != nil { + return 0, "", "", err + } + defer func() { _ = resp.Body.Close() }() + body, err := io.ReadAll(resp.Body) + return resp.StatusCode, resp.Header.Get("X-Cache"), string(body), err +} + +func ingestRow(t *testing.T, baseURL, table, user string) { + t.Helper() + resp, err := http.Post(baseURL+"/v1/ingest?table="+url.QueryEscape(table), "application/json", + strings.NewReader(fmt.Sprintf(`{"user_id":%q,"value":1}`, user))) + require.NoError(t, err) + defer func() { _ = resp.Body.Close() }() + require.Equal(t, http.StatusOK, resp.StatusCode) +} + +// Two WaveHouse processes share one Redis and one ClickHouse. A result one +// fills is a hit for the other, and an insert one process's worker makes +// invalidates what the other cached: the other's next query is a miss that +// returns the new row, well inside the TTL the stale entry was filed with. +func TestSharedCache_IngestOnOneInstanceInvalidatesAnother(t *testing.T) { + table := createTable(t, "user_id String, value Float64", "ORDER BY user_id") + _, redisAddr := startRedis(t) + prefix := fmt.Sprintf("it%d", cachePrefixes.Add(1)) + a := bootRedisApp(t, redisAddr, prefix, 2*time.Second) + b := bootRedisApp(t, redisAddr, prefix, 2*time.Second) + + rc, err := rueidis.NewClient(rueidis.ClientOption{InitAddress: []string{redisAddr}, DisableCache: true, ForceSingleClient: true}) + require.NoError(t, err) + t.Cleanup(rc.Close) + tableToken := func() string { + v, err := rc.Do(context.Background(), rc.B().Get().Key(prefix+":{0}:B:"+table).Build()).ToString() + if rueidis.IsRedisNil(err) { + return "" + } + require.NoError(t, err) + return v + } + + status, xc, body := structuredQuery(t, b, table) + require.Equal(t, http.StatusOK, status, body) + require.Equal(t, "MISS", xc) + status, xc, _ = structuredQuery(t, a, table) + require.Equal(t, http.StatusOK, status) + require.Equal(t, "HIT", xc, "a fills, b hits: one cache") + + // An insert's worker and its invalidation run on whichever process took + // the ingest, so the discriminating window is the stale entry's TTL: if + // the batch window and load push the new row past it, the round proves + // nothing and runs again with a fresh fill. + for round := 1; ; round++ { + user := fmt.Sprintf("user-%d", round) + status, xc, body = structuredQuery(t, b, table) + require.Equal(t, http.StatusOK, status, body) + require.NotContains(t, body, user) + filled := time.Now() + if xc == "HIT" { + // The previous round's fill: refresh it so the TTL window starts now. + _, err := rc.Do(context.Background(), rc.B().Flushdb().Build()).ToString() + require.NoError(t, err) + status, xc, body = structuredQuery(t, b, table) + require.Equal(t, http.StatusOK, status, body) + require.Equal(t, "MISS", xc) + filled = time.Now() + } + before := tableToken() + + ingestRow(t, a, table, user) + var seenAt time.Time + require.Eventually(t, func() bool { + var err error + status, xc, body, err = tryStructuredQuery(b, table) + if err == nil && status == http.StatusOK && strings.Contains(body, user) { + seenAt = time.Now() + return true + } + return false + }, 30*time.Second, 100*time.Millisecond, "b never served the row ingested through a") + assert.NotEqual(t, before, tableToken(), "a's worker bumped the table token in the shared server") + if seenAt.Sub(filled) < minCacheTTL-time.Second { + assert.Equal(t, "MISS", xc, "the first answer carrying the new row is b's refill") + status, xc, _ = structuredQuery(t, b, table) + require.Equal(t, http.StatusOK, status) + assert.Equal(t, "HIT", xc, "b's refill is cached again") + return + } + require.Less(t, round, 3, "b served the new row only once its stale entry could have expired, in every round: the invalidation never reached it, or ingest is too slow here to tell") + t.Logf("round %d: row landed %s after the fill, past the TTL floor; retrying", round, seenAt.Sub(filled)) + } +} + +// A Redis that stops answering costs queries nothing but the cache: they +// keep succeeding, straight from ClickHouse, each a miss; an ingest made +// meanwhile is visible at once. Once it answers again, the cache serves hits. +func TestSharedCache_RedisDownQueriesBypass(t *testing.T) { + ctx := context.Background() + table := createTable(t, "user_id String, value Float64", "ORDER BY user_id") + ctr, redisAddr := startRedis(t) + const timeout = 200 * time.Millisecond + a := bootRedisApp(t, redisAddr, fmt.Sprintf("it%d", cachePrefixes.Add(1)), timeout) + + status, xc, _ := structuredQuery(t, a, table) + require.Equal(t, http.StatusOK, status) + require.Equal(t, "MISS", xc) + status, xc, _ = structuredQuery(t, a, table) + require.Equal(t, http.StatusOK, status) + require.Equal(t, "HIT", xc) + + d, err := testcontainers.NewDockerClientWithOpts(ctx) + require.NoError(t, err) + t.Cleanup(func() { _ = d.Close() }) + _, err = d.ContainerPause(ctx, ctr.GetContainerID(), client.ContainerPauseOptions{}) + require.NoError(t, err) + paused := true + unpause := func() { + if paused { + paused = false + _, err := d.ContainerUnpause(ctx, ctr.GetContainerID(), client.ContainerUnpauseOptions{}) + require.NoError(t, err) + } + } + t.Cleanup(unpause) + + for range 8 { + start := time.Now() + status, xc, body := structuredQuery(t, a, table) + require.Equal(t, http.StatusOK, status, body) + assert.Equal(t, "MISS", xc) + // A lookup and a fill each wait at most the timeout; the query itself + // is a few ms. Generous for -race under load. + assert.Less(t, time.Since(start), 5*timeout+2*time.Second) + } + + ingestRow(t, a, table, "while-down") + require.Eventually(t, func() bool { + status, _, body, err := tryStructuredQuery(a, table) + return err == nil && status == http.StatusOK && strings.Contains(body, "while-down") + }, 30*time.Second, 200*time.Millisecond, "a row ingested while the cache is down is served") + + unpause() + require.Eventually(t, func() bool { + _, xc, body, err := tryStructuredQuery(a, table) + return err == nil && xc == "HIT" && strings.Contains(body, "while-down") + }, 30*time.Second, 200*time.Millisecond, "the cache serves hits again once the server answers") +} From 13c6cb3014d211ed9f41f47c98ad49746599d0b4 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 03:44:18 -0400 Subject: [PATCH 13/79] =?UTF-8?q?fix(cache):=20review=20round=20=E2=80=94?= =?UTF-8?q?=20e2e=20gate,=20trust=20boundary,=20#386,=20addr=20spaces?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Exclude LocalCache from the e2e gate now that e2e runs on Redis; document the shared server as a trust boundary, the noeviction and mutation-pipe (#386) staleness cases, and the connection commands an ACL user needs; refuse addresses with surrounding spaces. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- .testcoverage.yml | 5 +++++ docs/src/content/docs/api.md | 6 +++--- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/configuration.mdx | 6 +++--- docs/src/content/docs/deployment.md | 12 +++++++++++- internal/config/cache_redis.go | 6 ++++-- internal/config/cache_redis_test.go | 1 + 7 files changed, 28 insertions(+), 10 deletions(-) diff --git a/.testcoverage.yml b/.testcoverage.yml index aff1a694d..59ff6d0b7 100644 --- a/.testcoverage.yml +++ b/.testcoverage.yml @@ -73,3 +73,8 @@ exclude: - ^internal/settings/ - ^cmd/wavehouse/validate\.go$ - ^cmd/wavehouse/bootstrap\.go$ + # The in-process cache backend: the e2e stack runs cache.backend=redis + # (#613), so the binary carries LocalCache and its version index but e2e + # never reaches them. The unit suite and the integration suite's main + # app (cache.backend=local) cover them; the merged total still counts them. + - ^internal/cache/(local|version_manager)\.go$ diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index c82739c07..14f5d1308 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -422,7 +422,7 @@ ClickHouse's inline `FORMAT` clause (e.g. `SELECT 1 FORMAT CSV` or `… FORMAT P The proxy buffers the upstream response in memory before forwarding (no row-streaming yet), so a `SELECT *` from a large table can pin RAM on the API server. To avoid an admin OOMing themselves, responses larger than 64 MiB return 502 with a `clickhouse response exceeded N bytes` error. Narrow the query with `LIMIT`, or use a streaming client outside WaveHouse that talks to ClickHouse directly (the standard escape hatch — the same admin credentials work). ::: -This endpoint **does not cache, does not singleflight, and emits `Cache-Control: no-store`** — every request goes straight to ClickHouse, mutation or read, and downstream HTTP caches are explicitly told not to store the response. Raw SQL is an admin escape hatch with infrequent, ad-hoc traffic, so the L1/singleflight machinery would only add complexity without a real hit-rate win. Use [`POST /v1/query?table={table}`](#post-v1querytabletable--structured-query) or [`GET/POST /v1/pipes/{name}`](#getpost-v1pipesname--execute-named-pipe) for the cached read paths (dashboards, high-QPS clients, etc.) — both share an in-process L1 (Ristretto) with singleflight coalescing. +This endpoint **does not cache, does not singleflight, and emits `Cache-Control: no-store`** — every request goes straight to ClickHouse, mutation or read, and downstream HTTP caches are explicitly told not to store the response. Raw SQL is an admin escape hatch with infrequent, ad-hoc traffic, so the L1/singleflight machinery would only add complexity without a real hit-rate win. Use [`POST /v1/query?table={table}`](#post-v1querytabletable--structured-query) or [`GET/POST /v1/pipes/{name}`](#getpost-v1pipesname--execute-named-pipe) for the cached read paths (dashboards, high-QPS clients, etc.) — both go through the query cache ([`cache.backend`](/configuration#backends): in-process, or a Redis shared by every instance) with singleflight coalescing. :::note[Admin only] The route is mounted under `/v1/ops/*`, behind the `RequireAdmin` gate: only a caller whose JWT role equals the policy `admin_role` (`"admin"` by default) — or who presents the non-JWT [operator key](#authentication) — may use it. A tokenless request (or a valid token without a role claim) resolves to the `default_role` (not the admin role unless `default_role` is deliberately set to it — a loudly-warned dev-only setting) and is rejected with `403`; a present-but-invalid token — expired, malformed, bad signature — keeps its stashed verification error and fails loud with `401` instead. Raw SQL has no per-statement scope check (a full SQL parser would be needed to authorize predicates), so the role gate is the entire authorization story, shared with the rest of `/v1/ops/*` (see [Admin Endpoints](#admin-endpoints)). The normal surfaces for non-admin callers are `POST /v1/ingest?table={table}` for writes, `POST /v1/query?table={table}` for structured reads, and `GET/POST /v1/pipes/{name}` for pre-defined queries — none of which expose raw SQL. @@ -538,7 +538,7 @@ Table, column, and alias names may contain any characters ClickHouse accepts — **Response:** -JSON array of result rows. Top-level `DateTime`/`DateTime64` values are returned in canonical RFC 3339 UTC (`2026-06-21T04:00:00.123Z`) — `Nullable` timestamp columns included (a SQL `NULL` renders as JSON `null`), while timestamps nested inside `Array`/`Map`/`Tuple` columns are rendered in the column's declared zone, else the ClickHouse server's, as the driver returns them — byte-identical to the [SSE stream](#get-v1stream--server-sent-events-stream) for values [canonicalized at ingest](#timestamp-canonicalization) (a fail-open pass-through that ClickHouse accepted still comes back canonical here, though it streamed in the producer's spelling). The response carries an `X-Cache: HIT` or `X-Cache: MISS` header — this endpoint shares the in-process L1 (Ristretto) + singleflight machinery (unlike `/v1/ops/query`, which always hits ClickHouse), keyed by [tenant](/deployment#multi-tenant-deployments): a request is never served from, or coalesced with, another tenant's. +JSON array of result rows. Top-level `DateTime`/`DateTime64` values are returned in canonical RFC 3339 UTC (`2026-06-21T04:00:00.123Z`) — `Nullable` timestamp columns included (a SQL `NULL` renders as JSON `null`), while timestamps nested inside `Array`/`Map`/`Tuple` columns are rendered in the column's declared zone, else the ClickHouse server's, as the driver returns them — byte-identical to the [SSE stream](#get-v1stream--server-sent-events-stream) for values [canonicalized at ingest](#timestamp-canonicalization) (a fail-open pass-through that ClickHouse accepted still comes back canonical here, though it streamed in the producer's spelling). The response carries an `X-Cache: HIT` or `X-Cache: MISS` header — this endpoint shares the query cache + singleflight machinery (unlike `/v1/ops/query`, which always hits ClickHouse), keyed by [tenant](/deployment#multi-tenant-deployments): a request is never served from, or coalesced with, another tenant's. The inbound request body is capped at 1 MiB; a body over the cap is rejected with `413`. A query AST is bounded by nature (far under 1 MiB even with a large `in`-list), and the cap blocks a single-request memory-exhaustion vector on this public endpoint. Set a tighter or higher outer limit at your [reverse proxy](/reverse-proxy#request-body-size-limits) — but it can only narrow the effective limit, not raise it past this cap. @@ -575,7 +575,7 @@ Executes a pre-defined named query (pipe) with parameter binding. Parameters can **Response:** -JSON array of result rows, with `X-Cache: HIT` or `X-Cache: MISS` indicating whether the row came from the in-process L1. +JSON array of result rows, with `X-Cache: HIT` or `X-Cache: MISS` indicating whether the rows came from the query cache. The POST parameter body is capped at 1 MiB; a body over the cap is rejected with `413` (the same 1 MiB parameter/AST-body cap as [`POST /v1/query`](#post-v1querytabletable--structured-query) — see [reverse proxy → body limits](/reverse-proxy#request-body-size-limits)). A malformed-but-within-cap body is ignored rather than rejected, since parameters may legitimately come from the query string alone. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 6c59e5f16..a16d01b9d 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -112,7 +112,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on, each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope: one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)). `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. - **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: one version per tenant, per (tenant, table) and per (tenant, table, scope), each keyed by its name alone and bumped in place, so the index holds one entry per live tenant, table and scope however often each is bumped ([#262](https://github.com/Wave-RF/WaveHouse/issues/262)). A query key folds the tenant's version and, for each dependency, its tenant's, table's and scope's, so bumping a table (a scopeless write) orphans every scope of it, and bumping one scope orphans that scope and the whole-table view — scope is reserved and empty today, so every write is the whole-table bump — all without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results. A tenant's version is a *generation*, unique within the process and handed out by the first key built for the tenant; `BumpTenant` (behind `InvalidateTenant`) drops the tenant's whole index, so the next key gets a fresh generation no cached entry folds, orphaning every cached result of the tenant in one step — a pipe result with no dependencies, and a table no bump ever keyed, included — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). `LocalCache.Prune` does the same for every tenant no longer served, which `internal/app` runs after each settings reload, so a tenant removed or rejected stops holding its index. A table bump drops the table's scope versions with it, since every key they were folded into also folds the old table version; and a bump of a tenant with no index is a no-op, since no key folds its next generation yet. The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. -- **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET`, `MGET` and `PING`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token (a string of another length, a list, a hash) is replaced, so a token lost to eviction, expiry or a restart is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing. `cache.backend: redis` selects it: `internal/app`'s `wireCache` maps the boot config's `cache.redis` block onto `RedisConfig`, reading the TLS files, and releases it with the other components. +- **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — its data commands are only `GET`, `SET`, `MGET` and `PING`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token (a string of another length, a list, a hash) is replaced, so a token lost to eviction, expiry or a restart is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing. `cache.backend: redis` selects it: `internal/app`'s `wireCache` maps the boot config's `cache.redis` block onto `RedisConfig`, reading the TLS files, and releases it with the other components. - **redis_codec.go** — the key schema (`tokenKeys`, `bumpKeys`, `valueKey`) and the value frame: format, flags, expiry, the token list, the payload, zstd-compressed from `CompressMinBytes` when that is smaller. A stored value is capped at `MaxValueBytes` and a decoded one at eight times that, which refuses a zip bomb planted in a shared server; an unknown format, as a newer process writes during a rolling upgrade, is a miss. - **breaker.go** — the circuit breaker: `BreakerThreshold` (5) consecutive failures open it, and while open every operation skips the server at once; after `BreakerOpenFor` (5 s) one background `PING` decides whether it closes. Only a transport failure or a timeout counts against the server: any reply, an error reply or one the backend cannot use included, counts as a success, for operations and the probe alike, and a caller that gave up first counts as nothing. - **pending.go** — invalidations the server did not take, kept per token key (repeats coalesce) and retried by a background loop with backoff from 100 ms to 10 s until they land. Landing late is still correct: a fresh token orphans the pre-write entries and any fill made meanwhile. Past `PendingMax` (100,000) keys the set collapses to one tenant bump per affected tenant. `Close` makes one last attempt at the bumps still owed, past the breaker (an open one is why they are owed); what that attempt cannot deliver is lost, and the entries they would orphan are served until their TTL — the same failure as a worker stopping between an insert and its invalidation. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 82744a359..f03a2854f 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -122,11 +122,11 @@ Each tenant's queue has its own disk budget, `mq.max_bytes_gb`, a hot-reloadable | --- | --- | ------- | ----------- | | `cache.l1_max_cost` | `WH_CACHE_L1_MAX_COST` | `67108864` | Maximum size in bytes (~64 MB) of the `local` backend's in-process cache. The time-range bucket structured queries normalize to is `query.timestamp_bucket_seconds` in the [Settings Directory](/settings-directory#configjson-keys). | -The `redis` backend's settings, read only when `cache.backend` is `redis`. It runs on Redis, Valkey, Dragonfly, ElastiCache (including Serverless) and MemoryDB: it sends only `GET`, `SET`, `MGET` and `PING`. [Deployment](/deployment#multiple-instances-and-the-shared-cache) covers sizing, `maxmemory-policy` and what a reader on another instance can see. +The `redis` backend's settings, read only when `cache.backend` is `redis`. It is tested on Redis, Valkey, Dragonfly and a Redis Cluster node, and its data commands are only `GET`, `SET`, `MGET` and `PING` (no scripts, no client tracking), which ElastiCache and MemoryDB also serve. An ACL user also needs the connection commands the client sends when it dials: `HELLO`, `CLIENT`, `SELECT`, and `CLUSTER` in cluster mode; without them it is refused (`NOPERM`) and the cache stays bypassed. Whoever can write to the server can replace cached query results, which are served after the access policy has already been applied, so treat the server as part of WaveHouse's trust boundary (see [Deployment](/deployment#multiple-instances-and-the-shared-cache)). [Deployment](/deployment#multiple-instances-and-the-shared-cache) covers sizing, `maxmemory-policy` and what a reader on another instance can see. | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | -| `cache.redis.addrs` | `WH_CACHE_REDIS_ADDRS` | — | **Required** with `backend: redis.` `host:port` of the server; several are a cluster's seed nodes or the sentinels. Comma-separated in the env var. | +| `cache.redis.addrs` | `WH_CACHE_REDIS_ADDRS` | — | **Required** with `backend: redis`. `host:port` of the server; several are a cluster's seed nodes or the sentinels. Comma-separated in the env var. | | `cache.redis.mode` | `WH_CACHE_REDIS_MODE` | `standalone` | `standalone`, `cluster` or `sentinel`. | | `cache.redis.sentinel_master` | `WH_CACHE_REDIS_SENTINEL_MASTER` | — | The master set name. Required with `mode: sentinel`. | | `cache.redis.username` | `WH_CACHE_REDIS_USERNAME` | — | ACL user. Empty uses the server's `default` user. | @@ -138,7 +138,7 @@ The `redis` backend's settings, read only when `cache.backend` is `redis`. It ru | `cache.redis.tls.key_file` | `WH_CACHE_REDIS_TLS_KEY_FILE` | — | The client certificate's private key (PEM). | | `cache.redis.tls.server_name` | `WH_CACHE_REDIS_TLS_SERVER_NAME` | — | Name to verify the server's certificate against, when it differs from the address. | | `cache.redis.tls.insecure_skip_verify` | `WH_CACHE_REDIS_TLS_INSECURE_SKIP_VERIFY` | `false` | Accept any server certificate. Logged at `WARN` at boot: whoever can intercept the connection can read and replace cached results. | -| `cache.redis.key_prefix` | `WH_CACHE_REDIS_KEY_PREFIX` | `wh` | Leads every key, so several deployments can share one server. No `{` or `}`. | +| `cache.redis.key_prefix` | `WH_CACHE_REDIS_KEY_PREFIX` | `wh` | Leads every key, so several deployments can share one server, provided you trust each as much as the others: any of them can overwrite what the rest serve. No `{` or `}`. | | `cache.redis.timeout` | `WH_CACHE_REDIS_TIMEOUT` | `100ms` | Per operation. A lookup or fill that takes longer is a miss or a skipped fill, never a failed query. | | `cache.redis.dial_timeout` | `WH_CACHE_REDIS_DIAL_TIMEOUT` | `1s` | Per connection attempt. | | `cache.redis.max_value_bytes` | `WH_CACHE_REDIS_MAX_VALUE_BYTES` | `1048576` | Largest result stored, after compression (1 MiB). A larger one is returned to the caller but not cached. | diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 63f1b1a9c..b94beb18c 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -151,6 +151,12 @@ WH_AUTH_JWT_SECRET= # as an admin secret — inject from your secret store, serve only over TLS. WH_AUTH_OPERATOR_KEY= +# Optional shared query cache for several instances (see Multiple instances +# and the shared cache below); the password is a secret like the ones above. +# WH_CACHE_BACKEND=redis +# WH_CACHE_REDIS_ADDRS=redis:6379 +# WH_CACHE_REDIS_PASSWORD= + # Settings directory (required): roles.json, policies.json, pipes.json, # config.json — the hot-reloadable configuration: the access-control policy # and its roles, the named pipes, and the tunables including the ClickHouse @@ -407,9 +413,13 @@ The query-result cache is the layer that can be shared today. With the default ` - **The server is unreachable from the inserting instance.** The invalidation is kept and retried until it lands (`wavehouse_cache_invalidations_pending` counts what is owed). Meanwhile other instances that can still reach the server keep serving the older results, for as long as the outage lasts and at most until each entry's TTL. An instance that stops while invalidations are still owed loses them, with the same bound. The same thing happens today when a process stops between an insert and its invalidation. - **A failover to a replica that had not yet received the latest token writes** can bring back entries filed under the older tokens, bounded by the replication lag at the moment of failover and those entries' TTL. WaveHouse never reads from replicas. +- **The server is full and `maxmemory-policy` is `noeviction`.** It refuses the token writes, so invalidations are kept and retried, and until one lands every instance serves the results from before the insert, up to their TTL. +- **A pipe that writes** (an `INSERT` in `pipes.json`) has its result cached like a read, so a repeated identical call is answered from the cache and the write does not run again ([#386](https://github.com/Wave-RF/WaveHouse/issues/386)). With a shared cache that holds on every instance, until the entry's TTL. - **Admin writes through `POST /v1/ops/query`** do not invalidate the cache ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). With a shared cache, the stale results they leave are served by every instance, not only one. -**Sizing the server.** Every key WaveHouse writes has a TTL, and a version token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses, never bring back an entry it had invalidated. So set `maxmemory` and let the server evict: `maxmemory-policy allkeys-lru` (or `allkeys-lfu`, `volatile-lru`, `volatile-lfu`). Under `noeviction`, a full server refuses the writes. Fills then fail (counted by `wavehouse_cache_set_failures_total{reason="oom"}`), lookups keep working, and invalidations are kept and retried. A stored result is capped at `cache.redis.max_value_bytes` (1 MiB compressed). A tenant's version tokens share one hash tag, so each lookup reads them in one `MGET` in cluster mode as well. The results themselves carry no hash tag and spread across shards. Persistence is not needed: an empty server after a restart is a cold cache, not a wrong one. +**Sizing the server.** Every key WaveHouse writes has a TTL, and a version token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses, never bring back an entry it had invalidated. So set `maxmemory` and let the server evict: `maxmemory-policy allkeys-lru` (or `allkeys-lfu`, `volatile-lru`, `volatile-lfu`). Under `noeviction`, a full server refuses the writes. Fills then fail (counted by `wavehouse_cache_set_failures_total{reason="oom"}`), only lookups whose version tokens already exist keep working, and invalidations are kept and retried, so the pre-insert results above stay served: avoid `noeviction`. A stored result is capped at `cache.redis.max_value_bytes` (1 MiB compressed). A tenant's version tokens share one hash tag, so each lookup reads them in one `MGET` in cluster mode as well. The results themselves carry no hash tag and spread across shards. Persistence is not needed: an empty server after a restart is a cold cache, not a wrong one. + +**The server is inside the trust boundary.** A cached result is served after the access policy has filtered it, so whoever can write to the server can change what any caller reads. Keep it on a private network, require a password or ACL user (`WH_CACHE_REDIS_PASSWORD`), use TLS across links you do not trust, and share it only with deployments you trust as much as this one. **Coalescing stays per instance.** `singleflight` collapses identical concurrent queries within each instance, so a cold hot query costs at most one ClickHouse query per instance, not one per request. diff --git a/internal/config/cache_redis.go b/internal/config/cache_redis.go index f0af67137..3ae3114ea 100644 --- a/internal/config/cache_redis.go +++ b/internal/config/cache_redis.go @@ -55,8 +55,7 @@ type CacheRedisTLS struct { InsecureSkipVerify bool `yaml:"insecure_skip_verify" env:"WH_CACHE_REDIS_TLS_INSECURE_SKIP_VERIFY"` } -// hasAddrs reports whether any address is set. An env file's blank -// `WH_CACHE_REDIS_ADDRS=` loads as one empty address, which is none. +// hasAddrs reports whether any address is set; a YAML `addrs: [""]` is none. func (r CacheRedisConfig) hasAddrs() bool { return len(r.Addrs) > 1 || len(r.Addrs) == 1 && r.Addrs[0] != "" } @@ -66,6 +65,9 @@ func (r CacheRedisConfig) validate() error { return errors.New("cache.backend=redis needs cache.redis.addrs (WH_CACHE_REDIS_ADDRS): the server's host:port, or a cluster's seeds, or the sentinels") } for _, a := range r.Addrs { + if strings.TrimSpace(a) != a { + return fmt.Errorf("cache.redis.addrs (WH_CACHE_REDIS_ADDRS) %q: no spaces around an address", a) + } if _, _, err := net.SplitHostPort(a); err != nil { return fmt.Errorf("cache.redis.addrs (WH_CACHE_REDIS_ADDRS) %q: want host:port: %w", a, err) } diff --git a/internal/config/cache_redis_test.go b/internal/config/cache_redis_test.go index 6d989c737..e1b5783a8 100644 --- a/internal/config/cache_redis_test.go +++ b/internal/config/cache_redis_test.go @@ -189,6 +189,7 @@ func TestValidate_CacheRedis(t *testing.T) { }{ {"defaults", func(*CacheRedisConfig) {}, ""}, {"no addrs", func(r *CacheRedisConfig) { r.Addrs = nil }, "cache.backend=redis needs cache.redis.addrs (WH_CACHE_REDIS_ADDRS)"}, + {"addr with space", func(r *CacheRedisConfig) { r.Addrs = []string{"a:6379", " b:6379"} }, "no spaces around an address"}, {"addr without port", func(r *CacheRedisConfig) { r.Addrs = []string{"redis"} }, `cache.redis.addrs (WH_CACHE_REDIS_ADDRS) "redis": want host:port`}, {"mode", func(r *CacheRedisConfig) { r.Mode = "replica" }, `cache.redis.mode (WH_CACHE_REDIS_MODE) "replica": valid: standalone, cluster, sentinel`}, {"sentinel without master", func(r *CacheRedisConfig) { r.Mode = RedisSentinel }, "needs cache.redis.sentinel_master"}, From 1134d5cf1383e7819d70c4c9075fb15490c89bc0 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 04:17:20 -0400 Subject: [PATCH 14/79] fix(pipes): run write pipes every call, uncached and uncoalesced A pipe whose bound SQL is a write went to ClickHouse through Exec but still had its [] cached and identical in-flight calls coalesced, so a repeat within the TTL answered 200 without writing (#386). With a shared cache that holds on every instance. The handler now classifies the bound SQL with isMutation, the classifier executeCHQuery routes Exec by, and a write skips the cache lookup, fill and singleflight and answers X-Cache: BYPASS. Reads are unchanged. Fixes #386. Part of #613. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- AGENTS.md | 2 +- CHANGELOG.md | 1 + docs/src/content/docs/api.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/pipes.mdx | 10 ++- internal/api/pipes.go | 52 +++++++---- internal/api/pipes_test.go | 120 ++++++++++++++++++++++++++ 7 files changed, 166 insertions(+), 23 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index d3e34545e..880bc93c7 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -63,7 +63,7 @@ The invariant index — what must stay true. Full narrative and rationale live i 10. **Active Sweeper** — purges NATS messages that are both ACKed (written to CH) and older than the gap window; SSE gap-fill uses `DeliverByStartTime`, no in-process ring buffer. 11. **Hasura-style access control: fail-closed (security)** — `policy.IsAdmin` (role == `admin_role`, **exact case-sensitive**, default `"admin"`) is the single admin check, shared by `Evaluate`/`ResolveRole`/`Validate`/the `/v1/ops` gate/`RoleAllowed`. Empty/absent role matches nothing (no `"*"` wildcard); `Validate` rejects empty role keys; a `nil` policy (deleted) denies **everyone incl. admin** via a role — a total lockout for token-based callers, so recovery is writing `policies.json` and reloading, never an implicit admin grant (**exception:** the operator key's `auth.IsOperator` bit passes the `/v1/ops` gate even under a `nil` policy — a deliberate break-glass that can `POST /v1/ops/settings/reload` over HTTP, see #7). Over a nested settings directory the `/v1/ops` gate reads no policy at all — those routes reach every tenant, so the operator key alone passes and an admin-role token gets `403`; `api.NewRouter` decides that from the registry's shape, not from what was wired. `default_role` is the one sanctioned roleless exception (`ResolveRole` maps empty → it pre-eval); `default_role == admin_role` is permitted but dev-only and loudly warned (`policy.DefaultRoleGrantsAdmin`). Preserve when touching `internal/policy` (policy twin of #13; see #159). Detail: architecture.md § `policy/`. 12. **Structured queries: column authz fail-closed (security)** — `POST /v1/query?table={table}`: typed AST validated against schema, permission-enforced, timestamp-bucketed for cache, `DefaultMaxRows` (10,000) cap. Every column reference — projection, aggregation args, `filters`, `group_by`, `order_by`, `time_range` — is authorized inside `query.Build` (the single chokepoint that enumerates them all), so no clause can skip the role's `allow_columns`/`deny_columns` check (#223). A `select_all` read by a *column-restricted* role expands to its allowed columns via `policy.AllowedProjection`, never a bare `SELECT *`; *unrestricted*/admin roles keep `SELECT *` (`policy.RestrictsColumns` decides). Omitting `columns` selects nothing (`ErrEmptyProjection` → `200 []`); `["*"]` is the literal column `*` (schema-gated, not a wildcard); a table-granted role with no readable columns fails closed (`ErrNoReadableColumns` → `403`). Structured and live-stream (`stream.projectIndices`) reads share the one per-column decision `policy.IsColumnAllowed`, so column visibility can't drift. Row visibility has the same one-source guarantee (#319): `Evaluate` resolves a role's row-`filter` once (`resolvePredicates`), and both surfaces consume that single resolution — the query path renders it to SQL (`predicatesToSQL`), the stream evaluates it in memory per subscriber (`ResolvedPermissions.RowVisible`, whose type-aware comparison fails closed on anything it can't prove about the ingested payload — `policy.ColumnSpec`, with `DateTime`/`DateTime64` operands compared as instants through the ingest grammar (`discovery.Column.TimeParser`) and claim constants rendered canonically and digit-exact by the one shared rule `policy.CanonicalScalar` (#457 — which also refuses a float64 at/past 2^53 rather than match a neighboring ID, and whose ok=false — an absent claim, a structured value, no canonical form — makes the predicate match no rows on BOTH surfaces: `1 = 0` in SQL, every row withheld in memory); numeric comparison runs in the column's STORAGE domain (`policy.NumericSpec`, classified by `discovery.NumericStorageOf` — Float width rounding, Decimal scale truncation, integer exactness, both operands narrowed as ClickHouse narrows stored value and bound constant, out-of-range operands refused rather than modeled; the `tests/integration` differential oracle holds in-range verdicts equal to a live ClickHouse's and the never-admit-where-SQL-hides direction for the refused out-of-range ones); an event whose insert later fails into the DLQ is the one residual payload-vs-stored asymmetry, documented in the access-control enforcement caution) — so row visibility can't drift either. Preserve when touching `internal/query` or the structured-query handler. Detail: architecture.md § `query/`. -13. **Named query pipes: fail-closed (security)** — pre-defined SQL templates (Tinybird-style) with param binding + caching; `GET/POST /v1/pipes/{name}` sit outside `RequireAdmin`, so per-pipe `allowed_roles` is the *only* execute-path gate, via `policy.RoleAllowed`: exact allowlist membership (no `"*"`), admin always passes, empty/absent role and empty-string entries authorize nobody, and no `allowed_roles` → admin-only. Preserve and exercise via `testutil.RunRoleMatrix` / `StandardRoleMatrix` (see #159). Detail: architecture.md § `pipes/`. +13. **Named query pipes: fail-closed (security)** — pre-defined SQL templates (Tinybird-style) with param binding + caching — reads only: a pipe whose SQL `isMutation` classifies as a write bypasses the cache and singleflight, since a cached or coalesced write is a dropped one (#386); `GET/POST /v1/pipes/{name}` sit outside `RequireAdmin`, so per-pipe `allowed_roles` is the *only* execute-path gate, via `policy.RoleAllowed`: exact allowlist membership (no `"*"`), admin always passes, empty/absent role and empty-string entries authorize nobody, and no `allowed_roles` → admin-only. Preserve and exercise via `testutil.RunRoleMatrix` / `StandardRoleMatrix` (see #159). Detail: architecture.md § `pipes/`. 14. **TypeScript SDK** — `@wavehouse/sdk`: typed query builder, real-time SSE over `fetch`, live queries (incrementable/decomposable/poll aggregation), codegen CLI. Exactly one runtime dependency — `eventsource-parser` (SSE framing, itself dependency-free); adding a second needs the same scrutiny the first got. The canonical client (see §SDK Sync). 15. **Observability invariants** — stdout always 100% (sampling is OTLP-push-only); WARN+ERROR always export at 100% (a non-configurable floor — don't expose it); gRPC OTel exporters dial lazily so an unreachable collector never blocks startup; the OTel Prometheus exporter uses a **private** `prometheus.Registry`. The OTLP endpoint/TLS/custom-CA/mTLS/headers are delegated to the OpenTelemetry SDK's standard `OTEL_EXPORTER_OTLP_*` env vars — `InitProvider` passes **no** endpoint/header options. Known gap, intentionally not patched in WaveHouse app code: the pinned gRPC logs exporter (`otlploggrpc` v0.19/v0.20) ignores the env TLS-cert vars, so a custom/private CA and mutual TLS apply to traces/metrics but **not** the logs signal (public-CA/system-roots TLS and plaintext still work for logs) — upstream bug open-telemetry/opentelemetry-go#6661. A malformed `OTEL_EXPORTER_OTLP_HEADERS` is logged and skipped by the SDK (fail-soft), not fatal. Preserve when touching the logger/sampler/provider. Detail: architecture.md § `observability/`. 16. **Bearer-token-only CORS posture (security)** — Bearer JWT on every request, no cookies/sessions; `corsMiddleware` deliberately **never** emits `Access-Control-Allow-Credentials` (not needed, and `*` + credentials is a spec violation browsers reject). `cors.allowed_origins` (settings directory, per tenant: a tenant route is decorated from the list of the tenant it names, everything else from tenant `0`'s — `corsOrigins`) controls who can *read* responses, not cookie scope; CSRF protection is structural. Don't reintroduce cookie auth or `Allow-Credentials` without a design discussion — answers GitHub #29/#30. Code: `internal/api/router.go`. diff --git a/CHANGELOG.md b/CHANGELOG.md index 83f2fd832..9db5534c0 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -78,6 +78,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed +- **A pipe that writes runs on every call instead of being answered from the cache** (`internal/api/pipes.go` (+ tests), `docs/src/content/docs/{pipes.mdx,api.md,architecture.md}`, `AGENTS.md`): fixes [#386](https://github.com/Wave-RF/WaveHouse/issues/386). `/v1/pipes/{name}` sent a write's SQL to ClickHouse through `Exec`, but still cached the `[]` it returned and coalesced identical calls in flight, so a repeat within the TTL answered `200` without executing and concurrent identical calls became one write — silently dropped writes, and with a shared cache ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)) on every instance. A pipe whose bound SQL `isMutation` classifies as a write — the same classifier that picks `Exec` — now skips the cache lookup, the fill and singleflight, and answers `X-Cache: BYPASS`. Classification stays automatic rather than a declared pipe property, so an operator cannot forget to mark one, and costs no ClickHouse round trip. Read pipes are unchanged. Not in this fix: a write pipe still does not invalidate cached reads of the table it writes ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). - **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/app/wire.go`, `docs/src/content/docs/{architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` and pipe execution rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL. The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index 1634aaabc..2e5f49696 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -575,7 +575,7 @@ Executes a pre-defined named query (pipe) with parameter binding. Parameters can **Response:** -JSON array of result rows, with `X-Cache: HIT` or `X-Cache: MISS` indicating whether the row came from the in-process L1. +JSON array of result rows, with `X-Cache: HIT` or `X-Cache: MISS` indicating whether the row came from the in-process L1. A pipe whose SQL is a write (`INSERT`, `ALTER`, `WITH … INSERT`, …) bypasses the cache and singleflight: it executes on every call, identical calls in flight are not coalesced, and the response is `[]` with `X-Cache: BYPASS` — see [Pipes that write](/pipes#pipes-that-write). The POST parameter body is capped at 1 MiB; a body over the cap is rejected with `413` (the same 1 MiB parameter/AST-body cap as [`POST /v1/query`](#post-v1querytabletable--structured-query) — see [reverse proxy → body limits](/reverse-proxy#request-body-size-limits)). A malformed-but-within-cap body is ignored rather than rejected, since parameters may legitimately come from the query string alone. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 3ba9a9eb0..57c37df75 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -78,7 +78,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request - **router.go** — Route definitions. Public: `/livez`, `/readyz`, and the content-free `/v1/health` SDK ping (plus the permanent `/healthz` alias and the deprecated `/health`, `/ready` aliases). Policy-gated: `/v1/ingest?table={table}`, `/v1/query?table={table}` (structured), `/v1/pipes/{name}` (named pipes), `/v1/stream`. Admin-only (`RequireAdmin` — role == `policy.admin_role`, or a request bearing the operator key's operator bit, which passes even under a nil policy; over a nested settings directory `NewRouter` mounts the gate with no policy at all, whatever `Dependencies.PolicySource` was wired, so the operator key alone passes): `/v1/ops/schema/*`, `/v1/ops/dlq/stats`, `GET /v1/ops/pipes[/{name}]`, `/v1/ops/settings/reload`, `/v1/ops/query` (raw SQL — same gate as the rest of `/v1/ops/*`). - **auth middleware** — the JWT/JWKS authentication middleware is its own package, [`auth/`](#auth--authentication); the router runs it on every `/v1/*` route. - **tenant.go** — `TenantMW` resolves the request's tenant ahead of the auth middleware on every `/v1` route outside `/v1/ops/*`: the [`X-Tenant-ID`](/deployment#multi-tenant-deployments) header (absent means `tenant.Default`), validated by `tenant.Parse` (`400`), looked up in the `settings.Registry` (`resolveStore`: `404` for an id it does not hold, a bare `503` for a tenant whose folder was rejected — the findings stay out of a body answered before authentication), and the resolved `*settings.Store` stored in the request context (`WithStore` / `StoreFromContext` — here rather than in `tenant/`, because `settings` names `tenant.ID`). A handler reads the store once and passes it down as an argument — the per-tenant getters it holds take it as a parameter (`(*settings.Store).Policy`, `.DedupeFor`, `.DefaultMaxRows`, … in production) — and nothing below a handler reads the context; a tenant route reached without a resolved store answers `500` rather than fall back to a tenant. The probes, `/version`, the metrics path, and `/v1/ops/*` are tenant-exempt; an ops route that addresses one tenant — the admin pipe reads, the schema routes, the raw-SQL proxy, the settings reload and the DLQ stats — names it in `?tenant=` (`opsTenant`, and `opsStore` over it for the routes that need the tenant's store; the DLQ stats need none, since the MQ holds the queue), parsed strictly so that a query `url.ParseQuery` would half-read is a `400` rather than a read of the default tenant, which is what it means when absent on the reads (on the reload, absent is the whole directory). -- **pipes.go** — Named query pipe handlers: admin listing (`GET /v1/ops/pipes[/{name}]`, read per request from its `pipes.Source`) and execution with parameter binding. `pipes.json` is the only write path. +- **pipes.go** — Named query pipe handlers: admin listing (`GET /v1/ops/pipes[/{name}]`, read per request from its `pipes.Source`) and execution with parameter binding. A read is cached and coalesced; a write — bound SQL that `isMutation` (`clickhouse_exec.go`) classifies as one — bypasses both and runs every call. `pipes.json` is the only write path. - **structured_query.go** — Handler for `POST /v1/query?table={table}`: validates query AST, enforces permissions, builds and executes SQL. - **ingest.go** — Accepts `POST /v1/ingest?table={table}` in three body shapes: one flat JSON object, a JSON array of them, or NDJSON. The **required** `Content-Type` chooses the format *family* — `application/json` versus the four NDJSON spellings — and within the JSON family the body's first non-whitespace byte picks array versus single object; the bytes never choose the family. Anything that is not exactly one readable media type is a `415`, decided before the body is read: the header is parsed per RFC 9110 §8.3, and because `Content-Type` is a singleton field, repeated header lines must all resolve to the same format and a value carrying a comma is refused unless the value as a whole parses as one media type — a comma inside a *quoted* parameter value is data, so `application/json; a=", application/x-ndjson; b="` is accepted. It then reads the whole (`MaxBytesReader`-capped) body into a pooled buffer and runs the per-format record readers over those bytes, so the `413` lands before any record is processed and peak memory per request is O(body) rather than O(record). Then it validates each record against the discovered schema, optional dedup, and publishes each row through `mq.Publisher` on `mq.Topic{Tenant, Table, Scope}` (the request's tenant, read off its resolved store — `store.Tenant()` — and raw names; the subject it becomes is `internal/mq`'s; a full queue comes back as `mq.ErrQueueFull`, which is the `503` + `Retry-After`). When dedup is on, a row missing the configured `id_field` can't be deduped: it is logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total` (labeled by `table`), then published un-deduped — or rejected when `dedupe.require_id` is set ([#219](https://github.com/Wave-RF/WaveHouse/issues/219)). - **query.go** — Proxies raw SQL for `POST /v1/ops/query` straight to the `?tenant=`'s ClickHouse HTTP interface (`chconn.Pools.Target` by the resolved store's tenant; the zero target — no pool — is a `503` with `Retry-After`). **Not cached** — sets `Cache-Control: no-store` so every request hits ClickHouse; DateTime is rendered ISO-8601 via `date_time_output_format=iso` (the Go-side type conversion lives in the structured-query / pipes path, not here). diff --git a/docs/src/content/docs/pipes.mdx b/docs/src/content/docs/pipes.mdx index b8c7efa04..80052f44e 100644 --- a/docs/src/content/docs/pipes.mdx +++ b/docs/src/content/docs/pipes.mdx @@ -7,7 +7,7 @@ sidebar: A **named pipe** is a saved SQL query, registered under a name, that callers run by name with parameters — without ever sending raw SQL. They turn an ad-hoc query into a stable, cached, access-controlled endpoint: you write the SQL once as an operator in the settings directory's [`pipes.json`](/settings-directory#pipesjson), expose it at `GET/POST /v1/pipes/{name}`, and clients supply only the declared parameters. -Pipes are the right tool when a query is reusable and shouldn't live in client code — dashboards, reports, public APIs over curated slices of data. They sit on the **cached read path** (shared L1 + singleflight, same as structured queries), and authorize through a simple per-pipe allowlist rather than the full [policy engine](/access-control). +Pipes are the right tool when a query is reusable and shouldn't live in client code — dashboards, reports, public APIs over curated slices of data. They sit on the **cached read path** (shared L1 + singleflight, same as structured queries; a [pipe that writes](#pipes-that-write) bypasses both), and authorize through a simple per-pipe allowlist rather than the full [policy engine](/access-control). ## Anatomy of a pipe @@ -169,7 +169,13 @@ curl -X POST http://localhost:8080/v1/pipes/top_pages \ -d '{"start_date": "2024-01-01", "limit": 20}' ``` -The response is a JSON array of rows. Results flow through the shared in-process L1 cache (Ristretto) with singleflight coalescing, so concurrent identical calls hit ClickHouse once; an `X-Cache: HIT` or `X-Cache: MISS` header tells you which path served the response. +The response is a JSON array of rows. Results flow through the shared in-process L1 cache (Ristretto) with singleflight coalescing, so concurrent identical calls hit ClickHouse once; an `X-Cache: HIT` or `X-Cache: MISS` header tells you which path served the response. A [pipe that writes](#pipes-that-write) skips both and answers `X-Cache: BYPASS`. + +### Pipes that write + +A pipe's SQL may be a write — `INSERT`, `ALTER … DELETE`, `CREATE`, and the rest of ClickHouse's statements that return no rows, including a `WITH … INSERT`. Such a pipe runs on **every** call: it never reads or fills the cache and is never coalesced with an identical call in flight, so ten identical calls are ten writes. The response is `[]` with `X-Cache: BYPASS`. WaveHouse classifies the statement from its leading keyword (after any `WITH` list), with the same classifier that sends it to ClickHouse as a write, so no pipe property marks it. + +Two things a write pipe does not do yet: it does not invalidate cached reads of the table it writes — a structured query or read pipe over that table can serve pre-write rows until its TTL ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)) — and its rows do not reach [`/v1/stream`](/api#get-v1stream--server-sent-events-stream) subscribers, which only the [ingest pipeline](/ingest-pipeline) feeds. For writes that should be seen at once, use [`POST /v1/ingest`](/api#post-v1ingesttabletable--ingest-data). | Status | Body | Cause | | ------ | ---- | ----- | diff --git a/internal/api/pipes.go b/internal/api/pipes.go index 68df3b307..cf5c88381 100644 --- a/internal/api/pipes.go +++ b/internal/api/pipes.go @@ -161,6 +161,23 @@ func (h *PipesHandler) Execute(w http.ResponseWriter, r *http.Request) { return } + // A pipe that writes runs on every call: a cached or coalesced response + // would answer a repeat without executing it, silently dropping the write + // (#386) — on every instance once the cache is shared. isMutation is the + // classifier executeCHQuery routes Exec by, so what bypasses here is + // exactly what runs as a write. + if isMutation(sql) { + data, _, err := h.run(r.Context(), store, conn, sql, params) + if err != nil { + writeJSONError(w, http.StatusInternalServerError, err.Error()) + return + } + w.Header().Set("Content-Type", "application/json") + w.Header().Set("X-Cache", "BYPASS") + _, _ = w.Write(data) //nolint:gosec // G705: JSON the handler marshalled from the exec result + return + } + // Cache. A pipe can read several tables, but the current pipe impl doesn't // expose its table/scope dependencies, so we pass no deps: the result folds // the tenant's version alone, so InvalidateTenant orphans it but no insert @@ -182,28 +199,12 @@ func (h *PipesHandler) Execute(w http.ResponseWriter, r *http.Request) { // Execute with singleflight. v, err, _ := h.sf.Do(cacheKey, func() (interface{}, error) { - queryCtx, cancel := context.WithTimeout(r.Context(), timeoutOf(h.queryTimeout, store)) - defer cancel() - - start := time.Now() - - rows, err := executeCHQuery(queryCtx, conn, sql, params) - queryDuration := time.Since(start) - if err != nil { - // TODO: depending on the error, we may actually want to cache it - return nil, err - } - - data, err := json.Marshal(rows) + data, queryDuration, err := h.run(r.Context(), store, conn, sql, params) if err != nil { - // TODO: eventually we want CSV support etc return nil, err } - - ttl := cache.QueryTimeToTTL(queryDuration) - if h.Cache != nil { - _ = h.Cache.Set(r.Context(), snap, data, ttl) + _ = h.Cache.Set(r.Context(), snap, data, cache.QueryTimeToTTL(queryDuration)) } return data, nil }) @@ -216,3 +217,18 @@ func (h *PipesHandler) Execute(w http.ResponseWriter, r *http.Request) { w.Header().Set("X-Cache", "MISS") _, _ = w.Write(v.([]byte)) //nolint:gosec // G705: the tenant id on the key only selects the entry; the bytes are JSON the handler marshalled from ClickHouse rows } + +// run executes a pipe's bound SQL under the tenant's query timeout and +// returns the rows as JSON with how long ClickHouse took. +func (h *PipesHandler) run(ctx context.Context, store *settings.Store, conn driver.Conn, sql string, params []any) ([]byte, time.Duration, error) { + queryCtx, cancel := context.WithTimeout(ctx, timeoutOf(h.queryTimeout, store)) + defer cancel() + start := time.Now() + rows, err := executeCHQuery(queryCtx, conn, sql, params) + queryDuration := time.Since(start) + if err != nil { + return nil, 0, err + } + data, err := json.Marshal(rows) + return data, queryDuration, err +} diff --git a/internal/api/pipes_test.go b/internal/api/pipes_test.go index db7ff2c41..c554790d9 100644 --- a/internal/api/pipes_test.go +++ b/internal/api/pipes_test.go @@ -7,10 +7,15 @@ import ( "net/http" "net/http/httptest" "strings" + "sync" + "sync/atomic" "testing" + "testing/synctest" "time" + "github.com/ClickHouse/clickhouse-go/v2/lib/driver" "github.com/Wave-RF/WaveHouse/internal/auth" + "github.com/Wave-RF/WaveHouse/internal/cache" "github.com/Wave-RF/WaveHouse/internal/pipes" "github.com/Wave-RF/WaveHouse/internal/policy" "github.com/Wave-RF/WaveHouse/internal/settings" @@ -511,3 +516,118 @@ func TestPipesHandler_Execute_NoAllowedRoles_AdminAllowed(t *testing.T) { "admin bypasses the allowlist on a pipe with no allowed_roles") assert.NotEqual(t, http.StatusNotFound, w.Code) } + +// writeConn counts Exec and Query calls. With gate set, every Exec reports +// itself on entered and holds until gate is closed, so a test can hold +// requests in flight together. +type writeConn struct { + driver.Conn + execs, queries atomic.Int32 + entered, gate chan struct{} +} + +func (c *writeConn) Exec(context.Context, string, ...any) error { + c.execs.Add(1) + if c.gate != nil { + c.entered <- struct{}{} + <-c.gate + } + return nil +} + +func (c *writeConn) Query(context.Context, string, ...any) (driver.Rows, error) { + c.queries.Add(1) + return &chainEmptyRows{}, nil +} + +// pipeCallAs runs the pipe name as the writer role and returns the recorder. +func pipeCallAs(t *testing.T, h *PipesHandler, name string) *httptest.ResponseRecorder { + t.Helper() + w := httptest.NewRecorder() + r := pipesRequest(t, http.MethodPost, "/v1/pipes/"+name, name, map[string]any{"msg": "hello"}) + h.Execute(w, withTenant(r.WithContext(auth.WithRole(r.Context(), "writer")))) + return w +} + +func writerPipesHandler(t *testing.T, conn driver.Conn, c cache.Cache, queries ...*pipes.NamedQuery) *PipesHandler { + t.Helper() + for _, q := range queries { + q.AllowedRoles = []string{"writer"} + } + timeout := func(*settings.Store) time.Duration { return 5 * time.Second } + return NewPipesHandler(staticPipes(queries...), staticPolicy(&policy.Policy{}), fixedConn(conn), c, timeout) +} + +// #386: a pipe that writes executes on every call. Served from the cache, a +// repeat would answer 200 with the first call's `[]` and never reach +// ClickHouse — the write silently dropped. +func TestPipesHandler_Execute_MutationRunsEveryCall(t *testing.T) { + t.Parallel() + for name, sql := range map[string]string{ + "insert": "INSERT INTO audit_log VALUES ({{msg}}, now())", + "insert after cte": "WITH m AS (SELECT {{msg}} AS msg) INSERT INTO audit_log SELECT msg, now() FROM m", + "alter delete": "ALTER TABLE audit_log DELETE WHERE msg = {{msg}}", + } { + t.Run(name, func(t *testing.T) { + t.Parallel() + l1, err := cache.NewLocal(1 << 20) + require.NoError(t, err) + t.Cleanup(func() { _ = l1.Close() }) + conn := &writeConn{} + h := writerPipesHandler(t, conn, l1, &pipes.NamedQuery{Name: "log", SQL: sql}) + + for range 3 { + w := pipeCallAs(t, h, "log") + require.Equal(t, http.StatusOK, w.Code, "body: %s", w.Body.String()) + assert.Equal(t, "BYPASS", w.Header().Get("X-Cache")) + assert.JSONEq(t, `[]`, w.Body.String()) + l1.Wait() + } + assert.Equal(t, int32(3), conn.execs.Load(), "every call must reach ClickHouse") + assert.Zero(t, conn.queries.Load()) + }) + } +} + +// Identical mutation calls in flight together are each executed: coalescing +// them would run one write for all of them. Under synctest, Wait returns once +// every request is inside Exec or parked on another's flight. +func TestPipesHandler_Execute_ConcurrentMutationsNotCoalesced(t *testing.T) { + synctest.Test(t, func(t *testing.T) { + const calls = 3 + conn := &writeConn{entered: make(chan struct{}, calls), gate: make(chan struct{})} + h := writerPipesHandler(t, conn, nil, &pipes.NamedQuery{Name: "log", SQL: "INSERT INTO audit_log VALUES ({{msg}}, now())"}) + var wg sync.WaitGroup + for range calls { + wg.Go(func() { + w := pipeCallAs(t, h, "log") + assert.Equal(t, http.StatusOK, w.Code, "body: %s", w.Body.String()) + }) + } + synctest.Wait() + assert.Len(t, conn.entered, calls, "writes in flight once every request is blocked") + close(conn.gate) + wg.Wait() + assert.Equal(t, int32(calls), conn.execs.Load()) + }) +} + +// A read pipe keeps its cache, including one whose table name starts with a +// write verb: the classifier reads the statement, not the words in it. +func TestPipesHandler_Execute_ReadPipeStaysCached(t *testing.T) { + t.Parallel() + l1, err := cache.NewLocal(1 << 20) + require.NoError(t, err) + t.Cleanup(func() { _ = l1.Close() }) + conn := &writeConn{} + h := writerPipesHandler(t, conn, l1, &pipes.NamedQuery{Name: "recent", SQL: "SELECT * FROM insert_log WHERE msg = {{msg}}"}) + + for _, want := range []string{"MISS", "HIT", "HIT"} { + w := pipeCallAs(t, h, "recent") + require.Equal(t, http.StatusOK, w.Code, "body: %s", w.Body.String()) + assert.Equal(t, want, w.Header().Get("X-Cache")) + l1.Wait() + } + assert.Equal(t, int32(1), conn.queries.Load()) + assert.Zero(t, conn.execs.Load()) +} From 74a7bf43a45c08848a5d70bc62a5e9f47fd75682 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 04:25:31 -0400 Subject: [PATCH 15/79] fix(pipes): no-store on write pipes; reconcile the mutation-path docs A write pipe answers Cache-Control: no-store so an HTTP cache in front of a GET cannot drop the write. api.md, architecture.md and AGENTS.md no longer call /v1/ops/query the only non-insert write path, pipes.mdx says allowed_roles is a write pipe's only gate, and its section moves below the execution error table. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- AGENTS.md | 2 +- CHANGELOG.md | 2 +- docs/src/content/docs/api.md | 6 +++--- docs/src/content/docs/architecture.md | 5 +++-- docs/src/content/docs/pipes.mdx | 14 ++++++++------ internal/api/pipes.go | 4 +++- internal/api/pipes_test.go | 1 + 7 files changed, 20 insertions(+), 14 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 880bc93c7..a6f462e39 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -37,7 +37,7 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability — boot is the validator, there is no dry run - **`dedupe/`** — `Deduplicator` interface → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) -- **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) +- **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy; the only other write path is an admin-authored pipe (#386). A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) - **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, token encoding, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`); `internal/app` constructs it and hands everything else a `mq.Broker` - **`observability/`** — OpenTelemetry pipeline: `InitProvider` wires trace/metric/log providers via OTLP gRPC (each signal independently gated). A top-level `Prometheus` config block drives an optional `/metrics` scrape endpoint that runs independently of OTLP push — standalone (Alloy/Mimir scrape, no collector), alongside OTLP, or off. `NewLogger` produces a slog handler that fans out to stdout AND OTLP (stdout always 100%, OTLP sample-rate-aware). `TraceHandler` injects trace_id/span_id from active spans. `tracer.go` provides W3C trace context propagation over message headers (`InjectHeaders`/`ExtractHeaders` on a plain header map; `internal/mq` injects on every publish and extracts on the `Subscribe` path, so this package never sees a NATS type). - **`pipes/`** — Named query pipes: `NamedQuery` type + `BindParams` + `Source` (read per request; `settings.Store` in production, `Static(q...)` in tests) diff --git a/CHANGELOG.md b/CHANGELOG.md index 9db5534c0..976e847ab 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -78,7 +78,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed -- **A pipe that writes runs on every call instead of being answered from the cache** (`internal/api/pipes.go` (+ tests), `docs/src/content/docs/{pipes.mdx,api.md,architecture.md}`, `AGENTS.md`): fixes [#386](https://github.com/Wave-RF/WaveHouse/issues/386). `/v1/pipes/{name}` sent a write's SQL to ClickHouse through `Exec`, but still cached the `[]` it returned and coalesced identical calls in flight, so a repeat within the TTL answered `200` without executing and concurrent identical calls became one write — silently dropped writes, and with a shared cache ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)) on every instance. A pipe whose bound SQL `isMutation` classifies as a write — the same classifier that picks `Exec` — now skips the cache lookup, the fill and singleflight, and answers `X-Cache: BYPASS`. Classification stays automatic rather than a declared pipe property, so an operator cannot forget to mark one, and costs no ClickHouse round trip. Read pipes are unchanged. Not in this fix: a write pipe still does not invalidate cached reads of the table it writes ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). +- **A pipe that writes runs on every call instead of being answered from the cache** (`internal/api/pipes.go` (+ tests), `docs/src/content/docs/{pipes.mdx,api.md,architecture.md}`, `AGENTS.md`): fixes [#386](https://github.com/Wave-RF/WaveHouse/issues/386). `/v1/pipes/{name}` sent a write's SQL to ClickHouse through `Exec`, but still cached the `[]` it returned and coalesced identical calls in flight, so a repeat within the TTL answered `200` without executing and concurrent identical calls became one write — silently dropped writes, and with a shared cache ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)) on every instance. A pipe whose bound SQL `isMutation` classifies as a write — the same classifier that picks `Exec` — now skips the cache lookup, the fill and singleflight, and answers `X-Cache: BYPASS` with `Cache-Control: no-store`, so an HTTP cache in front of a `GET` cannot drop the write either. Classification stays automatic rather than a declared pipe property, so an operator cannot forget to mark one, and costs no ClickHouse round trip. Read pipes are unchanged. Not in this fix: a write pipe still does not invalidate cached reads of the table it writes ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). - **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/app/wire.go`, `docs/src/content/docs/{architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` and pipe execution rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL. The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index 2e5f49696..f80465711 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -212,9 +212,9 @@ The inbound request body is capped at 16 MiB; a body over the cap is rejected wi The `{table}` URL query must match a table that exists in ClickHouse. WaveHouse discovers table schemas on startup and refreshes them periodically. :::note[Insert-only] -The ingest pipeline accepts only inserts. All other mutations — `DELETE`, `UPDATE`, `TRUNCATE`, `DROP`, `ALTER`, `REPLACE`, etc. — must be issued through [`POST /v1/ops/query`](#post-v1opsquery--query-clickhouse), which is restricted to the admin role (`admin_role`, the same gate as the rest of `/v1/ops/*`). +The ingest pipeline accepts only inserts. All other mutations — `DELETE`, `UPDATE`, `TRUNCATE`, `DROP`, `ALTER`, `REPLACE`, etc. — must be issued through [`POST /v1/ops/query`](#post-v1opsquery--query-clickhouse), which is restricted to the admin role (`admin_role`, the same gate as the rest of `/v1/ops/*`). The one other route is a [pipe that writes](/pipes#pipes-that-write): the admin authors its statement in `pipes.json`, and the roles in its `allowed_roles` run it with parameter values only. -The policy engine authorizes mutations by inspecting the columns being written. That works for inserts but not for predicate-driven mutations like `DELETE … WHERE` — there's no way to prove the predicate matches only rows the caller is allowed to touch. Routing those statements through the admin-gated raw-SQL surface keeps the policy contract honest. +The policy engine authorizes mutations by inspecting the columns being written. That works for inserts but not for predicate-driven mutations like `DELETE … WHERE` — there's no way to prove the predicate matches only rows the caller is allowed to touch. Routing those statements through the admin-gated raw-SQL surface, or through a pipe whose predicate the admin wrote, keeps the policy contract honest. ::: **Request:** @@ -428,7 +428,7 @@ This endpoint **does not cache, does not singleflight, and emits `Cache-Control: The route is mounted under `/v1/ops/*`, behind the `RequireAdmin` gate: only a caller whose JWT role equals the policy `admin_role` (`"admin"` by default) — or who presents the non-JWT [operator key](#authentication) — may use it. A tokenless request (or a valid token without a role claim) resolves to the `default_role` (not the admin role unless `default_role` is deliberately set to it — a loudly-warned dev-only setting) and is rejected with `403`; a present-but-invalid token — expired, malformed, bad signature — keeps its stashed verification error and fails loud with `401` instead. Raw SQL has no per-statement scope check (a full SQL parser would be needed to authorize predicates), so the role gate is the entire authorization story, shared with the rest of `/v1/ops/*` (see [Admin Endpoints](#admin-endpoints)). The normal surfaces for non-admin callers are `POST /v1/ingest?table={table}` for writes, `POST /v1/query?table={table}` for structured reads, and `GET/POST /v1/pipes/{name}` for pre-defined queries — none of which expose raw SQL. ::: -`/v1/ops/query` is the only sanctioned surface for non-insert mutations (the ingest pipeline is insert-only). Granting raw-SQL access to a non-admin role via the policy engine is no longer supported: authenticate with the admin role (`admin_role`). +`/v1/ops/query` is the only surface for ad-hoc non-insert mutations (the ingest pipeline is insert-only; a [pipe that writes](/pipes#pipes-that-write) runs only the statement an admin authored). Granting raw-SQL access to a non-admin role via the policy engine is no longer supported: authenticate with the admin role (`admin_role`). An optional `?tenant=` names the [tenant](/deployment#the-nested-settings-directory) whose ClickHouse the SQL runs against — its own database, credentials and HTTP wiring; without it the SQL runs against tenant `0`'s, which is the whole settings directory unless it is nested. The parameter is parsed as strictly as on the [schema routes](#get-v1opsschema--list-all-table-schemas): `400` for a query string that does not parse or an empty, repeated or malformed id, `404` for an unknown tenant, `503` for one whose settings folder was rejected — all decided before the body is read. A tenant on no ClickHouse pool ([no pool could be opened for it](/settings-directory#clickhouse), such as one the connection ceiling refused) answers `503` `{"error":"no ClickHouse connection is open for this tenant"}` with `Retry-After: 30`. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 57c37df75..e8af1eeed 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -246,8 +246,9 @@ Ingest worker pipeline (StartIngestWorker): (Insert-only pipeline. The wire format `EventMessage` carries only {table_name, scope, received_timestamp, format, columns, row}; non-insert mutations - DELETE/UPDATE/TRUNCATE/DROP/etc. must go through POST /v1/ops/query — the - /v1/ops/* RequireAdmin gate rejects non-admin callers at the API layer, so + DELETE/UPDATE/TRUNCATE/DROP/etc. must go through POST /v1/ops/query (or an + admin-authored write pipe) — the /v1/ops/* RequireAdmin gate rejects + non-admin callers at the API layer, so a no/invalid-token request (resolved to default_role, not admin in a production config) cannot reach the proxy.) diff --git a/docs/src/content/docs/pipes.mdx b/docs/src/content/docs/pipes.mdx index 80052f44e..4886e69bf 100644 --- a/docs/src/content/docs/pipes.mdx +++ b/docs/src/content/docs/pipes.mdx @@ -171,12 +171,6 @@ curl -X POST http://localhost:8080/v1/pipes/top_pages \ The response is a JSON array of rows. Results flow through the shared in-process L1 cache (Ristretto) with singleflight coalescing, so concurrent identical calls hit ClickHouse once; an `X-Cache: HIT` or `X-Cache: MISS` header tells you which path served the response. A [pipe that writes](#pipes-that-write) skips both and answers `X-Cache: BYPASS`. -### Pipes that write - -A pipe's SQL may be a write — `INSERT`, `ALTER … DELETE`, `CREATE`, and the rest of ClickHouse's statements that return no rows, including a `WITH … INSERT`. Such a pipe runs on **every** call: it never reads or fills the cache and is never coalesced with an identical call in flight, so ten identical calls are ten writes. The response is `[]` with `X-Cache: BYPASS`. WaveHouse classifies the statement from its leading keyword (after any `WITH` list), with the same classifier that sends it to ClickHouse as a write, so no pipe property marks it. - -Two things a write pipe does not do yet: it does not invalidate cached reads of the table it writes — a structured query or read pipe over that table can serve pre-write rows until its TTL ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)) — and its rows do not reach [`/v1/stream`](/api#get-v1stream--server-sent-events-stream) subscribers, which only the [ingest pipeline](/ingest-pipeline) feeds. For writes that should be seen at once, use [`POST /v1/ingest`](/api#post-v1ingesttabletable--ingest-data). - | Status | Body | Cause | | ------ | ---- | ----- | | 404 | `{"error":"pipe not found"}` | No pipe registered under that name | @@ -184,6 +178,14 @@ Two things a write pipe does not do yet: it does not invalidate cached reads of | 400 | `{"error":"missing required parameter: x"}` | A required parameter wasn't supplied | | 400 | `{"error":"parameter \"x\": unsupported parameter type object"}` | A non-scalar value with no SQL form — a JSON object (directly, or nested in an array). An empty array is likewise rejected (`array parameter must not be empty`). | +### Pipes that write + +A pipe's SQL may be a write — `INSERT`, `ALTER … DELETE`, `CREATE`, and the rest of ClickHouse's statements that return no rows, including a `WITH … INSERT`. Such a pipe runs on **every** call: it never reads or fills the cache and is never coalesced with an identical call in flight, so ten identical calls are ten writes. The response is `[]` with `X-Cache: BYPASS` and `Cache-Control: no-store`, so an HTTP cache in front of a `GET` does not answer a repeat either. WaveHouse classifies the statement from its leading keyword (after any `WITH` list), with the same classifier that sends it to ClickHouse as a write, so no pipe property marks it. + +`allowed_roles` is a write pipe's only gate: the [policy engine](/access-control)'s insert rules do not apply to it, so any role you list — including a [`default_role`](/access-control#default_role--public-unauthenticated-access) that anonymous callers resolve to — can run the write. The admin fixes the statement and its predicate when authoring the pipe; callers supply only literal values. + +Two things a write pipe does not do yet: it does not invalidate cached reads of the table it writes — a structured query or read pipe over that table can serve pre-write rows until its TTL ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)) — and its rows do not reach [`/v1/stream`](/api#get-v1stream--server-sent-events-stream) subscribers, which only the [ingest pipeline](/ingest-pipeline) feeds ([#362](https://github.com/Wave-RF/WaveHouse/issues/362)). For writes that should be seen at once, use [`POST /v1/ingest`](/api#post-v1ingesttabletable--ingest-data). + ## End-to-end example Ship a curated "top pages" endpoint that the public dashboard can call with no token. diff --git a/internal/api/pipes.go b/internal/api/pipes.go index cf5c88381..0c5be81bf 100644 --- a/internal/api/pipes.go +++ b/internal/api/pipes.go @@ -165,7 +165,8 @@ func (h *PipesHandler) Execute(w http.ResponseWriter, r *http.Request) { // would answer a repeat without executing it, silently dropping the write // (#386) — on every instance once the cache is shared. isMutation is the // classifier executeCHQuery routes Exec by, so what bypasses here is - // exactly what runs as a write. + // exactly what runs as a write. no-store keeps an HTTP cache in front of + // a GET from answering a repeat the same way. if isMutation(sql) { data, _, err := h.run(r.Context(), store, conn, sql, params) if err != nil { @@ -174,6 +175,7 @@ func (h *PipesHandler) Execute(w http.ResponseWriter, r *http.Request) { } w.Header().Set("Content-Type", "application/json") w.Header().Set("X-Cache", "BYPASS") + w.Header().Set("Cache-Control", "no-store") _, _ = w.Write(data) //nolint:gosec // G705: JSON the handler marshalled from the exec result return } diff --git a/internal/api/pipes_test.go b/internal/api/pipes_test.go index c554790d9..3211c57c9 100644 --- a/internal/api/pipes_test.go +++ b/internal/api/pipes_test.go @@ -580,6 +580,7 @@ func TestPipesHandler_Execute_MutationRunsEveryCall(t *testing.T) { w := pipeCallAs(t, h, "log") require.Equal(t, http.StatusOK, w.Code, "body: %s", w.Body.String()) assert.Equal(t, "BYPASS", w.Header().Get("X-Cache")) + assert.Equal(t, "no-store", w.Header().Get("Cache-Control")) assert.JSONEq(t, `[]`, w.Body.String()) l1.Wait() } From 7bb8c271a17f902cdf354e314ae0727573d6ed43 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 04:36:35 -0400 Subject: [PATCH 16/79] docs(pipes): name the operator as a write pipe's author; no-store in api.md Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- AGENTS.md | 2 +- docs/src/content/docs/api.md | 8 ++++---- docs/src/content/docs/architecture.md | 11 ++++++----- docs/src/content/docs/pipes.mdx | 2 +- 4 files changed, 12 insertions(+), 11 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index a6f462e39..9726c0a5c 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -37,7 +37,7 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability — boot is the validator, there is no dry run - **`dedupe/`** — `Deduplicator` interface → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) -- **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy; the only other write path is an admin-authored pipe (#386). A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) +- **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy; the only other write path is an operator-authored pipe (#386). A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) - **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, token encoding, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`); `internal/app` constructs it and hands everything else a `mq.Broker` - **`observability/`** — OpenTelemetry pipeline: `InitProvider` wires trace/metric/log providers via OTLP gRPC (each signal independently gated). A top-level `Prometheus` config block drives an optional `/metrics` scrape endpoint that runs independently of OTLP push — standalone (Alloy/Mimir scrape, no collector), alongside OTLP, or off. `NewLogger` produces a slog handler that fans out to stdout AND OTLP (stdout always 100%, OTLP sample-rate-aware). `TraceHandler` injects trace_id/span_id from active spans. `tracer.go` provides W3C trace context propagation over message headers (`InjectHeaders`/`ExtractHeaders` on a plain header map; `internal/mq` injects on every publish and extracts on the `Subscribe` path, so this package never sees a NATS type). - **`pipes/`** — Named query pipes: `NamedQuery` type + `BindParams` + `Source` (read per request; `settings.Store` in production, `Static(q...)` in tests) diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index f80465711..ca03bacdb 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -212,9 +212,9 @@ The inbound request body is capped at 16 MiB; a body over the cap is rejected wi The `{table}` URL query must match a table that exists in ClickHouse. WaveHouse discovers table schemas on startup and refreshes them periodically. :::note[Insert-only] -The ingest pipeline accepts only inserts. All other mutations — `DELETE`, `UPDATE`, `TRUNCATE`, `DROP`, `ALTER`, `REPLACE`, etc. — must be issued through [`POST /v1/ops/query`](#post-v1opsquery--query-clickhouse), which is restricted to the admin role (`admin_role`, the same gate as the rest of `/v1/ops/*`). The one other route is a [pipe that writes](/pipes#pipes-that-write): the admin authors its statement in `pipes.json`, and the roles in its `allowed_roles` run it with parameter values only. +The ingest pipeline accepts only inserts. All other mutations — `DELETE`, `UPDATE`, `TRUNCATE`, `DROP`, `ALTER`, `REPLACE`, etc. — must be issued through [`POST /v1/ops/query`](#post-v1opsquery--query-clickhouse), which is restricted to the admin role (`admin_role`, the same gate as the rest of `/v1/ops/*`). The one other route is a [pipe that writes](/pipes#pipes-that-write): an operator authors its statement in `pipes.json`, and the roles in its `allowed_roles` run it with parameter values only. -The policy engine authorizes mutations by inspecting the columns being written. That works for inserts but not for predicate-driven mutations like `DELETE … WHERE` — there's no way to prove the predicate matches only rows the caller is allowed to touch. Routing those statements through the admin-gated raw-SQL surface, or through a pipe whose predicate the admin wrote, keeps the policy contract honest. +The policy engine authorizes mutations by inspecting the columns being written. That works for inserts but not for predicate-driven mutations like `DELETE … WHERE` — there's no way to prove the predicate matches only rows the caller is allowed to touch. Routing those statements through the admin-gated raw-SQL surface, or through a pipe whose predicate the operator wrote, keeps the policy contract honest. ::: **Request:** @@ -428,7 +428,7 @@ This endpoint **does not cache, does not singleflight, and emits `Cache-Control: The route is mounted under `/v1/ops/*`, behind the `RequireAdmin` gate: only a caller whose JWT role equals the policy `admin_role` (`"admin"` by default) — or who presents the non-JWT [operator key](#authentication) — may use it. A tokenless request (or a valid token without a role claim) resolves to the `default_role` (not the admin role unless `default_role` is deliberately set to it — a loudly-warned dev-only setting) and is rejected with `403`; a present-but-invalid token — expired, malformed, bad signature — keeps its stashed verification error and fails loud with `401` instead. Raw SQL has no per-statement scope check (a full SQL parser would be needed to authorize predicates), so the role gate is the entire authorization story, shared with the rest of `/v1/ops/*` (see [Admin Endpoints](#admin-endpoints)). The normal surfaces for non-admin callers are `POST /v1/ingest?table={table}` for writes, `POST /v1/query?table={table}` for structured reads, and `GET/POST /v1/pipes/{name}` for pre-defined queries — none of which expose raw SQL. ::: -`/v1/ops/query` is the only surface for ad-hoc non-insert mutations (the ingest pipeline is insert-only; a [pipe that writes](/pipes#pipes-that-write) runs only the statement an admin authored). Granting raw-SQL access to a non-admin role via the policy engine is no longer supported: authenticate with the admin role (`admin_role`). +`/v1/ops/query` is the only surface for ad-hoc non-insert mutations (the ingest pipeline is insert-only; a [pipe that writes](/pipes#pipes-that-write) runs only the statement an operator authored). Granting raw-SQL access to a non-admin role via the policy engine is no longer supported: authenticate with the admin role (`admin_role`). An optional `?tenant=` names the [tenant](/deployment#the-nested-settings-directory) whose ClickHouse the SQL runs against — its own database, credentials and HTTP wiring; without it the SQL runs against tenant `0`'s, which is the whole settings directory unless it is nested. The parameter is parsed as strictly as on the [schema routes](#get-v1opsschema--list-all-table-schemas): `400` for a query string that does not parse or an empty, repeated or malformed id, `404` for an unknown tenant, `503` for one whose settings folder was rejected — all decided before the body is read. A tenant on no ClickHouse pool ([no pool could be opened for it](/settings-directory#clickhouse), such as one the connection ceiling refused) answers `503` `{"error":"no ClickHouse connection is open for this tenant"}` with `Retry-After: 30`. @@ -575,7 +575,7 @@ Executes a pre-defined named query (pipe) with parameter binding. Parameters can **Response:** -JSON array of result rows, with `X-Cache: HIT` or `X-Cache: MISS` indicating whether the row came from the in-process L1. A pipe whose SQL is a write (`INSERT`, `ALTER`, `WITH … INSERT`, …) bypasses the cache and singleflight: it executes on every call, identical calls in flight are not coalesced, and the response is `[]` with `X-Cache: BYPASS` — see [Pipes that write](/pipes#pipes-that-write). +JSON array of result rows, with `X-Cache: HIT` or `X-Cache: MISS` indicating whether the row came from the in-process L1. A pipe whose SQL is a write (`INSERT`, `ALTER`, `WITH … INSERT`, …) bypasses the cache and singleflight: it executes on every call, identical calls in flight are not coalesced, and the response is `[]` with `X-Cache: BYPASS` and `Cache-Control: no-store` (so an HTTP cache in front of a `GET` cannot answer a repeat) — see [Pipes that write](/pipes#pipes-that-write). The POST parameter body is capped at 1 MiB; a body over the cap is rejected with `413` (the same 1 MiB parameter/AST-body cap as [`POST /v1/query`](#post-v1querytabletable--structured-query) — see [reverse proxy → body limits](/reverse-proxy#request-body-size-limits)). A malformed-but-within-cap body is ignored rather than rejected, since parameters may legitimately come from the query string alone. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index e8af1eeed..3aff5cc6e 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -247,7 +247,7 @@ Ingest worker pipeline (StartIngestWorker): (Insert-only pipeline. The wire format `EventMessage` carries only {table_name, scope, received_timestamp, format, columns, row}; non-insert mutations DELETE/UPDATE/TRUNCATE/DROP/etc. must go through POST /v1/ops/query (or an - admin-authored write pipe) — the /v1/ops/* RequireAdmin gate rejects + operator-authored write pipe) — the /v1/ops/* RequireAdmin gate rejects non-admin callers at the API layer, so a no/invalid-token request (resolved to default_role, not admin in a production config) cannot reach the proxy.) @@ -276,10 +276,11 @@ Client POST /v1/ops/query 401 when a stashed error shows the caller presented an invalid token, else 403. Raw SQL has no per-statement scope check (a full SQL parser would be needed to authorize predicates), so the role gate is the - entire authorization story. /v1/ops/query is the only sanctioned - surface for non-SELECT statements (DELETE/UPDATE/TRUNCATE/DROP/ALTER/…); - non-admin callers use `POST /v1/ingest?table={table}` for writes and - the structured query endpoint or named pipes for reads. + entire authorization story. /v1/ops/query is the only surface for + ad-hoc non-SELECT statements (DELETE/UPDATE/TRUNCATE/DROP/ALTER/…); + non-admin callers use `POST /v1/ingest?table={table}` or a write pipe + that lists their role for writes, and the structured query endpoint or + named pipes for reads. → Decode {"sql": "..."} from the request body. → POST the SQL verbatim to ClickHouse's HTTP interface at ://:/?default_format=JSON diff --git a/docs/src/content/docs/pipes.mdx b/docs/src/content/docs/pipes.mdx index 4886e69bf..d7101edf1 100644 --- a/docs/src/content/docs/pipes.mdx +++ b/docs/src/content/docs/pipes.mdx @@ -182,7 +182,7 @@ The response is a JSON array of rows. Results flow through the shared in-process A pipe's SQL may be a write — `INSERT`, `ALTER … DELETE`, `CREATE`, and the rest of ClickHouse's statements that return no rows, including a `WITH … INSERT`. Such a pipe runs on **every** call: it never reads or fills the cache and is never coalesced with an identical call in flight, so ten identical calls are ten writes. The response is `[]` with `X-Cache: BYPASS` and `Cache-Control: no-store`, so an HTTP cache in front of a `GET` does not answer a repeat either. WaveHouse classifies the statement from its leading keyword (after any `WITH` list), with the same classifier that sends it to ClickHouse as a write, so no pipe property marks it. -`allowed_roles` is a write pipe's only gate: the [policy engine](/access-control)'s insert rules do not apply to it, so any role you list — including a [`default_role`](/access-control#default_role--public-unauthenticated-access) that anonymous callers resolve to — can run the write. The admin fixes the statement and its predicate when authoring the pipe; callers supply only literal values. +`allowed_roles` is a write pipe's only gate: the [policy engine](/access-control)'s insert rules do not apply to it, so any role you list — including a [`default_role`](/access-control#default_role--public-unauthenticated-access) that anonymous callers resolve to — can run the write. The operator fixes the statement and its predicate when authoring the pipe; callers supply only literal values. Two things a write pipe does not do yet: it does not invalidate cached reads of the table it writes — a structured query or read pipe over that table can serve pre-write rows until its TTL ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)) — and its rows do not reach [`/v1/stream`](/api#get-v1stream--server-sent-events-stream) subscribers, which only the [ingest pipeline](/ingest-pipeline) feeds ([#362](https://github.com/Wave-RF/WaveHouse/issues/362)). For writes that should be seen at once, use [`POST /v1/ingest`](/api#post-v1ingesttabletable--ingest-data). From ce14795fe88195f060ceb63f51f6e27816903892 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 13:04:51 -0400 Subject: [PATCH 17/79] refactor(keyenc): one escaping for subject and cache key tokens NATS subject tokens (internal/mq) and cache namespace tokens (query.SafeEncodeToken) each carried a copy of the same encoder. Both now call internal/keyenc: Escape keeps [A-Za-z0-9_] and writes every other byte as %XX, Unescape decodes as url.PathUnescape did, and Join/Split join escaped fields with a separator the escaping never emits. Output is byte-identical, pinned by golden subject tests and against v0.1.0's encoder for every byte value. Part of #613. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- AGENTS.md | 3 +- CHANGELOG.md | 2 + docs/src/content/docs/architecture.md | 7 +- internal/keyenc/keyenc.go | 123 ++++++++++++++++++++++ internal/keyenc/keyenc_test.go | 142 ++++++++++++++++++++++++++ internal/mq/mq.go | 5 +- internal/mq/subject.go | 30 +----- internal/mq/subject_test.go | 73 ++++++------- internal/query/ident.go | 26 +---- 9 files changed, 315 insertions(+), 96 deletions(-) create mode 100644 internal/keyenc/keyenc.go create mode 100644 internal/keyenc/keyenc_test.go diff --git a/AGENTS.md b/AGENTS.md index 3d780d92b..aca624a07 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -26,7 +26,7 @@ One binary: - **`cmd/wavehouse/`** — Standalone mode (all-in-one with embedded NATS, optional Pebble dedup): argv dispatch, the logger, `config.Load`, and the signal context; everything else is `internal/app` -Eighteen internal packages under `internal/` (plus `internal/testutil/` for shared test helpers): +Nineteen internal packages under `internal/` (plus `internal/testutil/` for shared test helpers): - **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers - **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` handing the sweeper each tenant's own gap window (a rejected tenant's as its folder last had it, unbounded for one rejected since boot) and the `mq.max_bytes_gb` reconcile each served tenant's byte budget, and `defaultPolicy` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, and the same hook's `Hub.Prune` ends the open streams of a tenant no longer served), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it @@ -38,6 +38,7 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`dedupe/`** — `Deduplicator` interface → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) +- **`keyenc/`** — the one escaping composite keys are built from: `Escape` keeps `[A-Za-z0-9_]` and writes every other byte as `%XX`, `Unescape` reverses it leniently (as `url.PathUnescape` does), `Join`/`Split` join escaped fields with a separator the escaping never emits. NATS subject tokens and the cache's namespace tokens use it; its output is pinned byte for byte, since v0.1.0's subjects carry it - **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, token encoding, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`); `internal/app` constructs it and hands everything else a `mq.Broker` - **`observability/`** — OpenTelemetry pipeline: `InitProvider` wires trace/metric/log providers via OTLP gRPC (each signal independently gated). A top-level `Prometheus` config block drives an optional `/metrics` scrape endpoint that runs independently of OTLP push — standalone (Alloy/Mimir scrape, no collector), alongside OTLP, or off. `NewLogger` produces a slog handler that fans out to stdout AND OTLP (stdout always 100%, OTLP sample-rate-aware). `TraceHandler` injects trace_id/span_id from active spans. `tracer.go` provides W3C trace context propagation over message headers (`InjectHeaders`/`ExtractHeaders` on a plain header map; `internal/mq` injects on every publish and extracts on the `Subscribe` path, so this package never sees a NATS type). - **`pipes/`** — Named query pipes: `NamedQuery` type + `BindParams` + `Source` (read per request; `settings.Store` in production, `Static(q...)` in tests) diff --git a/CHANGELOG.md b/CHANGELOG.md index 70ce51e1d..914164808 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -32,6 +32,8 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Changed +- **One escaping for composite keys** (`internal/keyenc` (new, + tests), `internal/mq/{mq,subject}.go` (+ tests), `internal/query/ident.go`, `AGENTS.md`, `docs/src/content/docs/architecture.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). NATS subject tokens and the cache's namespace tokens each carried a copy of the same encoder; both now call `internal/keyenc`, which the dedupe keys will use too. Behaviour-preserving: every subject and cache key is byte-identical (pinned by golden tests over dotted, wildcard, whitespace, `-`, `%`, `/`, brace, NUL and non-ASCII names, and against v0.1.0's encoder for every byte value), and a subject token decodes exactly as `url.PathUnescape` decoded it. + - **Each tenant has a message queue of its own** (`internal/mq/{mq,subject,embedded}.go` (+ tests), `internal/ingest/{sweeper,worker}.go` (+ tests), `internal/api/{dlq,ingest}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `internal/stream/{subscriber,hub}.go`, `internal/settings/{settings,store,registry}.go` (+ tests), `internal/testutil/{mocks,testutil}.go`, `clients/ts/src/{dlq,types}.ts` (+ tests), `docs/src/content/docs/{deployment,api,architecture,ingest-pipeline,durability,why-wavehouse}.md`, `docs/src/content/docs/sdk/{admin,reference,streaming}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `AGENTS.md`): story 5b of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). The embedded NATS server keeps each tenant's events on a pair of JetStream streams of its own — `INGEST_` (`ingest..>`, `DiscardNew`) at the tenant's own `mq.max_bytes_gb`, and `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed; subjects are unchanged, and nothing outside `internal/mq` names a stream. A tenant at its budget gets `503` while the others keep publishing, and the ingest worker's and the hub bridge's durables are held on every tenant's stream, each with its own ack floor and `MaxAckPending`, so one tenant's backlog holds back neither another's delivery nor its purge; the worker's prefetch and the hub bridge's fetch-ahead are each shared across the tenants' streams. A removed or rejected tenant's stream is still consumed, so its queued rows reach the worker and are parked on its own dead-letter queue. The sweeper purges each tenant's stream at that tenant's own `stream.gap_window_minutes` — a rejected tenant's as its folder last had it (`settings.Registry.Known` now yields each tenant's last adopted store), and all of its history if the folder has been rejected since boot, so its clients resume once the folder is fixed — keeping no acknowledged history for a removed tenant, and every served tenant's `mq.max_bytes_gb` is applied after each reload rather than tenant `0`'s alone (the boot warning about a nested directory with no tenant `0` is gone with it, and so is the tracking of tenant `0`'s last adopted store: a flat directory's ops gate, its one reader left, reads tenant `0` through the registry). One tenant's failed purge holds up no other tenant's, and the sweep logs it at `ERROR` unless every failure in it is a buffer consumer not created yet. A reload that shrinks a budget no longer deletes dead letters: a dead-letter stream holding more than a tenth of the new budget keeps what it holds, and that is logged — the interim guard of [#532](https://github.com/Wave-RF/WaveHouse/issues/532). JetStream's own check of the streams' caps against the disk, which made them fit 75% of the free disk at boot together, is lifted: a budget is a cap and never a reservation, what the budgets add up to against the disk is [#138](https://github.com/Wave-RF/WaveHouse/issues/138)'s, and a flat directory whose budget exceeds three quarters of the free disk now boots where it used to be refused. A queue that cannot be opened refuses a flat boot like any other store and, over a nested directory, costs its tenant alone — its ingest answers `503`, each reload trying again, and a publish at most once every five seconds, so its clients retrying never hold up another tenant's queue. `GET /v1/ops/dlq/stats` reads one tenant's dead-letter queue — the one `?tenant=` names, parsed strictly like the other admin reads, and tenant `0`'s without it, no longer the sum across tenants — answering for a rejected or removed tenant too and `404` for a tenant with no queue; the SDK's `wh.dlq.list()` and `.table()` take a `tenant` option. The streams an earlier build kept for every tenant together (`WAVEHOUSE`, `WAVEHOUSE_DLQ`) overlap every tenant's subjects and are deleted at boot, what they held with them, and a subject with no tenant token no longer reads as tenant `0`'s. - **Message-queue subjects lead with the tenant, and the async paths read it off each message** (`internal/mq/{mq,subject,embedded}.go`, `internal/api/{ingest,stream}.go`, `internal/stream/hub.go`, `internal/ingest/{worker,sweeper}.go`, `internal/app/wire.go`, `docs/src/content/docs/{architecture,ingest-pipeline,deployment,api}.md`, `docs/src/content/docs/{settings-directory,access-control}.mdx`, `AGENTS.md`): story 5a of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). `mq.Topic` gains a `Tenant`, the leading token of every subject — `ingest..
[.]`, `dlq..
[.]` — placed verbatim, since the tenant-id grammar makes it one token, so one wildcard selects a tenant's traffic (`ingest.acme.>`); a settings directory that holds the four files produces the same subjects with `0` as the token, and nothing else about them changes. The ingest and stream handlers address the request's tenant, which the resolved store carries (`settings.Store.Tenant`, from story 8). The stream hub indexes subscribers by the full topic and evaluates each event under its own tenant's policy, so a subscriber on one tenant's table never receives another tenant's rows for a table of the same name; a gap-fill and the opening schema frame read the connection's tenant. The ingest worker reads each message's tenant off its topic, batches per tenant table, and resolves the dead-letter switch under the row's own tenant — an envelope it cannot read is parked or dropped under the topic's tenant too — and bumps that tenant's cache namespaces (story 8's `invalidate` now receives the message's tenant rather than tenant `0`; the wiring's `sharedTables` still repeats each bump under every tenant the directory holds, since every tenant reads the same ClickHouse table until story 6). `Publish` and gap-fill refuse a topic whose tenant is empty or outside the grammar, so nothing lands on tenant `0` by omission. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index fe0c94c97..b987938c8 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -61,6 +61,7 @@ internal/ ├── dedupe/ Optional deduplication (Pebble) ├── discovery/ ClickHouse schema introspection and validation ├── ingest/ Batch buffering, DLQ, and Active Sweeper +├── keyenc/ The one escaping composite keys are built from (NATS subject tokens, cache namespace tokens) ├── mq/ MQ boundary: the only NATS/JetStream importer (owned message/consumer/stream types + embedded server) ├── observability/ OpenTelemetry pipeline (traces/metrics/logs + Prometheus exposition) ├── pipes/ Named query pipes (NamedQuery type, parameter binding, Source) @@ -146,7 +147,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ The **only** package that imports NATS/JetStream — a `depguard` rule in `.golangci.yml` fails `make lint` on any `github.com/nats-io` import in every package golangci-lint builds; the `integration`-tagged files under `tests/` sit outside its default build context, so the boundary there rests on convention (AGENTS.md Key Design Decision #20). Every other package talks to the broker through the types below, so a subject, stream, or broker change lands here once. - **mq.go** — The owned surface, stated as intent rather than broker mechanics. `Topic{Tenant, Table, Scope}` is the only address the rest of the process handles (a validated tenant id and raw names; comparable, so the SSE hub keys its index by the value). `Message` carries `Data`, its topic (`Topic()` decodes the delivered key on demand — the tenant included, which is how the hub bridge and the worker learn whose event it is; `TopicKey()` is the delivered form, for log lines), and the ack family (`DoubleAck(ctx)`, `Ack()`, `Nak()`); `Headers` is the message header map (`Add`/`Set`/`Get`, exact-key) that `PublishOpt`s such as `WithHeader` shape. Interfaces, each speaking per tenant and never per stream: `Publisher` (`ErrQueueFull` when the tenant's ingest queue is at its byte budget or not open yet — the API's 503 + `Retry-After`), `Subscriber` (every ingest event of every tenant, under a named durable consumer, its fetch-ahead split across the tenants — the hub bridge), `ConsumerManager` → `Consumer` (a durable explicit-ack consumer from a `ConsumerConfig`, whose `MaxAckPending` holds per tenant; `Consume` delivers each tenant's messages on a goroutine of that tenant's, in order, so a blocking handler is backpressure on its own tenant alone, spreads the prefetch across the tenants, and returns a `stop` plus a `failed` channel that reports delivery ending on its own — `ErrDeliveryEnded`, e.g. a deleted consumer, a closed connection, or a tenant's queue that could not be joined — since no message would ever say so) for the ingest worker, `DeadLetterer.DeadLetter` (park a message under its own topic, in its tenant's dead-letter queue; the caller acks), `DeadLetterStats.DeadLetterCounts` (one tenant's; `ErrNoDeadLetterQueue` when it has none), `Purger.PurgeAcked` (drop what is both acked by a consumer and stored before its tenant's cutoff, everything acked for a tenant given none; one error per failed tenant, joined — `ErrConsumerNotFound` for a queue the consumer has not been created on yet, the one failure the sweeper logs as a warning rather than an error) for the sweeper, and `Replayer.ReplaySince` for SSE gap-fill. `Broker` composes them with each tenant's byte budget (`SetMaxBytes`/`MaxBytes`), `Stats`, and `Close`; it is what `internal/app` holds. -- **subject.go** — The embedded broker's naming, private to the package: the stream names (`INGEST_` and `DLQ_` — prefixes that differ in their first letter, so no tenant id makes one kind's name the other's — and the one pair an earlier build kept for every tenant together, `WAVEHOUSE`/`WAVEHOUSE_DLQ`, which boot deletes), the `ingest.`/`dlq.` prefixes and `>` wildcards, the subject-token encoder (alphanumerics and `_` pass, everything else is percent-encoded, so a name can never split or wildcard a subject), and `Topic` ↔ subject conversion. A subject is `.
[.]`: the tenant verbatim — its grammar (`tenant.Parse`) makes it one token, and it is checked on the way to the wire, so a topic without one has no subject — then the table and scope as encoded tokens; tenant first so one wildcard selects a tenant's traffic (`ingest.acme.>`). A topic has the same tail on both streams, so parking on the DLQ is a prefix swap on the delivered subject — nothing is decoded or re-encoded — and the tail's first token picks the tenant's stream. +- **subject.go** — The embedded broker's naming, private to the package: the stream names (`INGEST_` and `DLQ_` — prefixes that differ in their first letter, so no tenant id makes one kind's name the other's — and the one pair an earlier build kept for every tenant together, `WAVEHOUSE`/`WAVEHOUSE_DLQ`, which boot deletes), the `ingest.`/`dlq.` prefixes and `>` wildcards, the subject tokens (`internal/keyenc`: alphanumerics and `_` pass, everything else is percent-encoded, so a name can never split or wildcard a subject), and `Topic` ↔ subject conversion. A subject is `.
[.]`: the tenant verbatim — its grammar (`tenant.Parse`) makes it one token, and it is checked on the way to the wire, so a topic without one has no subject — then the table and scope as encoded tokens; tenant first so one wildcard selects a tenant's traffic (`ingest.acme.>`). A topic has the same tail on both streams, so parking on the DLQ is a prefix swap on the delivered subject — nothing is decoded or re-encoded — and the tail's first token picks the tenant's stream. - **purge.go** — The Active Sweeper's arithmetic over JetStream sequences: purge target = `MIN(consumer ack floor + 1, first sequence stored at or after the cutoff)`, the latter found by binary search over message timestamps (~15 lookups). Every uncertainty resolves toward purging less: a sequence that holds no message is kept as a candidate bound rather than discarding the half below it, and a lookup that fails outright aborts that tenant's purge. It runs on each tenant's stream at that tenant's cutoff, and a failure on one tenant's stream is reported without stopping the sweep of the others. Healthy state keeps exactly the gap window; ClickHouse down freezes purging; a catastrophic outage fills the stream to `MaxBytes` and `DiscardNew` pushes back. - **embedded.go** — `EmbeddedNATS`, the one `Broker`: an in-process NATS server with JetStream, giving each tenant a queue of its own — stream `INGEST_` with subjects `ingest..>`, capped at the tenant's `mq.max_bytes_gb` (`DiscardNew`), and stream `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — with the durable consumers on the ingest one; nothing outside the package sees that layout. JetStream's own check of the streams' caps against the disk (75% of the free disk by default) is set out of reach, so a budget is a cap and never a reservation. Boot deletes the pair an earlier build kept for every tenant together — its subjects overlap every tenant's — and takes stock of the tenants' streams on disk with their budgets, so a consumer created later is held on every one, a tenant no longer served included. `SetMaxBytes` opens a tenant's queue the first time — its dead-letter stream first, so no row is queued that could not be parked — and every registered consumer joins it; a publish or park that finds a stream missing reopens it at the budget last asked for the tenant, and so does a publish to a queue the broker has not recorded open — an open that timed out can leave a stream JetStream creates after all, which no consumer holds — either refused as a full queue with none asked yet. Publishes and parks that find the same queue not open share one attempt (`singleflight`), and after one fails the tenant's publishes and parks are refused at once for five seconds rather than each trying again under the broker's lock, which every tenant's open, resize and reload takes; a reload retries regardless. After that `SetMaxBytes` applies a reloaded budget to the tenant's two live streams as a pair: if the DLQ update fails after the ingest one succeeded, the ingest resize is undone so both stay on the previous budget — best effort, since if that undo also fails the ingest stream stays at the new limit and the DLQ at the previous, and the error says so. A dead-letter stream is never capped below the bytes it holds, which `DiscardOld` would delete to fit ([#532](https://github.com/Wave-RF/WaveHouse/issues/532)): it keeps what it holds, and that is logged. Its JetStream calls are bounded to ten seconds — plus ten more for the consumers joining a queue it has just opened, and five for the rollback of a failed resize, each a budget of its own rather than the one that just expired — since a reload holds the settings store's lock while its hooks run; `MaxBytes` reports the budget last applied in full, so a failed resize is retried by the next reload. The consumers `CreateConsumer` and `Subscribe` build hold one durable on each tenant's stream, looked up before anything is written so a boot over many queues writes nothing it need not, each delivering on a goroutine of its own into the one handler. `PurgeAcked` and `DeadLetterCounts` run per tenant stream. `Stats` reports connection and inbound-message counters for `observability.RegisterSystemMetrics`. Trace context rides in the message headers: `Publish` applies `observability.InjectHeaders`, and a message delivered through `Subscribe` (the hub bridge) carries `observability.ExtractHeaders` on its `Ctx`; the worker's `Consumer` path skips the extraction, since it batches across messages and reads no per-message context. @@ -200,6 +201,10 @@ The hot-reloadable half of configuration: a directory of four JSON files (`confi - **chsql.go** — Dependency-free ClickHouse SQL helpers shared by `query/` and `policy/`, kept in their own package to break an import cycle. `QuoteIdent` is the single place every identifier — column, table, alias — becomes SQL text: always backtick-quoted and escaped, so any ClickHouse-legal name (dots, spaces, unicode, keywords) is safe. `BindUnsafe` reports whether a name contains a literal `?`, which would desync clickhouse-go's positional binder; such names are rejected fail-closed rather than silently mis-bound. +### `keyenc/` — Key Escaping + +- **keyenc.go** — The one escaping every composite key is built from, so a name can never be mistaken for a separator: `Escape` keeps ASCII letters, digits and `_` and writes every other byte as `%XX` (uppercase hex), `Unescape` decodes `%XX` in either case and takes any other byte as itself, and `Join`/`Split` join escaped fields with a separator the escaping never emits (`Join` panics on one it could). NATS subject tokens (`internal/mq`) and the cache's namespace tokens (`query.SafeEncodeToken`) both use it; its output is byte-identical to the subject-token encoding v0.1.0 shipped, which queued messages depend on. + ## Data Flows ### Ingest Path diff --git a/internal/keyenc/keyenc.go b/internal/keyenc/keyenc.go new file mode 100644 index 000000000..ddda72cac --- /dev/null +++ b/internal/keyenc/keyenc.go @@ -0,0 +1,123 @@ +// Package keyenc is the one escaping every composite WaveHouse key is built +// from: NATS subject tokens, cache namespace tokens and dedupe keys. A field +// keeps ASCII letters, digits and '_' as they are and writes every other byte +// as %XX (uppercase hex), so no separator, wildcard, whitespace, brace or +// non-ASCII byte ever appears in it unescaped, and any table name ClickHouse +// accepts encodes. +// +// The output is pinned byte for byte: NATS subjects have carried it since +// v0.1.0, and queued messages outlive the binary that wrote them. +package keyenc + +import ( + "errors" + "fmt" + "strings" +) + +const upperHex = "0123456789ABCDEF" + +// kept reports whether b is written as itself. +func kept(b byte) bool { + return (b >= 'a' && b <= 'z') || (b >= 'A' && b <= 'Z') || (b >= '0' && b <= '9') || b == '_' +} + +// Escape encodes s as one field. +func Escape(s string) string { + for i := 0; i < len(s); i++ { + if !kept(s[i]) { + return string(AppendEscape(make([]byte, 0, len(s)+2*(len(s)-i)), s)) + } + } + return s +} + +// AppendEscape appends Escape(s) to dst. +func AppendEscape(dst []byte, s string) []byte { + for i := 0; i < len(s); i++ { + b := s[i] + if kept(b) { + dst = append(dst, b) + } else { + dst = append(dst, '%', upperHex[b>>4], upperHex[b&0x0F]) + } + } + return dst +} + +// ErrBadEscape is a '%' not followed by two hex digits. +var ErrBadEscape = errors.New("keyenc: malformed escape") + +// Unescape reverses Escape. It decodes %XX in either hex case and takes any +// other byte as itself — what url.PathUnescape accepts — so a field some +// other writer left partly unescaped still reads. +func Unescape(s string) (string, error) { + i := strings.IndexByte(s, '%') + if i < 0 { + return s, nil + } + out := make([]byte, 0, len(s)) + out = append(out, s[:i]...) + for ; i < len(s); i++ { + if s[i] != '%' { + out = append(out, s[i]) + continue + } + if i+2 >= len(s) { + return "", fmt.Errorf("%w in %q", ErrBadEscape, s) + } + hi, ok1 := unhex(s[i+1]) + lo, ok2 := unhex(s[i+2]) + if !ok1 || !ok2 { + return "", fmt.Errorf("%w in %q", ErrBadEscape, s) + } + out = append(out, hi<<4|lo) + i += 2 + } + return string(out), nil +} + +func unhex(c byte) (byte, bool) { + switch { + case c >= '0' && c <= '9': + return c - '0', true + case c >= 'a' && c <= 'f': + return c - 'a' + 10, true + case c >= 'A' && c <= 'F': + return c - 'A' + 10, true + } + return 0, false +} + +// Join escapes each field and joins them with sep. It panics if sep is a +// byte Escape keeps, since a key split on it could then not be told apart. +func Join(sep byte, fields ...string) string { + return string(AppendJoin(nil, sep, fields...)) +} + +// AppendJoin appends Join(sep, fields...) to dst. +func AppendJoin(dst []byte, sep byte, fields ...string) []byte { + if kept(sep) || sep == '%' { + panic(fmt.Sprintf("keyenc: %q cannot separate fields", sep)) + } + for i, f := range fields { + if i > 0 { + dst = append(dst, sep) + } + dst = AppendEscape(dst, f) + } + return dst +} + +// Split reverses Join: the fields of key, each unescaped. +func Split(key string, sep byte) ([]string, error) { + parts := strings.Split(key, string(sep)) + for i, p := range parts { + f, err := Unescape(p) + if err != nil { + return nil, err + } + parts[i] = f + } + return parts, nil +} diff --git a/internal/keyenc/keyenc_test.go b/internal/keyenc/keyenc_test.go new file mode 100644 index 000000000..7c178c86b --- /dev/null +++ b/internal/keyenc/keyenc_test.go @@ -0,0 +1,142 @@ +package keyenc_test + +import ( + "bytes" + "fmt" + "net/url" + "strings" + "testing" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + + "github.com/Wave-RF/WaveHouse/internal/keyenc" +) + +// v010Escape is the encoder v0.1.0 shipped as query.SafeEncodeNATS, copied +// verbatim: the reference Escape must match byte for byte. +func v010Escape(raw string) string { + var buf bytes.Buffer + for i := 0; i < len(raw); i++ { + b := raw[i] + if (b >= 'a' && b <= 'z') || (b >= 'A' && b <= 'Z') || (b >= '0' && b <= '9') || b == '_' { + buf.WriteByte(b) + } else { + fmt.Fprintf(&buf, "%%%02X", b) + } + } + return buf.String() +} + +func TestEscape_Golden(t *testing.T) { + t.Parallel() + for raw, want := range map[string]string{ + "": "", + "my_table123": "my_table123", + "default.clicks": "default%2Eclicks", + "my table": "my%20table", + "a-b/c": "a%2Db%2Fc", + "a.*.>": "a%2E%2A%2E%3E", + "100%": "100%25", + "{acme}:x|y": "%7Bacme%7D%3Ax%7Cy", + "a\x00b": "a%00b", + "café": "caf%C3%A9", + "\xff": "%FF", + "tab\tnewline\n": "tab%09newline%0A", + "#hash": "%23hash", + "ABCxyz_0189": "ABCxyz_0189", + "evt-123": "evt%2D123", + "日本": "%E6%97%A5%E6%9C%AC", + } { + assert.Equal(t, want, keyenc.Escape(raw), "%q", raw) + assert.Equal(t, want, string(keyenc.AppendEscape([]byte("x"), raw))[1:], "%q", raw) + } +} + +// Every byte value, alone and between kept bytes, encodes as v0.1.0 did. +func TestEscape_MatchesV010EveryByte(t *testing.T) { + t.Parallel() + for b := 0; b < 256; b++ { + for _, s := range []string{string([]byte{byte(b)}), "a" + string([]byte{byte(b)}) + "Z"} { + require.Equal(t, v010Escape(s), keyenc.Escape(s), "byte %#x", b) + } + } +} + +func TestEscape_NoAllocWhenNothingToEscape(t *testing.T) { + s := "events_2026" + assert.Zero(t, testing.AllocsPerRun(100, func() { _ = keyenc.Escape(s) })) +} + +func TestUnescape(t *testing.T) { + t.Parallel() + for in, want := range map[string]string{ + "": "", + "plain": "plain", + "default%2Eclicks": "default.clicks", + "lower%2ecase": "lower.case", + "left-as-is": "left-as-is", + "%00%FF": "\x00\xff", + "a+b": "a+b", + } { + got, err := keyenc.Unescape(in) + require.NoError(t, err, "%q", in) + assert.Equal(t, want, got, "%q", in) + } + for _, bad := range []string{"%", "%2", "a%2Gb", "%%41", "x%"} { + _, err := keyenc.Unescape(bad) + require.ErrorIs(t, err, keyenc.ErrBadEscape, "%q", bad) + } +} + +func TestJoinSplit(t *testing.T) { + t.Parallel() + assert.Equal(t, "acme/clicks/evt%2D123", keyenc.Join('/', "acme", "clicks", "evt-123")) + assert.Equal(t, "a%2Fb/c", keyenc.Join('/', "a/b", "c"), "a separator inside a field is escaped") + assert.Equal(t, "a..", keyenc.Join('.', "a", "", "")) + assert.Equal(t, "p:a", string(keyenc.AppendJoin([]byte("p:"), '/', "a"))) + + for _, fields := range [][]string{{"acme", "a/b", "id"}, {"", "", ""}, {"%", "/", "%2F"}, {"x"}} { + got, err := keyenc.Split(keyenc.Join('/', fields...), '/') + require.NoError(t, err) + assert.Equal(t, fields, got) + } + _, err := keyenc.Split("a/%zz", '/') + require.ErrorIs(t, err, keyenc.ErrBadEscape) + + for _, sep := range []byte{'a', 'Z', '5', '_', '%'} { + assert.Panics(t, func() { keyenc.Join(sep, "x") }, "%q", sep) + } +} + +// Unescape accepts exactly what url.PathUnescape did, so readers moved onto +// it decode every subject they decoded before. +func FuzzUnescapeMatchesPathUnescape(f *testing.F) { + for _, s := range []string{"", "a%2Eb", "%2e", "%", "%zz", "a+b", "%E6%97%A5", "100%25"} { + f.Add(s) + } + f.Fuzz(func(t *testing.T, s string) { + want, wantErr := url.PathUnescape(s) + got, err := keyenc.Unescape(s) + if wantErr != nil { + require.Error(t, err) + return + } + require.NoError(t, err) + require.Equal(t, want, got) + }) +} + +func FuzzEscapeRoundTrip(f *testing.F) { + for _, s := range []string{"", "a.b", "\x00", "café", "%", "a/b c"} { + f.Add(s) + } + f.Fuzz(func(t *testing.T, s string) { + enc := keyenc.Escape(s) + require.Equal(t, v010Escape(s), enc) + require.False(t, strings.ContainsAny(enc, "./:{}|# *>\x00"), enc) + dec, err := keyenc.Unescape(enc) + require.NoError(t, err) + require.Equal(t, s, dec) + }) +} diff --git a/internal/mq/mq.go b/internal/mq/mq.go index 3f1c45c1f..9198b17b4 100644 --- a/internal/mq/mq.go +++ b/internal/mq/mq.go @@ -13,6 +13,7 @@ import ( "errors" "time" + "github.com/Wave-RF/WaveHouse/internal/keyenc" "github.com/Wave-RF/WaveHouse/internal/observability" "github.com/Wave-RF/WaveHouse/internal/tenant" ) @@ -38,9 +39,9 @@ type Topic struct { // as its table (parseTopicKey's fallback). Callers key their own maps by the // Topic value itself. func (t Topic) key() string { - key := string(t.Tenant) + "." + encodeToken(t.Table) + key := string(t.Tenant) + "." + keyenc.Escape(t.Table) if t.Scope != "" { - key += "." + encodeToken(t.Scope) + key += "." + keyenc.Escape(t.Scope) } return key } diff --git a/internal/mq/subject.go b/internal/mq/subject.go index 489ad2414..2919464a1 100644 --- a/internal/mq/subject.go +++ b/internal/mq/subject.go @@ -1,11 +1,10 @@ package mq import ( - "bytes" "fmt" - "net/url" "strings" + "github.com/Wave-RF/WaveHouse/internal/keyenc" "github.com/Wave-RF/WaveHouse/internal/tenant" ) @@ -59,29 +58,6 @@ func streamTenant(prefix, name string) (tenant.ID, bool) { return id, err == nil } -// encodeToken converts any table or scope name into a safe, single NATS -// subject token. It preserves alphanumerics and underscores, but -// percent-encodes everything else (so '.', ' ', '*' and '>' can never split -// or wildcard a subject). -func encodeToken(raw string) string { - var buf bytes.Buffer - for i := 0; i < len(raw); i++ { - b := raw[i] - if (b >= 'a' && b <= 'z') || (b >= 'A' && b <= 'Z') || (b >= '0' && b <= '9') || b == '_' { - buf.WriteByte(b) - } else { - fmt.Fprintf(&buf, "%%%02X", b) - } - } - return buf.String() -} - -// decodeToken reverses encodeToken. url.PathUnescape handles exactly the %XX -// form encodeToken writes. -func decodeToken(safe string) (string, error) { - return url.PathUnescape(safe) -} - // subject renders a caller's topic under prefix. The tenant is checked // against its grammar here, on the way to the wire: an empty one — a caller // that never set it — must not become a subject of some other tenant's, and @@ -118,10 +94,10 @@ func parseTopicKey(tail string) Topic { switch len(parts) { case 2, 3: id, idErr := tenant.Parse(parts[0]) - table, tableErr := decodeToken(parts[1]) + table, tableErr := keyenc.Unescape(parts[1]) scope, scopeErr := "", error(nil) if len(parts) == 3 { - scope, scopeErr = decodeToken(parts[2]) + scope, scopeErr = keyenc.Unescape(parts[2]) } if idErr == nil && tableErr == nil && scopeErr == nil { return Topic{Tenant: id, Table: table, Scope: scope} diff --git a/internal/mq/subject_test.go b/internal/mq/subject_test.go index 67536e4e0..c2a314b41 100644 --- a/internal/mq/subject_test.go +++ b/internal/mq/subject_test.go @@ -9,56 +9,41 @@ import ( "github.com/stretchr/testify/require" ) -func TestEncodeToken(t *testing.T) { +// The subjects are pinned byte for byte: an embedded broker holds messages +// under them across an upgrade, and the table token has been this encoding +// since v0.1.0. +func TestSubject_Golden(t *testing.T) { t.Parallel() - tests := []struct { - name string - raw string - expected string + for _, tt := range []struct { + topic Topic + want string }{ - {"safe string", "my_table123", "my_table123"}, - {"with dots", "default.clicks", "default%2Eclicks"}, - {"with spaces", "my table", "my%20table"}, - {"with dashes and slashes", "a-b/c", "a%2Db%2Fc"}, - {"wildcards cannot survive", "a.*.>", "a%2E%2A%2E%3E"}, - {"empty string", "", ""}, - {"only safe characters", "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789_", "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789_"}, - } - - for _, tt := range tests { - t.Run(tt.name, func(t *testing.T) { - t.Parallel() - assert.Equal(t, tt.expected, encodeToken(tt.raw)) - }) + {Topic{Tenant: "0", Table: "events"}, "0.events"}, + {Topic{Tenant: "acme-co", Table: "default.clicks", Scope: "org_1"}, "acme-co.default%2Eclicks.org_1"}, + {Topic{Tenant: "a", Table: "a.*.>", Scope: "*"}, "a.a%2E%2A%2E%3E.%2A"}, + {Topic{Tenant: "a", Table: "my table", Scope: "tab\there"}, "a.my%20table.tab%09here"}, + {Topic{Tenant: "a", Table: "table-with-dashes", Scope: "org-1"}, "a.table%2Dwith%2Ddashes.org%2D1"}, + {Topic{Tenant: "a", Table: "100%", Scope: "a/b"}, "a.100%25.a%2Fb"}, + {Topic{Tenant: "a", Table: "{acme}:x"}, "a.%7Bacme%7D%3Ax"}, + {Topic{Tenant: "a", Table: "nul\x00", Scope: "\xff"}, "a.nul%00.%FF"}, + {Topic{Tenant: "a", Table: "caf\u00e9", Scope: "\u65e5"}, "a.caf%C3%A9.%E6%97%A5"}, + {Topic{Tenant: "a", Table: ""}, "a."}, + {Topic{Tenant: "a", Table: "t", Scope: ""}, "a.t"}, + } { + assert.Equal(t, tt.want, tt.topic.key(), "%+v", tt.topic) + for _, prefix := range []string{ingestPrefix, dlqPrefix} { + subj, err := subject(prefix, tt.topic) + require.NoError(t, err) + assert.Equal(t, prefix+tt.want, subj) + } } } -func TestDecodeToken(t *testing.T) { +// A token another writer left partly unescaped, or escaped in lowercase, +// still reads as it always did. +func TestParseTopicKey_LenientTokens(t *testing.T) { t.Parallel() - tests := []struct { - name string - safe string - expected string - wantErr bool - }{ - {"safe string", "my_table123", "my_table123", false}, - {"encoded dots", "default%2Eclicks", "default.clicks", false}, - {"encoded spaces", "my%20table", "my table", false}, - {"invalid percent encoding", "default%2Gclicks", "", true}, // %2G is not valid hex - } - - for _, tt := range tests { - t.Run(tt.name, func(t *testing.T) { - t.Parallel() - got, err := decodeToken(tt.safe) - if tt.wantErr { - assert.Error(t, err) - } else { - require.NoError(t, err) - assert.Equal(t, tt.expected, got) - } - }) - } + assert.Equal(t, Topic{Tenant: "a", Table: "b-c", Scope: "d.e"}, parseTopicKey("a.b-c.d%2ee")) } func TestSubject_RoundTripsEveryTopic(t *testing.T) { diff --git a/internal/query/ident.go b/internal/query/ident.go index 320fe097b..35f032da9 100644 --- a/internal/query/ident.go +++ b/internal/query/ident.go @@ -1,24 +1,8 @@ package query -import ( - "bytes" - "fmt" -) +import "github.com/Wave-RF/WaveHouse/internal/keyenc" -// SafeEncodeToken converts any table or scope name into a single dot-free -// token, for composing the cache's dotted namespace keys. It preserves -// alphanumerics and underscores, but percent-encodes everything else. -func SafeEncodeToken(raw string) string { - var buf bytes.Buffer - for i := 0; i < len(raw); i++ { - b := raw[i] - // Pass through safe characters: a-z, A-Z, 0-9, and _ (underscore) - if (b >= 'a' && b <= 'z') || (b >= 'A' && b <= 'Z') || (b >= '0' && b <= '9') || b == '_' { - buf.WriteByte(b) - } else { - // Hex encode everything else (e.g., '.' becomes '%2E', ' ' becomes '%20') - fmt.Fprintf(&buf, "%%%02X", b) - } - } - return buf.String() -} +// SafeEncodeToken renders a table or scope name as one dot-free token of the +// cache's namespace keys: keyenc's escaping, the same bytes a NATS subject +// carries for the name. +func SafeEncodeToken(raw string) string { return keyenc.Escape(raw) } From 51db11caa49196195e89a32a3a92fa3d94d08f53 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 13:10:02 -0400 Subject: [PATCH 18/79] docs(keyenc): name only the keys built from it today Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- internal/keyenc/keyenc.go | 12 ++++++------ 1 file changed, 6 insertions(+), 6 deletions(-) diff --git a/internal/keyenc/keyenc.go b/internal/keyenc/keyenc.go index ddda72cac..807cf4354 100644 --- a/internal/keyenc/keyenc.go +++ b/internal/keyenc/keyenc.go @@ -1,9 +1,9 @@ -// Package keyenc is the one escaping every composite WaveHouse key is built -// from: NATS subject tokens, cache namespace tokens and dedupe keys. A field -// keeps ASCII letters, digits and '_' as they are and writes every other byte -// as %XX (uppercase hex), so no separator, wildcard, whitespace, brace or -// non-ASCII byte ever appears in it unescaped, and any table name ClickHouse -// accepts encodes. +// Package keyenc is the one escaping composite WaveHouse keys are built +// from: NATS subject tokens and cache namespace tokens. A field keeps ASCII +// letters, digits and '_' as they are and writes every other byte as %XX +// (uppercase hex), so no separator, wildcard, whitespace, brace or non-ASCII +// byte ever appears in it unescaped, and any table name ClickHouse accepts +// encodes. // // The output is pinned byte for byte: NATS subjects have carried it since // v0.1.0, and queued messages outlive the binary that wrote them. From 27f760893f52012183723fb9d1b80ad7db4bc2f2 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 13:11:38 -0400 Subject: [PATCH 19/79] docs(architecture): keyenc is not yet every composite key's escaping Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- docs/src/content/docs/architecture.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index b987938c8..43b66bd24 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -203,7 +203,7 @@ The hot-reloadable half of configuration: a directory of four JSON files (`confi ### `keyenc/` — Key Escaping -- **keyenc.go** — The one escaping every composite key is built from, so a name can never be mistaken for a separator: `Escape` keeps ASCII letters, digits and `_` and writes every other byte as `%XX` (uppercase hex), `Unescape` decodes `%XX` in either case and takes any other byte as itself, and `Join`/`Split` join escaped fields with a separator the escaping never emits (`Join` panics on one it could). NATS subject tokens (`internal/mq`) and the cache's namespace tokens (`query.SafeEncodeToken`) both use it; its output is byte-identical to the subject-token encoding v0.1.0 shipped, which queued messages depend on. +- **keyenc.go** — The one escaping composite keys are built from, so a name can never be mistaken for a separator: `Escape` keeps ASCII letters, digits and `_` and writes every other byte as `%XX` (uppercase hex), `Unescape` decodes `%XX` in either case and takes any other byte as itself, and `Join`/`Split` join escaped fields with a separator the escaping never emits (`Join` panics on one it could). NATS subject tokens (`internal/mq`) and the cache's namespace tokens (`query.SafeEncodeToken`) both use it; its output is byte-identical to the subject-token encoding v0.1.0 shipped, which queued messages depend on. ## Data Flows From 2c31eff00e5ed4c507c514bf950e0b1057c1f594 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 19:30:42 -0400 Subject: [PATCH 20/79] refactor(keyenc): keep '-', join keys, count dead letters per table Escape now keeps '-' alongside [A-Za-z0-9_], exactly the tenant-id grammar, so a tenant id is its own escaped form and dashed names read as themselves. Unescape is url.PathUnescape, so v0.1.0's %2D still decodes. Join/AppendJoin/Split are fixed: Split converted a separator of 0x80+ as a rune, and Join of no fields could not be told from one empty field; both now refuse those inputs. Topic keys are built with AppendJoin and parsed with Split, the tenant token still verbatim. Dead-letter counts are keyed by table, every scope of a table summed under it, through one deadLetterTables function; the ?table= filter keeps all of a table's scopes. A scoped message used to count under "table.scope", which a dotted table name could share. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01B1tJWUp6oaDoH1usLwGtLF --- AGENTS.md | 4 +- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 5 +- internal/keyenc/keyenc.go | 77 +++++++------------ internal/keyenc/keyenc_test.go | 102 +++++++++++++++++--------- internal/mq/deadletter.go | 18 +++++ internal/mq/deadletter_test.go | 24 ++++++ internal/mq/embedded.go | 21 +----- internal/mq/mq.go | 24 +++--- internal/mq/subject.go | 37 +++++----- internal/mq/subject_test.go | 10 ++- internal/query/ident_test.go | 2 +- 12 files changed, 186 insertions(+), 140 deletions(-) create mode 100644 internal/mq/deadletter.go create mode 100644 internal/mq/deadletter_test.go diff --git a/AGENTS.md b/AGENTS.md index 30c808c90..bf8dc6a7f 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -38,8 +38,8 @@ Nineteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`dedupe/`** — `Deduplicator` interface → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) -- **`keyenc/`** — the one escaping composite keys are built from: `Escape` keeps `[A-Za-z0-9_]` and writes every other byte as `%XX`, `Unescape` reverses it leniently (as `url.PathUnescape` does), `Join`/`Split` join escaped fields with a separator the escaping never emits. NATS subject tokens and the cache's namespace tokens use it; its output is pinned byte for byte, since v0.1.0's subjects carry it -- **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, token encoding, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`); `internal/app` constructs it and hands everything else a `mq.Broker` +- **`keyenc/`** — the one escaping composite keys are built from: `Escape` keeps `[A-Za-z0-9_-]` (exactly the tenant-id grammar, so a tenant id is its own escaped form) and writes every other byte as `%XX`, `Unescape` is `url.PathUnescape` (lenient: either hex case, and a byte left unescaped reads as itself, so v0.1.0's `%2D` still reads), `Join`/`AppendJoin` escape each field and put a separator between them (they panic on no fields, and on a separator the escaping could write or one outside ASCII) and `Split` reverses them. NATS subjects (`Join`/`Split` after the verbatim tenant) and the cache's namespace tokens use it; changing what it keeps orphans every stored key +- **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`, `deadletter.go`), whose subject tokens are escaped by the shared `internal/keyenc`; `internal/app` constructs it and hands everything else a `mq.Broker` - **`observability/`** — OpenTelemetry pipeline: `InitProvider` wires trace/metric/log providers via OTLP gRPC (each signal independently gated). A top-level `Prometheus` config block drives an optional `/metrics` scrape endpoint that runs independently of OTLP push — standalone (Alloy/Mimir scrape, no collector), alongside OTLP, or off. `NewLogger` produces a slog handler that fans out to stdout AND OTLP (stdout always 100%, OTLP sample-rate-aware). `TraceHandler` injects trace_id/span_id from active spans. `tracer.go` provides W3C trace context propagation over message headers (`InjectHeaders`/`ExtractHeaders` on a plain header map; `internal/mq` injects on every publish and extracts on the `Subscribe` path, so this package never sees a NATS type). - **`pipes/`** — Named query pipes: `NamedQuery` type + `BindParams` + `Source` (read per request; `settings.Store` in production, `Static(q...)` in tests) - **`policy/`** — Hasura-style access control, **role-first**: `TablePolicy` is `map[string]RolePermissions`, and a role's grant splits by operation into `SelectPermissions` (columns, row `filter`, aggregations, the `max_*` limits) and `InsertPermissions` (columns, `check`) — so a field only one side honors does not exist on the other. `Evaluate()` resolves ONE operation and leaves the other side **nil** (`Select *ResolvedSelect` / `Insert *ResolvedInsert`), which every accessor fails closed on — nil is "not resolved", distinct from an empty side, which is "unrestricted" (what the admin return builds). Claim templating (`{{ jwt.claim.path }}`) resolves during that call. Policies come from `Source`, a `func() *Policy` read per call (`settings.Store.Policy` in production, `Static(p)` in tests) diff --git a/CHANGELOG.md b/CHANGELOG.md index 96c8081cf..6283af538 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -32,7 +32,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Changed -- **One escaping for composite keys** (`internal/keyenc` (new, + tests), `internal/mq/{mq,subject}.go` (+ tests), `internal/query/ident.go`, `AGENTS.md`, `docs/src/content/docs/architecture.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). NATS subject tokens and the cache's namespace tokens each carried a copy of the same encoder; both now call `internal/keyenc`, which the dedupe keys will use too. Behaviour-preserving: every subject and cache key is byte-identical (pinned by golden tests over dotted, wildcard, whitespace, `-`, `%`, `/`, brace, NUL and non-ASCII names, and against v0.1.0's encoder for every byte value), and a subject token decodes exactly as `url.PathUnescape` decoded it. +- **One escaping for composite keys, `-` kept; dead-letter counts per table** (`internal/keyenc` (new, + tests), `internal/mq/{mq,subject,embedded,deadletter}.go` (+ tests), `internal/query/ident.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/architecture.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). NATS subject tokens and the cache's namespace tokens each carried a copy of the same encoder; both now call `internal/keyenc`, which the dedupe keys will use too, and subjects are built with its `Join` (each field escaped, with a separator the escaping never writes between them). The escaping now keeps `-` as well as ASCII letters, digits and `_` — exactly the tenant-id grammar — so a table or scope such as `my-table` is `my-table` in a subject rather than `my%2Dtable`; every other byte is escaped as before (pinned by golden tests, and against v0.1.0's encoder for every other byte value). Decoding is `url.PathUnescape`, as it was, so a subject written with v0.1.0's `%2D` reads as the same topic and a dead-letter count merges both forms. One thing notices the change: a `/v1/stream` client resuming across the upgrade (`Last-Event-ID` or `since`) on a table whose name holds `-` misses that table's events queued before the upgrade, since the replay filters on the table's exact subject. `GET /v1/ops/dlq/stats` now counts every scope of a table under the table itself, and `?table=` keeps all of its scopes; a scoped message used to count under `table.scope`, a name a dotted table could share. Scope is always empty today, so the response is unchanged. - **Each tenant has a message queue of its own** (`internal/mq/{mq,subject,embedded}.go` (+ tests), `internal/ingest/{sweeper,worker}.go` (+ tests), `internal/api/{dlq,ingest}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `internal/stream/{subscriber,hub}.go`, `internal/settings/{settings,store,registry}.go` (+ tests), `internal/testutil/{mocks,testutil}.go`, `clients/ts/src/{dlq,types}.ts` (+ tests), `docs/src/content/docs/{deployment,api,architecture,ingest-pipeline,durability,why-wavehouse}.md`, `docs/src/content/docs/sdk/{admin,reference,streaming}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `AGENTS.md`): story 5b of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). The embedded NATS server keeps each tenant's events on a pair of JetStream streams of its own — `INGEST_` (`ingest..>`, `DiscardNew`) at the tenant's own `mq.max_bytes_gb`, and `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed; subjects are unchanged, and nothing outside `internal/mq` names a stream. A tenant at its budget gets `503` while the others keep publishing, and the ingest worker's and the hub bridge's durables are held on every tenant's stream, each with its own ack floor and `MaxAckPending`, so one tenant's backlog holds back neither another's delivery nor its purge; the worker's prefetch and the hub bridge's fetch-ahead are each shared across the tenants' streams. A removed or rejected tenant's stream is still consumed, so its queued rows reach the worker and are parked on its own dead-letter queue. The sweeper purges each tenant's stream at that tenant's own `stream.gap_window_minutes` — a rejected tenant's as its folder last had it (`settings.Registry.Known` now yields each tenant's last adopted store), and all of its history if the folder has been rejected since boot, so its clients resume once the folder is fixed — keeping no acknowledged history for a removed tenant, and every served tenant's `mq.max_bytes_gb` is applied after each reload rather than tenant `0`'s alone (the boot warning about a nested directory with no tenant `0` is gone with it, and so is the tracking of tenant `0`'s last adopted store: a flat directory's ops gate, its one reader left, reads tenant `0` through the registry). One tenant's failed purge holds up no other tenant's, and the sweep logs it at `ERROR` unless every failure in it is a buffer consumer not created yet. A reload that shrinks a budget no longer deletes dead letters: a dead-letter stream holding more than a tenth of the new budget keeps what it holds, and that is logged — the interim guard of [#532](https://github.com/Wave-RF/WaveHouse/issues/532). JetStream's own check of the streams' caps against the disk, which made them fit 75% of the free disk at boot together, is lifted: a budget is a cap and never a reservation, what the budgets add up to against the disk is [#138](https://github.com/Wave-RF/WaveHouse/issues/138)'s, and a flat directory whose budget exceeds three quarters of the free disk now boots where it used to be refused. A queue that cannot be opened refuses a flat boot like any other store and, over a nested directory, costs its tenant alone — its ingest answers `503`, each reload trying again, and a publish at most once every five seconds, so its clients retrying never hold up another tenant's queue. `GET /v1/ops/dlq/stats` reads one tenant's dead-letter queue — the one `?tenant=` names, parsed strictly like the other admin reads, and tenant `0`'s without it, no longer the sum across tenants — answering for a rejected or removed tenant too and `404` for a tenant with no queue; the SDK's `wh.dlq.list()` and `.table()` take a `tenant` option. The streams an earlier build kept for every tenant together (`WAVEHOUSE`, `WAVEHOUSE_DLQ`) overlap every tenant's subjects and are deleted at boot, what they held with them, and a subject with no tenant token no longer reads as tenant `0`'s. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 43b66bd24..59b2fe3fe 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -147,7 +147,8 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ The **only** package that imports NATS/JetStream — a `depguard` rule in `.golangci.yml` fails `make lint` on any `github.com/nats-io` import in every package golangci-lint builds; the `integration`-tagged files under `tests/` sit outside its default build context, so the boundary there rests on convention (AGENTS.md Key Design Decision #20). Every other package talks to the broker through the types below, so a subject, stream, or broker change lands here once. - **mq.go** — The owned surface, stated as intent rather than broker mechanics. `Topic{Tenant, Table, Scope}` is the only address the rest of the process handles (a validated tenant id and raw names; comparable, so the SSE hub keys its index by the value). `Message` carries `Data`, its topic (`Topic()` decodes the delivered key on demand — the tenant included, which is how the hub bridge and the worker learn whose event it is; `TopicKey()` is the delivered form, for log lines), and the ack family (`DoubleAck(ctx)`, `Ack()`, `Nak()`); `Headers` is the message header map (`Add`/`Set`/`Get`, exact-key) that `PublishOpt`s such as `WithHeader` shape. Interfaces, each speaking per tenant and never per stream: `Publisher` (`ErrQueueFull` when the tenant's ingest queue is at its byte budget or not open yet — the API's 503 + `Retry-After`), `Subscriber` (every ingest event of every tenant, under a named durable consumer, its fetch-ahead split across the tenants — the hub bridge), `ConsumerManager` → `Consumer` (a durable explicit-ack consumer from a `ConsumerConfig`, whose `MaxAckPending` holds per tenant; `Consume` delivers each tenant's messages on a goroutine of that tenant's, in order, so a blocking handler is backpressure on its own tenant alone, spreads the prefetch across the tenants, and returns a `stop` plus a `failed` channel that reports delivery ending on its own — `ErrDeliveryEnded`, e.g. a deleted consumer, a closed connection, or a tenant's queue that could not be joined — since no message would ever say so) for the ingest worker, `DeadLetterer.DeadLetter` (park a message under its own topic, in its tenant's dead-letter queue; the caller acks), `DeadLetterStats.DeadLetterCounts` (one tenant's; `ErrNoDeadLetterQueue` when it has none), `Purger.PurgeAcked` (drop what is both acked by a consumer and stored before its tenant's cutoff, everything acked for a tenant given none; one error per failed tenant, joined — `ErrConsumerNotFound` for a queue the consumer has not been created on yet, the one failure the sweeper logs as a warning rather than an error) for the sweeper, and `Replayer.ReplaySince` for SSE gap-fill. `Broker` composes them with each tenant's byte budget (`SetMaxBytes`/`MaxBytes`), `Stats`, and `Close`; it is what `internal/app` holds. -- **subject.go** — The embedded broker's naming, private to the package: the stream names (`INGEST_` and `DLQ_` — prefixes that differ in their first letter, so no tenant id makes one kind's name the other's — and the one pair an earlier build kept for every tenant together, `WAVEHOUSE`/`WAVEHOUSE_DLQ`, which boot deletes), the `ingest.`/`dlq.` prefixes and `>` wildcards, the subject tokens (`internal/keyenc`: alphanumerics and `_` pass, everything else is percent-encoded, so a name can never split or wildcard a subject), and `Topic` ↔ subject conversion. A subject is `.
[.]`: the tenant verbatim — its grammar (`tenant.Parse`) makes it one token, and it is checked on the way to the wire, so a topic without one has no subject — then the table and scope as encoded tokens; tenant first so one wildcard selects a tenant's traffic (`ingest.acme.>`). A topic has the same tail on both streams, so parking on the DLQ is a prefix swap on the delivered subject — nothing is decoded or re-encoded — and the tail's first token picks the tenant's stream. +- **subject.go** — The embedded broker's naming, private to the package: the stream names (`INGEST_` and `DLQ_` — prefixes that differ in their first letter, so no tenant id makes one kind's name the other's — and the one pair an earlier build kept for every tenant together, `WAVEHOUSE`/`WAVEHOUSE_DLQ`, which boot deletes), the `ingest.`/`dlq.` prefixes and `>` wildcards, the subject tokens (`internal/keyenc`: ASCII letters, digits, `_` and `-` pass, everything else is percent-encoded, so a name can never split or wildcard a subject), and `Topic` ↔ subject conversion. A subject is `.
[.]`: the tenant verbatim — its grammar (`tenant.Parse`) makes it one token, and it is checked on the way to the wire, so a topic without one has no subject — then the table and scope as encoded tokens; tenant first so one wildcard selects a tenant's traffic (`ingest.acme.>`). A topic has the same tail on both streams, so parking on the DLQ is a prefix swap on the delivered subject — nothing is decoded or re-encoded — and the tail's first token picks the tenant's stream. +- **deadletter.go** — `deadLetterTables`, the per-table count `DeadLetterCounts` reports: a dead-letter stream's per-subject counts, each subject parsed back to its topic and counted under its table — every scope of a table under the table itself, so a dotted table name never shares a count with a table + scope pair — and a table filter keeps that table with all of its scopes. - **purge.go** — The Active Sweeper's arithmetic over JetStream sequences: purge target = `MIN(consumer ack floor + 1, first sequence stored at or after the cutoff)`, the latter found by binary search over message timestamps (~15 lookups). Every uncertainty resolves toward purging less: a sequence that holds no message is kept as a candidate bound rather than discarding the half below it, and a lookup that fails outright aborts that tenant's purge. It runs on each tenant's stream at that tenant's cutoff, and a failure on one tenant's stream is reported without stopping the sweep of the others. Healthy state keeps exactly the gap window; ClickHouse down freezes purging; a catastrophic outage fills the stream to `MaxBytes` and `DiscardNew` pushes back. - **embedded.go** — `EmbeddedNATS`, the one `Broker`: an in-process NATS server with JetStream, giving each tenant a queue of its own — stream `INGEST_` with subjects `ingest..>`, capped at the tenant's `mq.max_bytes_gb` (`DiscardNew`), and stream `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — with the durable consumers on the ingest one; nothing outside the package sees that layout. JetStream's own check of the streams' caps against the disk (75% of the free disk by default) is set out of reach, so a budget is a cap and never a reservation. Boot deletes the pair an earlier build kept for every tenant together — its subjects overlap every tenant's — and takes stock of the tenants' streams on disk with their budgets, so a consumer created later is held on every one, a tenant no longer served included. `SetMaxBytes` opens a tenant's queue the first time — its dead-letter stream first, so no row is queued that could not be parked — and every registered consumer joins it; a publish or park that finds a stream missing reopens it at the budget last asked for the tenant, and so does a publish to a queue the broker has not recorded open — an open that timed out can leave a stream JetStream creates after all, which no consumer holds — either refused as a full queue with none asked yet. Publishes and parks that find the same queue not open share one attempt (`singleflight`), and after one fails the tenant's publishes and parks are refused at once for five seconds rather than each trying again under the broker's lock, which every tenant's open, resize and reload takes; a reload retries regardless. After that `SetMaxBytes` applies a reloaded budget to the tenant's two live streams as a pair: if the DLQ update fails after the ingest one succeeded, the ingest resize is undone so both stay on the previous budget — best effort, since if that undo also fails the ingest stream stays at the new limit and the DLQ at the previous, and the error says so. A dead-letter stream is never capped below the bytes it holds, which `DiscardOld` would delete to fit ([#532](https://github.com/Wave-RF/WaveHouse/issues/532)): it keeps what it holds, and that is logged. Its JetStream calls are bounded to ten seconds — plus ten more for the consumers joining a queue it has just opened, and five for the rollback of a failed resize, each a budget of its own rather than the one that just expired — since a reload holds the settings store's lock while its hooks run; `MaxBytes` reports the budget last applied in full, so a failed resize is retried by the next reload. The consumers `CreateConsumer` and `Subscribe` build hold one durable on each tenant's stream, looked up before anything is written so a boot over many queues writes nothing it need not, each delivering on a goroutine of its own into the one handler. `PurgeAcked` and `DeadLetterCounts` run per tenant stream. `Stats` reports connection and inbound-message counters for `observability.RegisterSystemMetrics`. Trace context rides in the message headers: `Publish` applies `observability.InjectHeaders`, and a message delivered through `Subscribe` (the hub bridge) carries `observability.ExtractHeaders` on its `Ctx`; the worker's `Consumer` path skips the extraction, since it batches across messages and reads no per-message context. @@ -203,7 +204,7 @@ The hot-reloadable half of configuration: a directory of four JSON files (`confi ### `keyenc/` — Key Escaping -- **keyenc.go** — The one escaping composite keys are built from, so a name can never be mistaken for a separator: `Escape` keeps ASCII letters, digits and `_` and writes every other byte as `%XX` (uppercase hex), `Unescape` decodes `%XX` in either case and takes any other byte as itself, and `Join`/`Split` join escaped fields with a separator the escaping never emits (`Join` panics on one it could). NATS subject tokens (`internal/mq`) and the cache's namespace tokens (`query.SafeEncodeToken`) both use it; its output is byte-identical to the subject-token encoding v0.1.0 shipped, which queued messages depend on. +- **keyenc.go** — The one escaping composite keys are built from, so a name can never be mistaken for a separator: `Escape` keeps ASCII letters, digits, `_` and `-` — exactly the tenant-id grammar, so a tenant id is its own escaped form — and writes every other byte as `%XX` (uppercase hex); `Unescape` is `url.PathUnescape`, which decodes `%XX` in either case and takes any other byte as itself, so v0.1.0's `%2D` for `-` still reads. `Join`/`AppendJoin` escape each field and put a separator between them, panicking on no fields and on a separator the escaping could write or one outside ASCII, and `Split` reverses them. NATS subject tokens (`internal/mq`) and the cache's namespace tokens (`query.SafeEncodeToken`) both use it. Keys built from it are stored, so changing what it keeps orphans them. ## Data Flows diff --git a/internal/keyenc/keyenc.go b/internal/keyenc/keyenc.go index 807cf4354..59bac7e1c 100644 --- a/internal/keyenc/keyenc.go +++ b/internal/keyenc/keyenc.go @@ -1,17 +1,19 @@ // Package keyenc is the one escaping composite WaveHouse keys are built // from: NATS subject tokens and cache namespace tokens. A field keeps ASCII -// letters, digits and '_' as they are and writes every other byte as %XX +// letters, digits, '_' and '-' as they are and writes every other byte as %XX // (uppercase hex), so no separator, wildcard, whitespace, brace or non-ASCII // byte ever appears in it unescaped, and any table name ClickHouse accepts -// encodes. +// encodes. The bytes it keeps are exactly a tenant id's (tenant.Parse), so a +// tenant id is its own escaped form. // -// The output is pinned byte for byte: NATS subjects have carried it since -// v0.1.0, and queued messages outlive the binary that wrote them. +// Keys built from it are stored — queued under NATS subjects, held in caches +// — so a change to what it keeps orphans them. v0.1.0 escaped '-' as %2D; +// Unescape still reads that form. package keyenc import ( - "errors" "fmt" + "net/url" "strings" ) @@ -19,7 +21,7 @@ const upperHex = "0123456789ABCDEF" // kept reports whether b is written as itself. func kept(b byte) bool { - return (b >= 'a' && b <= 'z') || (b >= 'A' && b <= 'Z') || (b >= '0' && b <= '9') || b == '_' + return (b >= 'a' && b <= 'z') || (b >= 'A' && b <= 'Z') || (b >= '0' && b <= '9') || b == '_' || b == '-' } // Escape encodes s as one field. @@ -45,60 +47,33 @@ func AppendEscape(dst []byte, s string) []byte { return dst } -// ErrBadEscape is a '%' not followed by two hex digits. -var ErrBadEscape = errors.New("keyenc: malformed escape") - -// Unescape reverses Escape. It decodes %XX in either hex case and takes any -// other byte as itself — what url.PathUnescape accepts — so a field some -// other writer left partly unescaped still reads. +// Unescape reverses Escape. It is url.PathUnescape: %XX in either hex case +// decodes, and any other byte reads as itself, so a field another writer +// left partly unescaped — v0.1.0's %2D included — still reads. func Unescape(s string) (string, error) { - i := strings.IndexByte(s, '%') - if i < 0 { - return s, nil - } - out := make([]byte, 0, len(s)) - out = append(out, s[:i]...) - for ; i < len(s); i++ { - if s[i] != '%' { - out = append(out, s[i]) - continue - } - if i+2 >= len(s) { - return "", fmt.Errorf("%w in %q", ErrBadEscape, s) - } - hi, ok1 := unhex(s[i+1]) - lo, ok2 := unhex(s[i+2]) - if !ok1 || !ok2 { - return "", fmt.Errorf("%w in %q", ErrBadEscape, s) - } - out = append(out, hi<<4|lo) - i += 2 - } - return string(out), nil + return url.PathUnescape(s) } -func unhex(c byte) (byte, bool) { - switch { - case c >= '0' && c <= '9': - return c - '0', true - case c >= 'a' && c <= 'f': - return c - 'a' + 10, true - case c >= 'A' && c <= 'F': - return c - 'A' + 10, true +// checkSep panics unless sep can separate escaped fields: a byte Escape never +// writes, and ASCII, so the key stays valid UTF-8. +func checkSep(sep byte) { + if kept(sep) || sep == '%' || sep >= 0x80 { + panic(fmt.Sprintf("keyenc: %q cannot separate fields", sep)) } - return 0, false } -// Join escapes each field and joins them with sep. It panics if sep is a -// byte Escape keeps, since a key split on it could then not be told apart. +// Join escapes each field and joins them with sep. It panics on no fields, +// whose key would be one empty field's, and on a separator Escape could +// write. func Join(sep byte, fields ...string) string { return string(AppendJoin(nil, sep, fields...)) } // AppendJoin appends Join(sep, fields...) to dst. func AppendJoin(dst []byte, sep byte, fields ...string) []byte { - if kept(sep) || sep == '%' { - panic(fmt.Sprintf("keyenc: %q cannot separate fields", sep)) + checkSep(sep) + if len(fields) == 0 { + panic("keyenc: Join needs at least one field") } for i, f := range fields { if i > 0 { @@ -109,9 +84,11 @@ func AppendJoin(dst []byte, sep byte, fields ...string) []byte { return dst } -// Split reverses Join: the fields of key, each unescaped. +// Split reverses Join: the fields of key, each unescaped. It panics on a +// separator Join would refuse. func Split(key string, sep byte) ([]string, error) { - parts := strings.Split(key, string(sep)) + checkSep(sep) + parts := strings.Split(key, string([]byte{sep})) for i, p := range parts { f, err := Unescape(p) if err != nil { diff --git a/internal/keyenc/keyenc_test.go b/internal/keyenc/keyenc_test.go index 7c178c86b..5b1630d6f 100644 --- a/internal/keyenc/keyenc_test.go +++ b/internal/keyenc/keyenc_test.go @@ -3,7 +3,6 @@ package keyenc_test import ( "bytes" "fmt" - "net/url" "strings" "testing" @@ -11,10 +10,11 @@ import ( "github.com/stretchr/testify/require" "github.com/Wave-RF/WaveHouse/internal/keyenc" + "github.com/Wave-RF/WaveHouse/internal/tenant" ) // v010Escape is the encoder v0.1.0 shipped as query.SafeEncodeNATS, copied -// verbatim: the reference Escape must match byte for byte. +// verbatim. Escape differs from it only in keeping '-'. func v010Escape(raw string) string { var buf bytes.Buffer for i := 0; i < len(raw); i++ { @@ -35,7 +35,7 @@ func TestEscape_Golden(t *testing.T) { "my_table123": "my_table123", "default.clicks": "default%2Eclicks", "my table": "my%20table", - "a-b/c": "a%2Db%2Fc", + "a-b/c": "a-b%2Fc", "a.*.>": "a%2E%2A%2E%3E", "100%": "100%25", "{acme}:x|y": "%7Bacme%7D%3Ax%7Cy", @@ -45,7 +45,7 @@ func TestEscape_Golden(t *testing.T) { "tab\tnewline\n": "tab%09newline%0A", "#hash": "%23hash", "ABCxyz_0189": "ABCxyz_0189", - "evt-123": "evt%2D123", + "evt-123": "evt-123", "日本": "%E6%97%A5%E6%9C%AC", } { assert.Equal(t, want, keyenc.Escape(raw), "%q", raw) @@ -53,21 +53,52 @@ func TestEscape_Golden(t *testing.T) { } } -// Every byte value, alone and between kept bytes, encodes as v0.1.0 did. -func TestEscape_MatchesV010EveryByte(t *testing.T) { +// Every byte value, alone and between kept bytes, encodes as v0.1.0 did, +// but for '-'. +func TestEscape_MatchesV010ButDash(t *testing.T) { t.Parallel() - for b := 0; b < 256; b++ { + for b := range 256 { for _, s := range []string{string([]byte{byte(b)}), "a" + string([]byte{byte(b)}) + "Z"} { - require.Equal(t, v010Escape(s), keyenc.Escape(s), "byte %#x", b) + want := v010Escape(s) + if b == '-' { + want = s + } + require.Equal(t, want, keyenc.Escape(s), "byte %#x", b) } } } +// A tenant id is its own escaped form: the kept bytes are its grammar. +func TestEscape_KeepsExactlyTheTenantGrammar(t *testing.T) { + t.Parallel() + for b := range 256 { + s := string([]byte{byte(b)}) + _, err := tenant.Parse(s) + assert.Equal(t, err == nil, keyenc.Escape(s) == s, "byte %#x", b) + } +} + func TestEscape_NoAllocWhenNothingToEscape(t *testing.T) { - s := "events_2026" + s := "events_2026-09" assert.Zero(t, testing.AllocsPerRun(100, func() { _ = keyenc.Escape(s) })) } +// Distinct names never share an escaped form, even names that look escaped: +// '%' is itself escaped. +func TestEscape_LookalikesStayDistinct(t *testing.T) { + t.Parallel() + names := []string{"b-c", "b%2Dc", "b%2dc", "b.c", "b%2Ec"} + seen := map[string]string{} + for _, n := range names { + e := keyenc.Escape(n) + require.NotContains(t, seen, e, "%q and %q", seen[e], n) + seen[e] = n + back, err := keyenc.Unescape(e) + require.NoError(t, err) + assert.Equal(t, n, back) + } +} + func TestUnescape(t *testing.T) { t.Parallel() for in, want := range map[string]string{ @@ -75,7 +106,7 @@ func TestUnescape(t *testing.T) { "plain": "plain", "default%2Eclicks": "default.clicks", "lower%2ecase": "lower.case", - "left-as-is": "left-as-is", + "evt%2D123": "evt-123", // v0.1.0's form "%00%FF": "\x00\xff", "a+b": "a+b", } { @@ -85,55 +116,58 @@ func TestUnescape(t *testing.T) { } for _, bad := range []string{"%", "%2", "a%2Gb", "%%41", "x%"} { _, err := keyenc.Unescape(bad) - require.ErrorIs(t, err, keyenc.ErrBadEscape, "%q", bad) + require.Error(t, err, "%q", bad) } } func TestJoinSplit(t *testing.T) { t.Parallel() - assert.Equal(t, "acme/clicks/evt%2D123", keyenc.Join('/', "acme", "clicks", "evt-123")) + assert.Equal(t, "acme/clicks/evt-123", keyenc.Join('/', "acme", "clicks", "evt-123")) assert.Equal(t, "a%2Fb/c", keyenc.Join('/', "a/b", "c"), "a separator inside a field is escaped") assert.Equal(t, "a..", keyenc.Join('.', "a", "", "")) + assert.Equal(t, "", keyenc.Join('.', ""), "one empty field") assert.Equal(t, "p:a", string(keyenc.AppendJoin([]byte("p:"), '/', "a"))) - for _, fields := range [][]string{{"acme", "a/b", "id"}, {"", "", ""}, {"%", "/", "%2F"}, {"x"}} { + for _, fields := range [][]string{{"acme", "a/b", "id"}, {"", "", ""}, {""}, {"%", "/", "%2F"}, {"x"}} { got, err := keyenc.Split(keyenc.Join('/', fields...), '/') require.NoError(t, err) assert.Equal(t, fields, got) } _, err := keyenc.Split("a/%zz", '/') - require.ErrorIs(t, err, keyenc.ErrBadEscape) + require.Error(t, err) +} - for _, sep := range []byte{'a', 'Z', '5', '_', '%'} { - assert.Panics(t, func() { keyenc.Join(sep, "x") }, "%q", sep) - } +func TestJoin_RefusesZeroFields(t *testing.T) { + t.Parallel() + assert.Panics(t, func() { keyenc.Join('/') }) + assert.Panics(t, func() { keyenc.AppendJoin(nil, '/') }) } -// Unescape accepts exactly what url.PathUnescape did, so readers moved onto -// it decode every subject they decoded before. -func FuzzUnescapeMatchesPathUnescape(f *testing.F) { - for _, s := range []string{"", "a%2Eb", "%2e", "%", "%zz", "a+b", "%E6%97%A5", "100%25"} { - f.Add(s) - } - f.Fuzz(func(t *testing.T, s string) { - want, wantErr := url.PathUnescape(s) - got, err := keyenc.Unescape(s) - if wantErr != nil { - require.Error(t, err) - return +// A separator Escape could write, or one outside ASCII, is refused by Join +// and Split alike; every other byte separates. +func TestSeparators(t *testing.T) { + t.Parallel() + for b := range 256 { + sep := byte(b) + if keyenc.Escape(string([]byte{sep})) == string([]byte{sep}) || sep == '%' || sep >= 0x80 { + assert.Panics(t, func() { keyenc.Join(sep, "x") }, "%#x", b) + assert.Panics(t, func() { _, _ = keyenc.Split("x", sep) }, "%#x", b) + continue } - require.NoError(t, err) - require.Equal(t, want, got) - }) + fields := []string{"a", string([]byte{sep}), "b" + string([]byte{sep}) + "c"} + got, err := keyenc.Split(keyenc.Join(sep, fields...), sep) + require.NoError(t, err, "%#x", b) + assert.Equal(t, fields, got, "%#x", b) + } } func FuzzEscapeRoundTrip(f *testing.F) { - for _, s := range []string{"", "a.b", "\x00", "café", "%", "a/b c"} { + for _, s := range []string{"", "a.b", "\x00", "café", "%", "a/b c", "b-c", "b%2Dc"} { f.Add(s) } f.Fuzz(func(t *testing.T, s string) { enc := keyenc.Escape(s) - require.Equal(t, v010Escape(s), enc) + require.Equal(t, strings.ReplaceAll(v010Escape(s), "%2D", "-"), enc) require.False(t, strings.ContainsAny(enc, "./:{}|# *>\x00"), enc) dec, err := keyenc.Unescape(enc) require.NoError(t, err) diff --git a/internal/mq/deadletter.go b/internal/mq/deadletter.go new file mode 100644 index 000000000..e6cc98e11 --- /dev/null +++ b/internal/mq/deadletter.go @@ -0,0 +1,18 @@ +package mq + +// deadLetterTables counts a dead-letter stream's parked messages per table, +// from its per-subject counts under prefix. Every scope of a table counts +// under the table itself, so no table + scope pair can share a count with a +// dotted table name. A non-empty table keeps that table alone, all of its +// scopes included. +func deadLetterTables(subjects map[string]uint64, prefix, table string) map[string]uint64 { + tables := make(map[string]uint64, len(subjects)) + for subj, n := range subjects { + t := parseTopicKey(topicKey(prefix, subj)) + if table != "" && t.Table != table { + continue + } + tables[t.Table] += n + } + return tables +} diff --git a/internal/mq/deadletter_test.go b/internal/mq/deadletter_test.go new file mode 100644 index 000000000..aaed64fef --- /dev/null +++ b/internal/mq/deadletter_test.go @@ -0,0 +1,24 @@ +package mq + +import ( + "testing" + + "github.com/stretchr/testify/assert" +) + +func TestDeadLetterTables(t *testing.T) { + t.Parallel() + subjects := map[string]uint64{ + "dlq.0.a%2Eb": 1, // table "a.b" + "dlq.0.a.b": 2, // table "a", scope "b" + "dlq.0.a": 4, // table "a", unscoped + "dlq.0.my-t": 8, + "dlq.0.my%2Dt.org-1": 16, // v0.1.0's escaping of '-', scoped + "dlq.0.clicks.org%2E": 32, + } + assert.Equal(t, map[string]uint64{"a.b": 1, "a": 6, "my-t": 24, "clicks": 32}, + deadLetterTables(subjects, dlqPrefix, ""), "a dotted table never shares a count with a table + scope") + assert.Equal(t, map[string]uint64{"a": 6}, deadLetterTables(subjects, dlqPrefix, "a"), "the filter keeps every scope of its table") + assert.Equal(t, map[string]uint64{"a.b": 1}, deadLetterTables(subjects, dlqPrefix, "a.b")) + assert.Empty(t, deadLetterTables(subjects, dlqPrefix, "never_failed")) +} diff --git a/internal/mq/embedded.go b/internal/mq/embedded.go index 989714920..f314840dd 100644 --- a/internal/mq/embedded.go +++ b/internal/mq/embedded.go @@ -1003,9 +1003,9 @@ func (e *EmbeddedNATS) PurgeAcked(ctx context.Context, consumer string, olderTha } // DeadLetterCounts reads tenant id's dead-letter stream's per-subject counts -// and keys them by table. The table filter matches that table's unscoped -// subject, so it is applied to the parsed topic rather than as a subject -// filter; a scoped topic counts under "table.scope". +// and keys them by table (deadLetterTables). The table filter matches every +// scope of that table, so it is applied to the parsed topic rather than as a +// subject filter. func (e *EmbeddedNATS) DeadLetterCounts(ctx context.Context, id tenant.ID, table string) (DeadLetterCounts, error) { if _, err := tenant.Parse(string(id)); err != nil { return DeadLetterCounts{}, fmt.Errorf("tenant: %w", err) @@ -1023,20 +1023,7 @@ func (e *EmbeddedNATS) DeadLetterCounts(ctx context.Context, id tenant.ID, table return DeadLetterCounts{}, fmt.Errorf("dlq stream info: %w", err) } - counts := DeadLetterCounts{Tables: make(map[string]uint64, len(state.Subjects)), Total: state.Msgs} - for subj, n := range state.Subjects { - t := parseTopicKey(topicKey(dlqPrefix, subj)) - if table != "" && (t.Table != table || t.Scope != "") { - continue - } - name := t.Table - if t.Scope != "" { - // TODO(#235): break scopes out rather than fold them into the name. - name += "." + t.Scope - } - counts.Tables[name] += n - } - return counts, nil + return DeadLetterCounts{Tables: deadLetterTables(state.Subjects, dlqPrefix, table), Total: state.Msgs}, nil } // ReplaySince creates an ephemeral consumer on topic's ingest subject, in its diff --git a/internal/mq/mq.go b/internal/mq/mq.go index 9198b17b4..57eae51c6 100644 --- a/internal/mq/mq.go +++ b/internal/mq/mq.go @@ -34,16 +34,16 @@ type Topic struct { // key is the injective string form of the topic that a subject's tail // carries: the tenant first, verbatim — its grammar makes it one token — then -// the table and scope as encoded tokens. A topic without a tenant has no -// subject, and its key parses back to a topic of no tenant with the whole key -// as its table (parseTopicKey's fallback). Callers key their own maps by the -// Topic value itself. +// the table and scope joined as escaped tokens (keyenc.AppendJoin). A topic +// without a tenant has no subject, and its key parses back to a topic of no +// tenant with the whole key as its table (parseTopicKey's fallback). Callers +// key their own maps by the Topic value itself. func (t Topic) key() string { - key := string(t.Tenant) + "." + keyenc.Escape(t.Table) - if t.Scope != "" { - key += "." + keyenc.Escape(t.Scope) + key := append([]byte(t.Tenant), '.') + if t.Scope == "" { + return string(keyenc.AppendJoin(key, '.', t.Table)) } - return key + return string(keyenc.AppendJoin(key, '.', t.Table, t.Scope)) } // Message represents a message received from the queue. @@ -247,8 +247,8 @@ type DeadLetterer interface { // DeadLetterCounts is what is parked on one tenant's dead-letter queue. type DeadLetterCounts struct { // Tables maps table name → parked messages, for the tables asked about. - // Scope is not broken out yet (it is inert until #235): a message parked - // under a scoped topic counts under "table.scope", not under its table. + // Every scope of a table counts under the table; scope is not broken out + // yet (it is inert until #235). Tables map[string]uint64 // Total is every parked message of the tenant, whatever the filter. Total uint64 @@ -263,8 +263,8 @@ var ErrNoDeadLetterQueue = errors.New("dead-letter queue not found") type DeadLetterStats interface { // DeadLetterCounts counts tenant id's parked messages per table — a // tenant served, rejected, or removed alike, for as long as its queue is - // kept. A non-empty table narrows Tables to that one (its unscoped - // messages). + // kept. A non-empty table narrows Tables to that one (all of its + // scopes). DeadLetterCounts(ctx context.Context, id tenant.ID, table string) (DeadLetterCounts, error) } diff --git a/internal/mq/subject.go b/internal/mq/subject.go index 2919464a1..2ee41ea31 100644 --- a/internal/mq/subject.go +++ b/internal/mq/subject.go @@ -84,24 +84,27 @@ func keyTenant(key string) (tenant.ID, bool) { return id, err == nil } -// parseTopicKey recovers the Topic from a subject tail. Three tokens are -// tenant, table and scope; two are tenant and table. A tail this package -// could not have written — one token, more than three, a token that does not -// decode, a tenant outside the grammar — cannot be split reliably, so the -// whole of it becomes the table of no tenant rather than being dropped. +// parseTopicKey recovers the Topic from a subject tail: the tenant token, +// read verbatim as keyTenant reads it, then one or two escaped tokens, table +// and scope (keyenc.Split). A tail this package could not have written — one +// token, more than three, a token that does not decode, a tenant outside the +// grammar — cannot be split reliably, so the whole of it becomes the table of +// no tenant rather than being dropped. func parseTopicKey(tail string) Topic { - parts := strings.Split(tail, ".") - switch len(parts) { - case 2, 3: - id, idErr := tenant.Parse(parts[0]) - table, tableErr := keyenc.Unescape(parts[1]) - scope, scopeErr := "", error(nil) - if len(parts) == 3 { - scope, scopeErr = keyenc.Unescape(parts[2]) - } - if idErr == nil && tableErr == nil && scopeErr == nil { - return Topic{Tenant: id, Table: table, Scope: scope} - } + first, rest, ok := strings.Cut(tail, ".") + if !ok { + return Topic{Table: tail} + } + id, idErr := tenant.Parse(first) + fields, fieldsErr := keyenc.Split(rest, '.') + if idErr != nil || fieldsErr != nil { + return Topic{Table: tail} + } + switch len(fields) { + case 1: + return Topic{Tenant: id, Table: fields[0]} + case 2: + return Topic{Tenant: id, Table: fields[0], Scope: fields[1]} } return Topic{Table: tail} } diff --git a/internal/mq/subject_test.go b/internal/mq/subject_test.go index c2a314b41..22d3bf9d8 100644 --- a/internal/mq/subject_test.go +++ b/internal/mq/subject_test.go @@ -10,8 +10,7 @@ import ( ) // The subjects are pinned byte for byte: an embedded broker holds messages -// under them across an upgrade, and the table token has been this encoding -// since v0.1.0. +// under them across an upgrade. func TestSubject_Golden(t *testing.T) { t.Parallel() for _, tt := range []struct { @@ -22,7 +21,7 @@ func TestSubject_Golden(t *testing.T) { {Topic{Tenant: "acme-co", Table: "default.clicks", Scope: "org_1"}, "acme-co.default%2Eclicks.org_1"}, {Topic{Tenant: "a", Table: "a.*.>", Scope: "*"}, "a.a%2E%2A%2E%3E.%2A"}, {Topic{Tenant: "a", Table: "my table", Scope: "tab\there"}, "a.my%20table.tab%09here"}, - {Topic{Tenant: "a", Table: "table-with-dashes", Scope: "org-1"}, "a.table%2Dwith%2Ddashes.org%2D1"}, + {Topic{Tenant: "a", Table: "table-with-dashes", Scope: "org-1"}, "a.table-with-dashes.org-1"}, {Topic{Tenant: "a", Table: "100%", Scope: "a/b"}, "a.100%25.a%2Fb"}, {Topic{Tenant: "a", Table: "{acme}:x"}, "a.%7Bacme%7D%3Ax"}, {Topic{Tenant: "a", Table: "nul\x00", Scope: "\xff"}, "a.nul%00.%FF"}, @@ -40,10 +39,12 @@ func TestSubject_Golden(t *testing.T) { } // A token another writer left partly unescaped, or escaped in lowercase, -// still reads as it always did. +// still reads as it always did — and so does v0.1.0's %2D for '-', so a +// message queued before '-' was kept reads as the same topic. func TestParseTopicKey_LenientTokens(t *testing.T) { t.Parallel() assert.Equal(t, Topic{Tenant: "a", Table: "b-c", Scope: "d.e"}, parseTopicKey("a.b-c.d%2ee")) + assert.Equal(t, parseTopicKey("a.table-with-dashes.org-1"), parseTopicKey("a.table%2Dwith%2Ddashes.org%2D1")) } func TestSubject_RoundTripsEveryTopic(t *testing.T) { @@ -121,6 +122,7 @@ func TestParseTopicKey_ForeignTailKeepsItself(t *testing.T) { "a.b.c.d", // more tokens than any topic renders "0.bad%2Gtoken", // a token that does not decode "a%2Eb.events", // a tenant outside the grammar + "a%2Db.events", // a tenant token is read verbatim, never decoded ".events", // a topic whose tenant was never set "events", // one token: no tenant leads it "bad%2G", // one token that does not decode diff --git a/internal/query/ident_test.go b/internal/query/ident_test.go index 572b005e7..ac4fd64a8 100644 --- a/internal/query/ident_test.go +++ b/internal/query/ident_test.go @@ -16,7 +16,7 @@ func TestEncodeTable(t *testing.T) { {"safe string", "my_table123", "my_table123"}, {"with dots", "default.clicks", "default%2Eclicks"}, {"with spaces", "my table", "my%20table"}, - {"with dashes and slashes", "a-b/c", "a%2Db%2Fc"}, + {"with dashes and slashes", "a-b/c", "a-b%2Fc"}, {"empty string", "", ""}, {"only safe characters", "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789_", "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789_"}, } From 999e92c5186831fced3ff000e00c4b0ab34b3c60 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 19:37:09 -0400 Subject: [PATCH 21/79] docs(keyenc): pin the %2D compatibility to builds since #612 A v0.1.0 queue is deleted at boot, so the lenient %2D decoding and the stream-resume caveat only concern queues an unreleased build since #612 wrote; say so in the changelog and the comments. List keyenc in the development.md tree and AGENTS.md's file structure, and correct the stale "L2" cache line beside it. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01B1tJWUp6oaDoH1usLwGtLF --- AGENTS.md | 3 ++- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/development.md | 3 ++- internal/keyenc/keyenc.go | 6 +++--- internal/keyenc/keyenc_test.go | 2 +- internal/mq/deadletter_test.go | 2 +- internal/mq/subject_test.go | 4 ++-- 8 files changed, 13 insertions(+), 11 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index bf8dc6a7f..592361133 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -38,7 +38,7 @@ Nineteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`dedupe/`** — `Deduplicator` interface → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) -- **`keyenc/`** — the one escaping composite keys are built from: `Escape` keeps `[A-Za-z0-9_-]` (exactly the tenant-id grammar, so a tenant id is its own escaped form) and writes every other byte as `%XX`, `Unescape` is `url.PathUnescape` (lenient: either hex case, and a byte left unescaped reads as itself, so v0.1.0's `%2D` still reads), `Join`/`AppendJoin` escape each field and put a separator between them (they panic on no fields, and on a separator the escaping could write or one outside ASCII) and `Split` reverses them. NATS subjects (`Join`/`Split` after the verbatim tenant) and the cache's namespace tokens use it; changing what it keeps orphans every stored key +- **`keyenc/`** — the one escaping composite keys are built from: `Escape` keeps `[A-Za-z0-9_-]` (exactly the tenant-id grammar, so a tenant id is its own escaped form) and writes every other byte as `%XX`, `Unescape` is `url.PathUnescape` (lenient: either hex case, and a byte left unescaped reads as itself, so a `%2D` an earlier build wrote still reads), `Join`/`AppendJoin` escape each field and put a separator between them (they panic on no fields, and on a separator the escaping could write or one outside ASCII) and `Split` reverses them. NATS subjects (`Join`/`Split` after the verbatim tenant) and the cache's namespace tokens use it; changing what it keeps orphans every stored key - **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`, `deadletter.go`), whose subject tokens are escaped by the shared `internal/keyenc`; `internal/app` constructs it and hands everything else a `mq.Broker` - **`observability/`** — OpenTelemetry pipeline: `InitProvider` wires trace/metric/log providers via OTLP gRPC (each signal independently gated). A top-level `Prometheus` config block drives an optional `/metrics` scrape endpoint that runs independently of OTLP push — standalone (Alloy/Mimir scrape, no collector), alongside OTLP, or off. `NewLogger` produces a slog handler that fans out to stdout AND OTLP (stdout always 100%, OTLP sample-rate-aware). `TraceHandler` injects trace_id/span_id from active spans. `tracer.go` provides W3C trace context propagation over message headers (`InjectHeaders`/`ExtractHeaders` on a plain header map; `internal/mq` injects on every publish and extracts on the `Subscribe` path, so this package never sees a NATS type). - **`pipes/`** — Named query pipes: `NamedQuery` type + `BindParams` + `Source` (read per request; `settings.Store` in production, `Static(q...)` in tests) @@ -435,6 +435,7 @@ internal/config/ → Configuration structs + loader internal/dedupe/ → Optional deduplication (interface + embedded/distributed) internal/discovery/ → ClickHouse schema introspection + ingest validation internal/ingest/ → Batch buffer with DLQ + Active Sweeper (NATS message lifecycle) +internal/keyenc/ → One escaping for composite keys (NATS subject tokens, cache namespace tokens) internal/mq/ → MQ boundary (the only NATS/JetStream importer: owned message/consumer/stream types + embedded server) internal/observability/ → OpenTelemetry pipeline (traces/metrics/logs providers, Prometheus exporter, slog fan-out, message-header trace propagation) internal/pipes/ → Named query pipes (types, parameter binding, Source) diff --git a/CHANGELOG.md b/CHANGELOG.md index 6283af538..7861daaa1 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -32,7 +32,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Changed -- **One escaping for composite keys, `-` kept; dead-letter counts per table** (`internal/keyenc` (new, + tests), `internal/mq/{mq,subject,embedded,deadletter}.go` (+ tests), `internal/query/ident.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/architecture.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). NATS subject tokens and the cache's namespace tokens each carried a copy of the same encoder; both now call `internal/keyenc`, which the dedupe keys will use too, and subjects are built with its `Join` (each field escaped, with a separator the escaping never writes between them). The escaping now keeps `-` as well as ASCII letters, digits and `_` — exactly the tenant-id grammar — so a table or scope such as `my-table` is `my-table` in a subject rather than `my%2Dtable`; every other byte is escaped as before (pinned by golden tests, and against v0.1.0's encoder for every other byte value). Decoding is `url.PathUnescape`, as it was, so a subject written with v0.1.0's `%2D` reads as the same topic and a dead-letter count merges both forms. One thing notices the change: a `/v1/stream` client resuming across the upgrade (`Last-Event-ID` or `since`) on a table whose name holds `-` misses that table's events queued before the upgrade, since the replay filters on the table's exact subject. `GET /v1/ops/dlq/stats` now counts every scope of a table under the table itself, and `?table=` keeps all of its scopes; a scoped message used to count under `table.scope`, a name a dotted table could share. Scope is always empty today, so the response is unchanged. +- **One escaping for composite keys, `-` kept; dead-letter counts per table** (`internal/keyenc` (new, + tests), `internal/mq/{mq,subject,embedded,deadletter}.go` (+ tests), `internal/query/ident.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{architecture,development}.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). NATS subject tokens and the cache's namespace tokens each carried a copy of the same encoder; both now call `internal/keyenc`, which the dedupe keys will use too, and subjects are built with its `Join` (each field escaped, with a separator the escaping never writes between them). The escaping now keeps `-` as well as ASCII letters, digits and `_` — exactly the tenant-id grammar — so a table or scope such as `my-table` is `my-table` in a subject rather than `my%2Dtable`; every other byte is escaped as before (pinned by golden tests, and against v0.1.0's encoder for every other byte value). Upgrading from v0.1.0 notices nothing further, since its queue is deleted at boot (below). A queue an unreleased build since [#612](https://github.com/Wave-RF/WaveHouse/pull/612) wrote still reads, because decoding is `url.PathUnescape` as it was: `%2D` decodes to `-`, and a dead-letter count merges both forms. Only a `/v1/stream` client resuming across such an upgrade (`Last-Event-ID` or `since`) on a table whose name holds `-` misses that table's events queued before it, since the replay filters on the table's exact subject. `GET /v1/ops/dlq/stats` now counts every scope of a table under the table itself, and `?table=` keeps all of its scopes; a scoped message used to count under `table.scope`, a name a dotted table could share. Scope is always empty today, so the response is unchanged. - **Each tenant has a message queue of its own** (`internal/mq/{mq,subject,embedded}.go` (+ tests), `internal/ingest/{sweeper,worker}.go` (+ tests), `internal/api/{dlq,ingest}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `internal/stream/{subscriber,hub}.go`, `internal/settings/{settings,store,registry}.go` (+ tests), `internal/testutil/{mocks,testutil}.go`, `clients/ts/src/{dlq,types}.ts` (+ tests), `docs/src/content/docs/{deployment,api,architecture,ingest-pipeline,durability,why-wavehouse}.md`, `docs/src/content/docs/sdk/{admin,reference,streaming}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `AGENTS.md`): story 5b of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). The embedded NATS server keeps each tenant's events on a pair of JetStream streams of its own — `INGEST_` (`ingest..>`, `DiscardNew`) at the tenant's own `mq.max_bytes_gb`, and `DLQ_` (`dlq..>`, `DiscardOld`) at a tenth of it — opened when the tenant is first served and kept, at the budget it last had, when its folder is rejected or removed; subjects are unchanged, and nothing outside `internal/mq` names a stream. A tenant at its budget gets `503` while the others keep publishing, and the ingest worker's and the hub bridge's durables are held on every tenant's stream, each with its own ack floor and `MaxAckPending`, so one tenant's backlog holds back neither another's delivery nor its purge; the worker's prefetch and the hub bridge's fetch-ahead are each shared across the tenants' streams. A removed or rejected tenant's stream is still consumed, so its queued rows reach the worker and are parked on its own dead-letter queue. The sweeper purges each tenant's stream at that tenant's own `stream.gap_window_minutes` — a rejected tenant's as its folder last had it (`settings.Registry.Known` now yields each tenant's last adopted store), and all of its history if the folder has been rejected since boot, so its clients resume once the folder is fixed — keeping no acknowledged history for a removed tenant, and every served tenant's `mq.max_bytes_gb` is applied after each reload rather than tenant `0`'s alone (the boot warning about a nested directory with no tenant `0` is gone with it, and so is the tracking of tenant `0`'s last adopted store: a flat directory's ops gate, its one reader left, reads tenant `0` through the registry). One tenant's failed purge holds up no other tenant's, and the sweep logs it at `ERROR` unless every failure in it is a buffer consumer not created yet. A reload that shrinks a budget no longer deletes dead letters: a dead-letter stream holding more than a tenth of the new budget keeps what it holds, and that is logged — the interim guard of [#532](https://github.com/Wave-RF/WaveHouse/issues/532). JetStream's own check of the streams' caps against the disk, which made them fit 75% of the free disk at boot together, is lifted: a budget is a cap and never a reservation, what the budgets add up to against the disk is [#138](https://github.com/Wave-RF/WaveHouse/issues/138)'s, and a flat directory whose budget exceeds three quarters of the free disk now boots where it used to be refused. A queue that cannot be opened refuses a flat boot like any other store and, over a nested directory, costs its tenant alone — its ingest answers `503`, each reload trying again, and a publish at most once every five seconds, so its clients retrying never hold up another tenant's queue. `GET /v1/ops/dlq/stats` reads one tenant's dead-letter queue — the one `?tenant=` names, parsed strictly like the other admin reads, and tenant `0`'s without it, no longer the sum across tenants — answering for a rejected or removed tenant too and `404` for a tenant with no queue; the SDK's `wh.dlq.list()` and `.table()` take a `tenant` option. The streams an earlier build kept for every tenant together (`WAVEHOUSE`, `WAVEHOUSE_DLQ`) overlap every tenant's subjects and are deleted at boot, what they held with them, and a subject with no tenant token no longer reads as tenant `0`'s. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 59b2fe3fe..7f0aa379b 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -204,7 +204,7 @@ The hot-reloadable half of configuration: a directory of four JSON files (`confi ### `keyenc/` — Key Escaping -- **keyenc.go** — The one escaping composite keys are built from, so a name can never be mistaken for a separator: `Escape` keeps ASCII letters, digits, `_` and `-` — exactly the tenant-id grammar, so a tenant id is its own escaped form — and writes every other byte as `%XX` (uppercase hex); `Unescape` is `url.PathUnescape`, which decodes `%XX` in either case and takes any other byte as itself, so v0.1.0's `%2D` for `-` still reads. `Join`/`AppendJoin` escape each field and put a separator between them, panicking on no fields and on a separator the escaping could write or one outside ASCII, and `Split` reverses them. NATS subject tokens (`internal/mq`) and the cache's namespace tokens (`query.SafeEncodeToken`) both use it. Keys built from it are stored, so changing what it keeps orphans them. +- **keyenc.go** — The one escaping composite keys are built from, so a name can never be mistaken for a separator: `Escape` keeps ASCII letters, digits, `_` and `-` — exactly the tenant-id grammar, so a tenant id is its own escaped form — and writes every other byte as `%XX` (uppercase hex); `Unescape` is `url.PathUnescape`, which decodes `%XX` in either case and takes any other byte as itself, so a `%2D` for `-` that an earlier build wrote still reads. `Join`/`AppendJoin` escape each field and put a separator between them, panicking on no fields and on a separator the escaping could write or one outside ASCII, and `Split` reverses them. NATS subject tokens (`internal/mq`) and the cache's namespace tokens (`query.SafeEncodeToken`) both use it. Keys built from it are stored, so changing what it keeps orphans them. ## Data Flows diff --git a/docs/src/content/docs/development.md b/docs/src/content/docs/development.md index 01f82e735..16b65a74a 100644 --- a/docs/src/content/docs/development.md +++ b/docs/src/content/docs/development.md @@ -454,13 +454,14 @@ WaveHouse/ │ ├── api/ # HTTP handlers, router, middleware │ ├── app/ # Process wiring (build every component, run under one errgroup, release in reverse) │ ├── auth/ # JWT/JWKS authentication middleware -│ ├── cache/ # L1 (Ristretto) + L2 caching +│ ├── cache/ # Query cache: Ristretto L1 + the tenant-led version index │ ├── chconn/ # ClickHouse pools, one per connection tuple (reconciled on settings reload) │ ├── chsql/ # Shared ClickHouse SQL helpers (quoting + bind-safety) │ ├── config/ # YAML + env var configuration │ ├── dedupe/ # Optional deduplication (Pebble) │ ├── discovery/ # ClickHouse schema introspection + validation │ ├── ingest/ # Batch buffering + DLQ + Active Sweeper +│ ├── keyenc/ # One escaping for composite keys (NATS subject tokens, cache namespace tokens) │ ├── mq/ # MQ boundary: the only NATS/JetStream importer │ ├── observability/ # OpenTelemetry pipeline (traces/metrics/logs + Prometheus) │ ├── pipes/ # Named query pipes (types + parameter binding) diff --git a/internal/keyenc/keyenc.go b/internal/keyenc/keyenc.go index 59bac7e1c..0e3434bc6 100644 --- a/internal/keyenc/keyenc.go +++ b/internal/keyenc/keyenc.go @@ -7,8 +7,8 @@ // tenant id is its own escaped form. // // Keys built from it are stored — queued under NATS subjects, held in caches -// — so a change to what it keeps orphans them. v0.1.0 escaped '-' as %2D; -// Unescape still reads that form. +// — so a change to what it keeps orphans them. Earlier builds escaped '-' as +// %2D; Unescape still reads that form. package keyenc import ( @@ -49,7 +49,7 @@ func AppendEscape(dst []byte, s string) []byte { // Unescape reverses Escape. It is url.PathUnescape: %XX in either hex case // decodes, and any other byte reads as itself, so a field another writer -// left partly unescaped — v0.1.0's %2D included — still reads. +// left partly unescaped — an earlier build's %2D included — still reads. func Unescape(s string) (string, error) { return url.PathUnescape(s) } diff --git a/internal/keyenc/keyenc_test.go b/internal/keyenc/keyenc_test.go index 5b1630d6f..877599e2c 100644 --- a/internal/keyenc/keyenc_test.go +++ b/internal/keyenc/keyenc_test.go @@ -106,7 +106,7 @@ func TestUnescape(t *testing.T) { "plain": "plain", "default%2Eclicks": "default.clicks", "lower%2ecase": "lower.case", - "evt%2D123": "evt-123", // v0.1.0's form + "evt%2D123": "evt-123", // an earlier build's form "%00%FF": "\x00\xff", "a+b": "a+b", } { diff --git a/internal/mq/deadletter_test.go b/internal/mq/deadletter_test.go index aaed64fef..aed760943 100644 --- a/internal/mq/deadletter_test.go +++ b/internal/mq/deadletter_test.go @@ -13,7 +13,7 @@ func TestDeadLetterTables(t *testing.T) { "dlq.0.a.b": 2, // table "a", scope "b" "dlq.0.a": 4, // table "a", unscoped "dlq.0.my-t": 8, - "dlq.0.my%2Dt.org-1": 16, // v0.1.0's escaping of '-', scoped + "dlq.0.my%2Dt.org-1": 16, // an earlier build's escaping of '-', scoped "dlq.0.clicks.org%2E": 32, } assert.Equal(t, map[string]uint64{"a.b": 1, "a": 6, "my-t": 24, "clicks": 32}, diff --git a/internal/mq/subject_test.go b/internal/mq/subject_test.go index 22d3bf9d8..28f4940be 100644 --- a/internal/mq/subject_test.go +++ b/internal/mq/subject_test.go @@ -39,8 +39,8 @@ func TestSubject_Golden(t *testing.T) { } // A token another writer left partly unescaped, or escaped in lowercase, -// still reads as it always did — and so does v0.1.0's %2D for '-', so a -// message queued before '-' was kept reads as the same topic. +// still reads as it always did — and so does an earlier build's %2D for '-', +// so a message it queued reads as the same topic. func TestParseTopicKey_LenientTokens(t *testing.T) { t.Parallel() assert.Equal(t, Topic{Tenant: "a", Table: "b-c", Scope: "d.e"}, parseTopicKey("a.b-c.d%2ee")) From cc516d8c33d6138759060898a7c3f04a7f746569 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 19:42:43 -0400 Subject: [PATCH 22/79] refactor(cache): namespaces carry raw names; the cache escapes its keys A cache.Namespace now carries the raw table and scope, and the version index builds its keys with keyenc.Join/AppendJoin: the table key and the namespace key are one '.'-joined level, and the query key escapes the sha and puts '|' between the escaped levels by hand. The structured-query read and the ingest worker's invalidation pass the names as they have them, so neither can forget to escape and no name reaches a key unescaped. query.SafeEncodeToken, whose only job was that pre-escaping, is removed. The conformance suite gains two cases every backend runs: a name holding a dot, a space or a '%' is read and bumped under one key, and dependency sets that would run together under an unescaped join (at a '.', ':', '|' or NUL) stay two entries. The second fails against the old unescaped index. An ingest test drives an envelope through parseMsg and flushTable into a real LocalCache for "default.clicks" and "my table", and an api test serves those tables from the cache through the handler and orphans them with the namespace the worker sends. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01B1tJWUp6oaDoH1usLwGtLF --- AGENTS.md | 4 +- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 8 ++-- internal/api/cache_tenant_test.go | 40 ++++++++++++++++- internal/api/structured_query.go | 9 ++-- internal/cache/version_manager.go | 30 ++++++++----- internal/cache/version_manager_test.go | 10 +++++ internal/ingest/worker.go | 14 +++--- internal/ingest/worker_test.go | 56 +++++++++++++++++++++--- internal/query/ident.go | 8 ---- internal/query/ident_test.go | 31 ------------- internal/testutil/cachetest/cachetest.go | 50 +++++++++++++++++++++ 12 files changed, 184 insertions(+), 78 deletions(-) delete mode 100644 internal/query/ident.go delete mode 100644 internal/query/ident_test.go diff --git a/AGENTS.md b/AGENTS.md index 9982ed3db..4eebcbdfb 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -31,14 +31,14 @@ Nineteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers - **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` handing the sweeper each tenant's own gap window (a rejected tenant's as its folder last had it, unbounded for one rejected since boot) and the `mq.max_bytes_gb` reconcile each served tenant's byte budget, and `defaultPolicy` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, and the same hook's `Hub.Prune` ends the open streams of a tenant no longer served), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it - **`auth/`** — JWT auth middleware: HMAC **or** JWKS verification with `alg` pinned to the active verifier, role extraction from a configurable claim path; always runs, never rejects (bad token → empty role + stashed reason). One verifier per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 9): `Authenticator` keys them by `tenant.ID` — the request store's `settings.Store.Tenant()`, through an injected `TenantSource`; `tenant.Default` on the tenant-exempt routes — built from each tenant's `auth` block by `Reconfigure`, dropped by `Prune` once the tenant stops being served (rejected or removed), released by `Close`; the secrets (`Config`) are boot-level and shared. A JWKS key set is fetched off the boot and reload paths: until one has been stored the verifier is pending and a token-bearing request gets `503` + `Retry-After` from `api.refuseUnverifiable` (`auth.ErrVerifierPending`), never a `default_role` evaluation; refresh is library-managed (Eric, 2026-09-22), response capped at 1 MiB; the operator key's admin role is the request tenant's -- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, `..
.
.` for a namespace — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version folded into every namespace and query key of one tenant, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read and `Set` files the fill under it, so a write landing mid-query orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) +- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, `..
.
.` for a namespace, each field escaped by `keyenc` where the key is built (a `Namespace` carries the raw table and scope, so no caller escapes) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version folded into every namespace and query key of one tenant, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read and `Set` files the fill under it, so a write landing mid-query orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) - **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability — boot is the validator, there is no dry run - **`dedupe/`** — `Deduplicator` interface → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) -- **`keyenc/`** — the one escaping composite keys are built from: `Escape` keeps `[A-Za-z0-9_-]` (exactly the tenant-id grammar, so a tenant id is its own escaped form) and writes every other byte as `%XX`, `Unescape` is `url.PathUnescape` (lenient: either hex case, and a byte left unescaped reads as itself, so v0.1.0's `%2D` still reads), `Join`/`AppendJoin` escape each field and put a separator between them (they panic on no fields, and on a separator the escaping could write or one outside ASCII) and `Split` reverses them. NATS subjects (`Join`/`Split` after the verbatim tenant) and the cache's namespace tokens use it; changing what it keeps orphans every stored key +- **`keyenc/`** — the one escaping composite keys are built from: `Escape` keeps `[A-Za-z0-9_-]` (exactly the tenant-id grammar, so a tenant id is its own escaped form) and writes every other byte as `%XX`, `Unescape` is `url.PathUnescape` (lenient: either hex case, and a byte left unescaped reads as itself, so v0.1.0's `%2D` still reads), `Join`/`AppendJoin` escape each field and put a separator between them (they panic on no fields, and on a separator the escaping could write or one outside ASCII) and `Split` reverses them. The package that builds a key takes raw names and escapes them itself, so no caller has to and no field reaches a key unescaped: NATS subjects (`Join`/`Split` after the verbatim tenant) and the cache's version-index keys (`cache.VersionManager`) use it; changing what it keeps orphans every stored key - **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`, `deadletter.go`), whose subject tokens are escaped by the shared `internal/keyenc`; `internal/app` constructs it and hands everything else a `mq.Broker` - **`observability/`** — OpenTelemetry pipeline: `InitProvider` wires trace/metric/log providers via OTLP gRPC (each signal independently gated). A top-level `Prometheus` config block drives an optional `/metrics` scrape endpoint that runs independently of OTLP push — standalone (Alloy/Mimir scrape, no collector), alongside OTLP, or off. `NewLogger` produces a slog handler that fans out to stdout AND OTLP (stdout always 100%, OTLP sample-rate-aware). `TraceHandler` injects trace_id/span_id from active spans. `tracer.go` provides W3C trace context propagation over message headers (`InjectHeaders`/`ExtractHeaders` on a plain header map; `internal/mq` injects on every publish and extracts on the `Subscribe` path, so this package never sees a NATS type). - **`pipes/`** — Named query pipes: `NamedQuery` type + `BindParams` + `Source` (read per request; `settings.Store` in production, `Static(q...)` in tests) diff --git a/CHANGELOG.md b/CHANGELOG.md index 27b7cdf46..976ec66a8 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -80,7 +80,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed -- **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/app/wire.go`, `docs/src/content/docs/{architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` and pipe execution rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL. The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. +- **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/query/ident.go` (removed, + tests), `internal/app/wire.go`, `docs/src/content/docs/{architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` and pipe execution rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL. The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. The cache now escapes table and scope names itself: a `Namespace` carries them raw and the version index builds its keys with `internal/keyenc`, so neither the structured-query read nor the ingest worker's invalidation escapes them (`query.SafeEncodeToken` is gone) and no name reaches a key unescaped. The suite pins that a name holding a dot, a space or a `%` is read and bumped under one key, and that names which would run together unescaped (`a.0.b` against `a` with scope `b.0.`) stay two entries. The keys themselves are unchanged apart from the query key's sha, now escaped too; they live in the process, so nothing stored is orphaned. - **An explicit `false`, `0` or `""` in `config.yaml` is no longer replaced by the key's default** (`internal/config/config.go`, `internal/config/defaults_test.go` (new), `docs/src/content/docs/configuration.mdx`, `AGENTS.md`): [#631](https://github.com/Wave-RF/WaveHouse/issues/631). Defaults lived in cleanenv `env-default` tags, which cleanenv applies after the YAML decode to any field still at its zero value, so it could not tell a key the file set to its zero value from one the file left out. `otel.traces.enabled: false`, `otel.metrics.enabled: false` and `otel.logs.enabled: false` came back `true`; `otel.traces.sample_rate: 0` and `otel.logs.sample_rate: 0` came back `1.0`; `server.shutdown_timeout: 0` came back `10`; `cache.l1_max_cost: 0`, `prometheus.path: ""` and `data_dir: ""` came back as their defaults; `server.port: 0` came back `8080`. All of it was silent. Defaults now live in one Go function, `defaults()`, which `Load` starts from before decoding the file and then applying `WH_*` variables, so the order is env > YAML > default and a key the file sets always wins. **Behaviour change if your file relied on the bug:** a zero you wrote now takes effect. A file that says `sample_rate: 0` now exports no traces (or no DEBUG/INFO logs), where it silently exported everything; a signal set `enabled: false` is now off; `shutdown_timeout: 0` now skips the drain. `cache.l1_max_cost: 0`, `server.port: 0`, and `data_dir: ""` now refuse boot (`cache init: MaxCost can't be zero`, `server.port 0 out of range`, `data_dir (WH_DATA_DIR) is required`) instead of running on the default; an empty `prometheus.path` refuses boot when `prometheus.enabled` is true. Delete the key to get the default back. Env vars are unchanged: they already honoured an explicit zero. New tests load through `config.Load` for every affected key (a YAML zero is kept, an absent key gets the default, env wins in both directions), refuse an `env-default` tag on any field, and pin each documented default in `configuration.mdx` to `defaults()`. - **An embedded queue store that cannot be created fails boot at once, naming the cause** (`internal/mq/embedded.go` (+ tests)): part of [#617](https://github.com/Wave-RF/WaveHouse/issues/617). A regular file at `/nats`, or a `nats` directory that could not be created there, failed JetStream in the background, so boot waited out the server's 5s readiness check and reported only `nats server not ready`. `NewEmbedded` now creates the directory first (at `0700`, as the server does) and refuses boot with the mkdir error. An existing but unwritable `nats` directory still takes the old path. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 9149835e4..b409c9e29 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -61,7 +61,7 @@ internal/ ├── dedupe/ Optional deduplication (Pebble) ├── discovery/ ClickHouse schema introspection and validation ├── ingest/ Batch buffering, DLQ, and Active Sweeper -├── keyenc/ The one escaping composite keys are built from (NATS subject tokens, cache namespace tokens) +├── keyenc/ The one escaping composite keys are built from (NATS subject tokens, cache version-index keys) ├── mq/ MQ boundary: the only NATS/JetStream importer (owned message/consumer/stream types + embedded server) ├── observability/ OpenTelemetry pipeline (traces/metrics/logs + Prometheus exposition) ├── pipes/ Named query pipes (NamedQuery type, parameter binding, Source) @@ -110,9 +110,9 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ ### `cache/` — Query Cache -- **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on, each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope: one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)). `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. +- **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on, each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope by their raw names, which the cache escapes where it builds a key: one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)). `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. -- **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: a namespace key is `..
.
.`, and a query key is folded with the tenant's version and each dependency's namespace key and namespace version, so bumping a table (a scopeless write) or one scope — scope is reserved and empty today, so every write is the whole-table bump — orphans every dependent entry without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results, and the flat directory's single tenant simply carries the `0` prefix. `BumpTenant` (behind `InvalidateTenant`) advances the tenant version that leads every namespace key of one tenant, orphaning its every namespace and every cached query in one step (the tenant version is folded into every query key, so a pipe result with no dependencies is orphaned too) — a table no bump ever keyed included, which an enumeration of the index would miss — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. +- **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: a namespace key is `..
.
.`, its fields joined with `keyenc.Join` so a dot in a table or scope name is escaped rather than read as a separator, and a query key, `|.||…` with the sha escaped, folds in the tenant's version and each dependency's namespace key and namespace version, so bumping a table (a scopeless write) or one scope — scope is reserved and empty today, so every write is the whole-table bump — orphans every dependent entry without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results, and the flat directory's single tenant simply carries the `0` prefix. `BumpTenant` (behind `InvalidateTenant`) advances the tenant version that leads every namespace key of one tenant, orphaning its every namespace and every cached query in one step (the tenant version is folded into every query key, so a pipe result with no dependencies is orphaned too) — a table no bump ever keyed included, which an enumeration of the index would miss — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. ### `config/` — Configuration @@ -204,7 +204,7 @@ The hot-reloadable half of configuration: a directory of four JSON files (`confi ### `keyenc/` — Key Escaping -- **keyenc.go** — The one escaping composite keys are built from, so a name can never be mistaken for a separator: `Escape` keeps ASCII letters, digits, `_` and `-` — exactly the tenant-id grammar, so a tenant id is its own escaped form — and writes every other byte as `%XX` (uppercase hex); `Unescape` is `url.PathUnescape`, which decodes `%XX` in either case and takes any other byte as itself, so v0.1.0's `%2D` for `-` still reads. `Join`/`AppendJoin` escape each field and put a separator between them, panicking on no fields and on a separator the escaping could write or one outside ASCII, and `Split` reverses them. NATS subject tokens (`internal/mq`) and the cache's namespace tokens (`query.SafeEncodeToken`) both use it. Keys built from it are stored, so changing what it keeps orphans them. +- **keyenc.go** — The one escaping composite keys are built from, so a name can never be mistaken for a separator: `Escape` keeps ASCII letters, digits, `_` and `-` — exactly the tenant-id grammar, so a tenant id is its own escaped form — and writes every other byte as `%XX` (uppercase hex); `Unescape` is `url.PathUnescape`, which decodes `%XX` in either case and takes any other byte as itself, so v0.1.0's `%2D` for `-` still reads. `Join`/`AppendJoin` escape each field and put a separator between them, panicking on no fields and on a separator the escaping could write or one outside ASCII, and `Split` reverses them. The package that builds a key takes raw names and escapes them itself, so no caller has to: NATS subject tokens (`internal/mq`) and the cache's version-index keys (`internal/cache`) both use it. Keys built from it are stored, so changing what it keeps orphans them. ## Data Flows diff --git a/internal/api/cache_tenant_test.go b/internal/api/cache_tenant_test.go index 2ab2fa9ce..742b55f25 100644 --- a/internal/api/cache_tenant_test.go +++ b/internal/api/cache_tenant_test.go @@ -16,6 +16,7 @@ import ( "github.com/stretchr/testify/require" "github.com/Wave-RF/WaveHouse/internal/cache" + "github.com/Wave-RF/WaveHouse/internal/discovery" "github.com/Wave-RF/WaveHouse/internal/pipes" "github.com/Wave-RF/WaveHouse/internal/policy" "github.com/Wave-RF/WaveHouse/internal/query" @@ -228,7 +229,7 @@ func (c *bumpingConn) Query(context.Context, string, ...any) (driver.Rows, error func TestCachedRoutes_BumpDuringQueryOrphansTheFill(t *testing.T) { bumps := map[string]func(ctx context.Context, c cache.Cache) error{ "structured query": func(ctx context.Context, c cache.Cache) error { - _, err := c.Invalidate(ctx, []cache.Namespace{{Tenant: tenant.Default, Table: query.SafeEncodeToken("clicks")}}) + _, err := c.Invalidate(ctx, []cache.Namespace{{Tenant: tenant.Default, Table: "clicks"}}) return err }, "pipe execute": func(ctx context.Context, c cache.Cache) error { return c.InvalidateTenant(ctx, tenant.Default) }, @@ -253,3 +254,40 @@ func TestCachedRoutes_BumpDuringQueryOrphansTheFill(t *testing.T) { }) } } + +// A structured query files its result under the table as the request names +// it, raw — the namespace the ingest worker bumps after an insert into that +// table (ingest's TestFlushTable_BumpsWhatTheReadFiles) — and the cache +// escapes both, so a name holding a dot or a space is served from the cache +// and orphaned by an insert like any other. +func TestStructuredQuery_RawTableNameMeetsTheInsertsBump(t *testing.T) { + t.Parallel() + tables := []string{"default.clicks", "my table"} + schemas := make([]*discovery.TableSchema, 0, len(tables)) + grants := make(map[string]policy.TablePolicy, len(tables)) + for _, name := range tables { + schemas = append(schemas, &discovery.TableSchema{Name: name, Columns: []discovery.Column{{Name: "page", Type: "String"}}}) + grants[name] = policy.TablePolicy{"viewer": {Select: &policy.SelectPermissions{AllowColumns: []string{"page"}}}} + } + l1, err := cache.NewLocal(1 << 20) + require.NoError(t, err) + t.Cleanup(func() { _ = l1.Close() }) + h := NewStructuredQueryHandler(fixedConn(&countingConn{}), l1, fixedRegistry(testutil.NewTestSchemaRegistry(t, schemas)), + staticPolicy(&policy.Policy{DefaultRole: "viewer", Tables: grants}), func(*settings.Store) int { return 60 }, + func(*settings.Store) time.Duration { return 5 * time.Second }, nil) + + for _, table := range tables { + xcache := func() string { + w := httptest.NewRecorder() + h.Handle(w, withTenant(structuredQueryRequest(t, table, query.StructuredQuery{SelectAll: true}))) + require.Equal(t, http.StatusOK, w.Code, "body: %s", w.Body.String()) + l1.Wait() + return w.Header().Get("X-Cache") + } + assert.Equal(t, "MISS", xcache(), table) + assert.Equal(t, "HIT", xcache(), table) + _, err := l1.Invalidate(t.Context(), []cache.Namespace{{Tenant: tenant.Default, Table: table}}) + require.NoError(t, err) + assert.Equal(t, "MISS", xcache(), "%s: the insert's bump orphans the cached result", table) + } +} diff --git a/internal/api/structured_query.go b/internal/api/structured_query.go index 921bc8f08..435c84779 100644 --- a/internal/api/structured_query.go +++ b/internal/api/structured_query.go @@ -174,13 +174,10 @@ func (h *StructuredQueryHandler) Handle(w http.ResponseWriter, r *http.Request) // TODO: impl scope scope := "" - safeTableName := query.SafeEncodeToken(table) // A structured query reads one table, so it depends on a single namespace: - // the request's tenant, the table, the scope. Encode the scope the way the - // ingest worker does (worker.go invalidate) so the read and invalidation - // sides build identical namespace keys once scope is implemented; - // SafeEncodeToken("") is "", so this is a no-op while scope is empty. - deps := []cache.Namespace{{Tenant: store.Tenant(), Table: safeTableName, Scope: query.SafeEncodeToken(scope)}} + // the request's tenant, the table, the scope — raw names, as the ingest + // worker's invalidation passes them; the cache escapes both sides alike. + deps := []cache.Namespace{{Tenant: store.Tenant(), Table: table, Scope: scope}} // Try cache. The snapshot is of the versions before the query runs, so a // write landing mid-query orphans the fill (#382). diff --git a/internal/cache/version_manager.go b/internal/cache/version_manager.go index d97bc70aa..48cf17c84 100644 --- a/internal/cache/version_manager.go +++ b/internal/cache/version_manager.go @@ -3,9 +3,11 @@ package cache import ( "fmt" "sort" + "strconv" "strings" "sync" + "github.com/Wave-RF/WaveHouse/internal/keyenc" "github.com/Wave-RF/WaveHouse/internal/tenant" ) @@ -34,7 +36,10 @@ func NewVersionManager() *VersionManager { // Namespace is one (tenant, table, scope) a cached result depends on. The // tenant leads every key built from it, so the same table under two tenants -// is two namespaces, versioned and bumped apart (#583 story 8). +// is two namespaces, versioned and bumped apart (#583 story 8). Table and +// Scope are raw names: the cache escapes them where it builds a key +// (keyenc), so no caller escapes and no separator in a name can run two +// fields together. type Namespace struct { Tenant tenant.ID Table string @@ -42,23 +47,24 @@ type Namespace struct { } // tableKeyLocked renders the table-versions key, -// "..
"; caller must hold vm.mu. A tenant id -// cannot contain a dot and callers encode the table dot-free, so the tokens -// can never run together. +// "..
", its fields joined by keyenc so no +// dot in a table name can run into the next field; caller must hold vm.mu. func (vm *VersionManager) tableKeyLocked(id tenant.ID, table string) string { - return fmt.Sprintf("%s.%d.%s", id, vm.tenantVersions[id], table) + return keyenc.Join('.', string(id), strconv.FormatUint(vm.tenantVersions[id], 10), table) } -// namespaceKeyLocked builds the namespace-table key; caller must hold vm.mu. +// namespaceKeyLocked builds the namespace-table key: the table key, then +// the table version and the scope, one more level of the same join; caller +// must hold vm.mu. func (vm *VersionManager) namespaceKeyLocked(ns Namespace) string { tk := vm.tableKeyLocked(ns.Tenant, ns.Table) - return fmt.Sprintf("%s.%d.%s", tk, vm.tableVersions[tk], ns.Scope) + return string(keyenc.AppendJoin([]byte(tk+"."), '.', strconv.FormatUint(vm.tableVersions[tk], 10), ns.Scope)) } // NamespaceKey renders the namespace-table key for ns at its tenant's and // table's current versions: -// "..
.." (scopeless -// scope is "", so e.g. ".0.
.."). +// "..
..", each field +// escaped (scopeless scope is "", so e.g. ".0.
.."). func (vm *VersionManager) NamespaceKey(ns Namespace) string { vm.mu.RLock() defer vm.mu.RUnlock() @@ -71,7 +77,9 @@ func (vm *VersionManager) NamespaceKey(ns Namespace) string { // a bump of the tenant or of any dependency misses the key — a result with no // deps (a pipe) is orphaned by BumpTenant too. A structured query passes one // Namespace; a pipe passes several. Deps are sorted so their order never -// changes the key. +// changes the key. The key nests two levels: the escaped sha and the +// '.'-joined tenant and dependency segments, separated by '|', which no +// escaped field or '.' join ever holds. func (vm *VersionManager) QueryKey(id tenant.ID, sha string, deps []Namespace) string { segs := make([]string, len(deps)) // Lock per dependency rather than across the whole loop: each dep's table + @@ -90,7 +98,7 @@ func (vm *VersionManager) QueryKey(id tenant.ID, sha string, deps []Namespace) s tv := vm.tenantVersions[id] vm.mu.RUnlock() sort.Strings(segs) - return fmt.Sprintf("%s|%s.%d|%s", sha, id, tv, strings.Join(segs, "|")) + return keyenc.Escape(sha) + "|" + keyenc.Join('.', string(id), strconv.FormatUint(tv, 10)) + "|" + strings.Join(segs, "|") } // BumpTable advances a tenant's table version, orphaning every namespace — and diff --git a/internal/cache/version_manager_test.go b/internal/cache/version_manager_test.go index e1a1c330b..91b59db52 100644 --- a/internal/cache/version_manager_test.go +++ b/internal/cache/version_manager_test.go @@ -19,6 +19,11 @@ func TestVersionManager_NamespaceKey(t *testing.T) { assert.Equal(t, "acme.0.users.0.org_1", vm.NamespaceKey(Namespace{Tenant: "acme", Table: "users", Scope: "org_1"})) assert.Equal(t, "0.0.users.0.", vm.NamespaceKey(Namespace{Tenant: tenant.Default, Table: "users"})) + // Names arrive raw and are escaped into the key, so a dot or a space in + // one is never read as the separator. + assert.Equal(t, "acme.0.default%2Eclicks.0.org%2E1", vm.NamespaceKey(Namespace{Tenant: "acme", Table: "default.clicks", Scope: "org.1"})) + assert.Equal(t, "acme.0.my%20table.0.", vm.NamespaceKey(Namespace{Tenant: "acme", Table: "my table"})) + // The table version is embedded in every namespace key for that tenant's // table, so a BumpTable is reflected across all its scopes at once — and // nowhere else: the same table under another tenant keeps its version. @@ -48,6 +53,11 @@ func TestVersionManager_QueryKey(t *testing.T) { // No deps (a pipe) still folds the tenant version. assert.Equal(t, "hash123|acme.0|", vm.QueryKey("acme", "hash123", nil)) + // The sha is a field like any other: escaped, so no '|' in it can pass + // for the separator. + assert.Equal(t, "acme%3Aquery%3Aab|acme.0|acme.0.my%20table.0..0", + vm.QueryKey("acme", "acme:query:ab", []Namespace{{Tenant: "acme", Table: "my table"}})) + // Dependency order must not change the key (segments are sorted). deps1 := []Namespace{{Tenant: "acme", Table: "a"}, {Tenant: "acme", Table: "b"}} deps2 := []Namespace{{Tenant: "acme", Table: "b"}, {Tenant: "acme", Table: "a"}} diff --git a/internal/ingest/worker.go b/internal/ingest/worker.go index f6cbaf6a7..1576e58c7 100644 --- a/internal/ingest/worker.go +++ b/internal/ingest/worker.go @@ -19,7 +19,6 @@ import ( "github.com/Wave-RF/WaveHouse/internal/chconn" "github.com/Wave-RF/WaveHouse/internal/chsql" "github.com/Wave-RF/WaveHouse/internal/mq" - "github.com/Wave-RF/WaveHouse/internal/query" "github.com/Wave-RF/WaveHouse/internal/tenant" "go.opentelemetry.io/otel" "go.opentelemetry.io/otel/attribute" @@ -747,7 +746,9 @@ func (w *IngestWorker) handleSuccess(ctx context.Context, tableName string, msgs // invalidate bumps the cache namespaces a batch of inserts into tableName // changed, for tenant id: the namespaces lead with the tenant (#583 story 8), // so the same table under another tenant keeps its cached results. id is the -// batch's tenant, read off each message's topic (story 5). +// batch's tenant, read off each message's topic (story 5). The table and +// scope go to the cache raw, as the structured-query read passes them; the +// cache escapes both sides alike. // // The set is the minimal one. Every msg here is for tableName, so a single // scopeless write bumps the whole table — which subsumes every scope — and @@ -755,24 +756,19 @@ func (w *IngestWorker) handleSuccess(ctx context.Context, tableName string, msgs // Doing this here (we already loop the batch once, and know it's one table) // keeps Cache.Invalidate a simple one-pass bump. func (w *IngestWorker) invalidate(ctx context.Context, id tenant.ID, tableName string, msgs []parsedMsg) { - encodedTable := query.SafeEncodeToken(tableName) seenScopes := make(map[string]struct{}, len(msgs)) namespaces := make([]cache.Namespace, 0, len(msgs)) for _, pm := range msgs { if pm.scope == "" { - namespaces = []cache.Namespace{{Tenant: id, Table: encodedTable}} + namespaces = []cache.Namespace{{Tenant: id, Table: tableName}} break } if _, exists := seenScopes[pm.scope]; exists { continue } seenScopes[pm.scope] = struct{}{} - namespaces = append(namespaces, cache.Namespace{ - Tenant: id, - Table: encodedTable, - Scope: query.SafeEncodeToken(pm.scope), - }) + namespaces = append(namespaces, cache.Namespace{Tenant: id, Table: tableName, Scope: pm.scope}) } if len(namespaces) == 0 { diff --git a/internal/ingest/worker_test.go b/internal/ingest/worker_test.go index e367327f4..b1e5fe4ce 100644 --- a/internal/ingest/worker_test.go +++ b/internal/ingest/worker_test.go @@ -457,7 +457,7 @@ func TestHandleSuccess(t *testing.T) { t.Parallel() // handleSuccess invalidates one namespace per distinct scope in the batch: the - // encoded table paired with the encoded scope (envelope.scope). Invalidate turns + // raw table paired with the raw scope (envelope.scope). Invalidate turns // an empty scope into a whole-table bump and a non-empty scope into a per-scope // bump, so the worker only needs to emit the table+scope pairs it saw. @@ -498,12 +498,12 @@ func TestHandleSuccess(t *testing.T) { wantNamespaces: []cache.Namespace{{Tenant: tenant.Default, Table: "events", Scope: ""}}, }, { - // Table and scope are percent-encoded so keys line up with the reader - // (internal/api/structured_query.go), which encodes the table too. - name: "table and scope are percent-encoded", + // Table and scope reach the cache raw, as the reader + // (internal/api/structured_query.go) passes them; the cache escapes. + name: "table and scope reach the cache raw", table: "events.staging", scopes: []string{"org.1"}, - wantNamespaces: []cache.Namespace{{Tenant: tenant.Default, Table: "events%2Estaging", Scope: "org%2E1"}}, + wantNamespaces: []cache.Namespace{{Tenant: tenant.Default, Table: "events.staging", Scope: "org.1"}}, }, { // Cache failure must not prevent ack — failure is logged, non-fatal. @@ -596,6 +596,52 @@ func TestInvalidate_ReachesOneTenantsEntries(t *testing.T) { assert.Equal(t, []byte("globex rows"), e.Value, "globex's entry survives acme's insert") } +// From the envelope the /v1/ingest producer publishes to the real cache: an +// insert into a table whose name holds a dot or a space — scoped or not — +// orphans the whole-table result a structured query on that table filed, +// under the raw name the request carries (the namespace internal/api's +// TestStructuredQuery_RawTableNameMeetsTheInsertsBump reads through), and +// leaves another table's. +func TestFlushTable_BumpsWhatTheReadFiles(t *testing.T) { + t.Parallel() + ok := &testutil.MockRoundTripper{Fn: func(*http.Request) (*http.Response, error) { + return &http.Response{StatusCode: http.StatusOK, Body: io.NopCloser(bytes.NewBufferString("OK"))}, nil + }} + for _, tc := range []struct{ table, scope string }{ + {"default.clicks", ""}, {"my table", ""}, {"default.clicks", "org.1"}, {"my table", "org 1"}, + } { + t.Run(tc.table+"/"+tc.scope, func(t *testing.T) { + t.Parallel() + l1, err := cache.NewLocal(1 << 20) + require.NoError(t, err) + t.Cleanup(func() { _ = l1.Close() }) + w, _, _, wait := newTestWorker(ok) + w.cache = l1 + + ctx := context.Background() + read := []cache.Namespace{{Tenant: tenant.Default, Table: tc.table}} + other := []cache.Namespace{{Tenant: tenant.Default, Table: "clicks"}} + for _, deps := range [][]cache.Namespace{read, other} { + _, snap, err := l1.Lookup(ctx, tenant.Default, "q", deps) + require.NoError(t, err) + require.NoError(t, l1.Set(ctx, snap, []byte(deps[0].Table), time.Minute)) + } + l1.Wait() + + msgs := parseAll(t, w, newIngestMsg(t, tc.table, tc.scope, map[string]any{"id": 1})) + w.flushTable(ctx, msgs[0].tableName, msgs) + wait() + + e, _, err := l1.Lookup(ctx, tenant.Default, "q", read) + require.NoError(t, err) + assert.Nil(t, e.Value, "the insert orphans the read's entry") + e, _, err = l1.Lookup(ctx, tenant.Default, "q", other) + require.NoError(t, err) + assert.Equal(t, []byte("clicks"), e.Value, "another table's entry survives") + }) + } +} + // --------------------------------------------------------------------------- // sendToDLQ — DLQ publish + headers + ack-only-on-success // --------------------------------------------------------------------------- diff --git a/internal/query/ident.go b/internal/query/ident.go deleted file mode 100644 index 35f032da9..000000000 --- a/internal/query/ident.go +++ /dev/null @@ -1,8 +0,0 @@ -package query - -import "github.com/Wave-RF/WaveHouse/internal/keyenc" - -// SafeEncodeToken renders a table or scope name as one dot-free token of the -// cache's namespace keys: keyenc's escaping, the same bytes a NATS subject -// carries for the name. -func SafeEncodeToken(raw string) string { return keyenc.Escape(raw) } diff --git a/internal/query/ident_test.go b/internal/query/ident_test.go deleted file mode 100644 index ac4fd64a8..000000000 --- a/internal/query/ident_test.go +++ /dev/null @@ -1,31 +0,0 @@ -package query - -import ( - "testing" - - "github.com/stretchr/testify/assert" -) - -func TestEncodeTable(t *testing.T) { - t.Parallel() - tests := []struct { - name string - raw string - expected string - }{ - {"safe string", "my_table123", "my_table123"}, - {"with dots", "default.clicks", "default%2Eclicks"}, - {"with spaces", "my table", "my%20table"}, - {"with dashes and slashes", "a-b/c", "a-b%2Fc"}, - {"empty string", "", ""}, - {"only safe characters", "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789_", "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789_"}, - } - - for _, tt := range tests { - t.Run(tt.name, func(t *testing.T) { - t.Parallel() - got := SafeEncodeToken(tt.raw) - assert.Equal(t, tt.expected, got) - }) - } -} diff --git a/internal/testutil/cachetest/cachetest.go b/internal/testutil/cachetest/cachetest.go index 10a74b393..acc24a134 100644 --- a/internal/testutil/cachetest/cachetest.go +++ b/internal/testutil/cachetest/cachetest.go @@ -46,6 +46,8 @@ func Run(t *testing.T, newCache func(t *testing.T) cache.Cache, opts Options) { {"tenant isolation", testTenantIsolation}, {"foreign dependency refused", testForeignDependency}, {"scope lattice", testScopeLattice}, + {"raw names read and bump alike", testRawNames}, + {"names never run together", testNamesApart}, {"invalidate counts namespaces", testInvalidateCount}, {"invalidate tenant orphans queries and pipes", testInvalidateTenant}, {"bump during the query orphans the fill", testBumpDuringQuery}, @@ -240,6 +242,54 @@ func testScopeLattice(t *testing.T, c cache.Cache) { requireHit(t, c, "orders/", acme, "q", orders) } +// Namespaces carry raw names, and the cache escapes them where it builds a +// key: a table or scope holding a dot, a space or a '%' is read and bumped +// under the one namespace both sides pass. +func testRawNames(t *testing.T, c cache.Cache) { + for _, table := range []string{"default.clicks", "my table", "100%"} { + whole, scoped := ns(acme, table, ""), ns(acme, table, "org.1") + fill(t, c, acme, "q", table, whole) + fill(t, c, acme, "q", table+"/org.1", scoped) + + invalidate(t, c, scoped) + assertMiss(t, c, acme, "q", scoped) + assertMiss(t, c, acme, "q", whole) + + fill(t, c, acme, "q", table, whole) + fill(t, c, acme, "q", table+"/org.1", scoped) + invalidate(t, c, whole) + assertMiss(t, c, acme, "q", whole) + assertMiss(t, c, acme, "q", scoped) + } +} + +// Two dependency sets whose names would run together under an unescaped +// join — at a '.', a ':', a '|' or a NUL, whichever a backend's key layout +// separates on — are two entries, and a bump of one leaves the other. +func testNamesApart(t *testing.T, c cache.Cache) { + pairs := []struct{ a, b []cache.Namespace }{ + {[]cache.Namespace{ns(acme, "a.0.b", "")}, []cache.Namespace{ns(acme, "a", "b.0.")}}, + {[]cache.Namespace{ns(acme, "a:b", "c")}, []cache.Namespace{ns(acme, "a", "b:c")}}, + {[]cache.Namespace{ns(acme, "a\x00b", "")}, []cache.Namespace{ns(acme, "a", "b\x00")}}, + {[]cache.Namespace{ns(acme, "x", ""), ns(acme, "y", "")}, []cache.Namespace{ns(acme, "x.0..0|acme.0.y", "")}}, + } + for i, p := range pairs { + sha := fmt.Sprintf("q%d", i) + fill(t, c, acme, sha, "a", p.a...) + fill(t, c, acme, sha, "b", p.b...) + requireHit(t, c, "a", acme, sha, p.a...) + + invalidate(t, c, p.a...) + assertMiss(t, c, acme, sha, p.a...) + requireHit(t, c, "b", acme, sha, p.b...) + + fill(t, c, acme, sha, "a", p.a...) + invalidate(t, c, p.b...) + assertMiss(t, c, acme, sha, p.b...) + requireHit(t, c, "a", acme, sha, p.a...) + } +} + func testInvalidateCount(t *testing.T, c cache.Cache) { n, err := c.Invalidate(context.Background(), []cache.Namespace{ns(acme, "events", ""), ns(globex, "events", "org_1")}) require.NoError(t, err) From 79ab364b1d70d16ff0325cffd8413c9bdde93892 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 19:45:58 -0400 Subject: [PATCH 23/79] test(mq): cover a byte left unescaped in the lenient-token case With '-' kept, "b-c" is what Escape writes, so the case no longer exercised a partly unescaped token; "b~c" does. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01B1tJWUp6oaDoH1usLwGtLF --- internal/mq/subject_test.go | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/internal/mq/subject_test.go b/internal/mq/subject_test.go index 28f4940be..6f859acdf 100644 --- a/internal/mq/subject_test.go +++ b/internal/mq/subject_test.go @@ -43,7 +43,7 @@ func TestSubject_Golden(t *testing.T) { // so a message it queued reads as the same topic. func TestParseTopicKey_LenientTokens(t *testing.T) { t.Parallel() - assert.Equal(t, Topic{Tenant: "a", Table: "b-c", Scope: "d.e"}, parseTopicKey("a.b-c.d%2ee")) + assert.Equal(t, Topic{Tenant: "a", Table: "b~c", Scope: "d.e"}, parseTopicKey("a.b~c.d%2ee")) assert.Equal(t, parseTopicKey("a.table-with-dashes.org-1"), parseTopicKey("a.table%2Dwith%2Ddashes.org%2D1")) } From 40d817d44dc4de8dba61ab804042262f3385eb26 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 19:47:09 -0400 Subject: [PATCH 24/79] refactor(cache): the Redis codec escapes the raw names it keys A Namespace now carries raw table and scope names, so the Redis codec builds its token keys with keyenc.AppendJoin after the fixed ":{tenant}:B:" and ":S:" pieces; the {tenant} hash tag is placed as it was, since a tenant id is its own escaped form. A ':' in a table or scope no longer reads as the separator: table "a:b" with scope "c" and table "a" with scope "b:c" had one scope token. valueKey hashed each dep's raw table and scope ended by NULs, so a name holding a NUL could hash the same input as another set of names: table "a\x00b" and table "a" with scope "b\x00" shared a value key. It now hashes the escaped sha and each dep's escaped, joined form, each ended by a NUL, which escaping never writes. The key layout is otherwise unchanged. The conformance cases added with the raw-name change run against Redis, Valkey, Dragonfly and a Redis Cluster node, and the old codec fails "names never run together" on all four. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01B1tJWUp6oaDoH1usLwGtLF --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 4 ++-- internal/cache/redis_codec.go | 24 +++++++++++-------- internal/cache/redis_codec_test.go | 33 +++++++++++++++++++++++---- 4 files changed, 46 insertions(+), 17 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 4227d58fa..3c1edc923 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A Redis-compatible shared cache backend, built but not yet selectable** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row open a circuit breaker that skips the server until a probe succeeds, and deferred invalidations are retried until they land, collapsing to one tenant-wide bump per tenant past 100,000 keys. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. Nothing selects it yet: the `cache.backend` switch and the `cache.redis.*` settings arrive with the wiring (E4), so every deployment still runs `LocalCache`. Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression and server-stops-answering cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. +- **A Redis-compatible shared cache backend, built but not yet selectable** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row open a circuit breaker that skips the server until a probe succeeds, and deferred invalidations are retried until they land, collapsing to one tenant-wide bump per tenant past 100,000 keys. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. Nothing selects it yet: the `cache.backend` switch and the `cache.redis.*` settings arrive with the wiring (E4), so every deployment still runs `LocalCache`. Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression and server-stops-answering cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index cb96ae2b8..975db37ee 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -113,8 +113,8 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on, each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope by their raw names, which the cache escapes where it builds a key: one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)). `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. - **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: a namespace key is `..
.
.`, its fields joined with `keyenc.Join` so a dot in a table or scope name is escaped rather than read as a separator, and a query key, `|.||…` with the sha escaped, folds in the tenant's version and each dependency's namespace key and namespace version, so bumping a table (a scopeless write) or one scope — scope is reserved and empty today, so every write is the whole-table bump — orphans every dependent entry without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results, and the flat directory's single tenant simply carries the `0` prefix. `BumpTenant` (behind `InvalidateTenant`) advances the tenant version that leads every namespace key of one tenant, orphaning its every namespace and every cached query in one step (the tenant version is folded into every query key, so a pipe result with no dependencies is orphaned too) — a table no bump ever keyed included, which an enumeration of the index would miss — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. -- **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET`, `MGET` and `PING`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token (a string of another length, a list, a hash) is replaced, so a token lost to eviction, expiry or a restart is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing. Nothing selects this backend yet — the boot config and wiring arrive with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s E4. -- **redis_codec.go** — the key schema (`tokenKeys`, `bumpKeys`, `valueKey`) and the value frame: format, flags, expiry, the token list, the payload, zstd-compressed from `CompressMinBytes` when that is smaller. A stored value is capped at `MaxValueBytes` and a decoded one at eight times that, which refuses a zip bomb planted in a shared server; an unknown format, as a newer process writes during a rolling upgrade, is a miss. +- **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET`, `MGET` and `PING`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot; the table and scope are escaped with `keyenc` after the fixed prefix, so a `:` in a name is never read as the separator. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, the hash over their escaped forms, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token (a string of another length, a list, a hash) is replaced, so a token lost to eviction, expiry or a restart is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing. Nothing selects this backend yet — the boot config and wiring arrive with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s E4. +- **redis_codec.go** — the key schema (`tokenKeys`, `bumpKeys`, `valueKey`, which take a `Namespace`'s raw names and escape them) and the value frame: format, flags, expiry, the token list, the payload, zstd-compressed from `CompressMinBytes` when that is smaller. A stored value is capped at `MaxValueBytes` and a decoded one at eight times that, which refuses a zip bomb planted in a shared server; an unknown format, as a newer process writes during a rolling upgrade, is a miss. - **breaker.go** — the circuit breaker: `BreakerThreshold` (5) consecutive failures open it, and while open every operation skips the server at once; after `BreakerOpenFor` (5 s) one background `PING` decides whether it closes. Only a transport failure or a timeout counts against the server: any reply, an error reply or one the backend cannot use included, counts as a success, for operations and the probe alike, and a caller that gave up first counts as nothing. - **pending.go** — invalidations the server did not take, kept per token key (repeats coalesce) and retried by a background loop with backoff from 100 ms to 10 s until they land. Landing late is still correct: a fresh token orphans the pre-write entries and any fill made meanwhile. Past `PendingMax` (100,000) keys the set collapses to one tenant bump per affected tenant. `Close` makes one last attempt at the bumps still owed, past the breaker (an open one is why they are owed); what that attempt cannot deliver is lost, and the entries they would orphan are served until their TTL — the same failure as a worker stopping between an insert and its invalidation. - **metrics.go** — the shared backend's instruments (meter `wavehouse-cache`, `backend="redis"`, no tenant attribute): `wavehouse_cache_lookups_total` by `result` (`hit`, `miss`, `stale` — filed under since-bumped tokens, `bypass` — server skipped, `error`), `wavehouse_cache_op_duration_seconds` by `op` (`lookup`, `set`, `invalidate`), `wavehouse_cache_breaker_open` (1 while bypassed, including before the first connection), `wavehouse_cache_invalidations_total` by `result` (`ok`, `deferred`), `wavehouse_cache_invalidations_pending` (bumps owed: entries they would orphan may be served stale meanwhile), `wavehouse_cache_value_bytes` (stored size), `wavehouse_cache_oversize_total` and `wavehouse_cache_set_failures_total` by `reason` (`oom`, `timeout`, `other`). diff --git a/internal/cache/redis_codec.go b/internal/cache/redis_codec.go index cf6f4b5ec..1ce16f0cb 100644 --- a/internal/cache/redis_codec.go +++ b/internal/cache/redis_codec.go @@ -13,6 +13,7 @@ import ( "github.com/klauspost/compress/zstd" + "github.com/Wave-RF/WaveHouse/internal/keyenc" "github.com/Wave-RF/WaveHouse/internal/tenant" ) @@ -44,17 +45,20 @@ func newToken() []byte { // tenantTokenKey is the key of tenant id's token. Every token key carries // the tenant as a hash tag, so all of a tenant's tokens share one cluster -// slot and a lookup reads them with one MGET. +// slot and a lookup reads them with one MGET. The tenant goes in verbatim: +// its grammar is keyenc's kept bytes, so it is its own escaped form. func tenantTokenKey(prefix string, id tenant.ID) string { return prefix + ":{" + string(id) + "}:T" } +// tableTokenKey and scopeTokenKey take raw names and escape them after the +// fixed prefix (keyenc), so no ':' in a name reads as the separator. func tableTokenKey(prefix string, id tenant.ID, table string) string { - return prefix + ":{" + string(id) + "}:B:" + table + return string(keyenc.AppendJoin([]byte(prefix+":{"+string(id)+"}:B:"), ':', table)) } func scopeTokenKey(prefix string, id tenant.ID, table, scope string) string { - return prefix + ":{" + string(id) + "}:S:" + table + ":" + scope + return string(keyenc.AppendJoin([]byte(prefix+":{"+string(id)+"}:S:"), ':', table, scope)) } // sortedDeps returns deps in canonical order without duplicates. @@ -91,16 +95,16 @@ func bumpKeys(prefix string, ns Namespace) []string { // valueKey names the entry for sha over deps. It carries no versions, so a // refill overwrites in place, and no hash tag, so one tenant's values spread -// across a cluster's shards. +// across a cluster's shards. It hashes the escaped sha and each dep's +// escaped, joined table and scope, each ended by a NUL, which escaping never +// writes, so no two sets of names hash the same input. func valueKey(prefix string, id tenant.ID, sha string, deps []Namespace) string { h := sha256.New() - h.Write([]byte(sha)) - h.Write([]byte{0}) + b := keyenc.AppendEscape(nil, sha) + h.Write(append(b, 0)) for _, d := range sortedDeps(deps) { - h.Write([]byte(d.Table)) - h.Write([]byte{0}) - h.Write([]byte(d.Scope)) - h.Write([]byte{0}) + b = keyenc.AppendJoin(b[:0], ':', d.Table, d.Scope) + h.Write(append(b, 0)) } return prefix + ":q:" + string(id) + ":" + hex.EncodeToString(h.Sum(nil)) } diff --git a/internal/cache/redis_codec_test.go b/internal/cache/redis_codec_test.go index 68c423d0f..c956a6097 100644 --- a/internal/cache/redis_codec_test.go +++ b/internal/cache/redis_codec_test.go @@ -28,6 +28,11 @@ func TestTokenKeys(t *testing.T) { []Namespace{{Tenant: "acme", Table: "events"}}, []string{"wh:{acme}:T", "wh:{acme}:B:events", "wh:{acme}:S:events:"}, }, + { + "names arrive raw and are escaped after the hash tag", + []Namespace{{Tenant: "acme", Table: "default.clicks", Scope: "org:1"}}, + []string{"wh:{acme}:T", "wh:{acme}:B:default%2Eclicks", "wh:{acme}:S:default%2Eclicks:org%3A1"}, + }, { "shared table tokens and duplicate deps appear once, sorted", []Namespace{ @@ -55,13 +60,26 @@ func TestBumpKeys(t *testing.T) { assert.Equal(t, []string{"p:{acme}:B:events"}, bumpKeys("p", Namespace{Tenant: "acme", Table: "events"})) assert.Equal(t, []string{"p:{acme}:S:events:org_1", "p:{acme}:S:events:"}, bumpKeys("p", Namespace{Tenant: "acme", Table: "events", Scope: "org_1"})) + assert.Equal(t, []string{"p:{acme}:S:my%20table:a%3Ab", "p:{acme}:S:my%20table:"}, + bumpKeys("p", Namespace{Tenant: "acme", Table: "my table", Scope: "a:b"})) + + // A ':' in a name never reads as the separator: table "a:b" with scope + // "c" and table "a" with scope "b:c" bump two tokens. + assert.NotEqual(t, + bumpKeys("p", Namespace{Tenant: "acme", Table: "a:b", Scope: "c"}), + bumpKeys("p", Namespace{Tenant: "acme", Table: "a", Scope: "b:c"})) } // Every key a bump writes is one some lookup reads: otherwise the bump // orphans nothing. func TestBumpKeysAreReadByLookups(t *testing.T) { t.Parallel() - for _, ns := range []Namespace{{Tenant: "acme", Table: "events"}, {Tenant: "acme", Table: "events", Scope: "org_1"}} { + for _, ns := range []Namespace{ + {Tenant: "acme", Table: "events"}, + {Tenant: "acme", Table: "events", Scope: "org_1"}, + {Tenant: "acme", Table: "default.clicks"}, + {Tenant: "acme", Table: "my table", Scope: "org:1"}, + } { read := map[string]bool{} for _, scope := range []string{"", ns.Scope} { for _, k := range tokenKeys("wh", "acme", []Namespace{{Tenant: "acme", Table: ns.Table, Scope: scope}}) { @@ -83,9 +101,16 @@ func TestValueKey(t *testing.T) { assert.Equal(t, k, valueKey("wh", "acme", "acme:query:abc", []Namespace{b, a, b}), "order and duplicates do not matter") assert.NotEqual(t, k, valueKey("wh", "acme", "acme:query:abc", []Namespace{a})) assert.NotEqual(t, k, valueKey("wh", "acme", "acme:query:abd", []Namespace{a, b})) - assert.NotEqual(t, - valueKey("wh", "acme", "q", []Namespace{{Tenant: "acme", Table: "ab", Scope: "c"}}), - valueKey("wh", "acme", "q", []Namespace{{Tenant: "acme", Table: "a", Scope: "bc"}})) + // Names that would run together unescaped hash apart: at the table and + // scope boundary, at a NUL, and across dependencies. + for _, pair := range [][2][]Namespace{ + {{{Tenant: "acme", Table: "ab", Scope: "c"}}, {{Tenant: "acme", Table: "a", Scope: "bc"}}}, + {{{Tenant: "acme", Table: "a\x00b"}}, {{Tenant: "acme", Table: "a", Scope: "b\x00"}}}, + {{{Tenant: "acme", Table: "a:b"}}, {{Tenant: "acme", Table: "a", Scope: "b"}}}, + {{{Tenant: "acme", Table: "a"}, {Tenant: "acme", Table: "b"}}, {{Tenant: "acme", Table: "a", Scope: "\x00b\x00"}}}, + } { + assert.NotEqual(t, valueKey("wh", "acme", "q", pair[0]), valueKey("wh", "acme", "q", pair[1]), "%q", pair) + } } func newTestCodec(t *testing.T, compressMin, maxDecoded int) *codec { From 972a53d4dff10c6e262b7a81dc7e98781ba10a2b Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 20:08:07 -0400 Subject: [PATCH 25/79] docs(cache): one name for keyenc's cache keys; the entry key's lead field MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit keyenc's cache user is "cache keys" everywhere (the package doc, AGENTS.md, architecture.md, development.md): it now escapes the caller's query key and joins the tenant segment as well as the version index's keys. AGENTS.md and architecture.md say the stored entry key leads with the caller's query key escaped whole, `|.||…`; only the singleflight key is literally `:query:`. The Cache doc puts the namespace count next to its noun, the pipe-result sentence in deployment.md and the sharedTables comment no longer reads "between those" as a time span, a test comment drops "table-keyed", and the Unreleased MQ-boundary entry no longer says the cache keeps query.SafeEncodeToken. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01B1tJWUp6oaDoH1usLwGtLF --- AGENTS.md | 6 +++--- CHANGELOG.md | 4 ++-- docs/src/content/docs/architecture.md | 8 ++++---- docs/src/content/docs/deployment.md | 2 +- docs/src/content/docs/development.md | 2 +- internal/app/app_test.go | 4 ++-- internal/app/wire.go | 7 ++++--- internal/keyenc/keyenc.go | 4 ++-- 8 files changed, 19 insertions(+), 18 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 1ee062cc7..1c167da27 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -31,14 +31,14 @@ Nineteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers - **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` handing the sweeper each tenant's own gap window (a rejected tenant's as its folder last had it, unbounded for one rejected since boot) and the `mq.max_bytes_gb` reconcile each served tenant's byte budget, and `defaultPolicy` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, and the same hook's `Hub.Prune` ends the open streams of a tenant no longer served), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `cmd/wavehouse` and `tests/integration` both boot through it - **`auth/`** — JWT auth middleware: HMAC **or** JWKS verification with `alg` pinned to the active verifier, role extraction from a configurable claim path; always runs, never rejects (bad token → empty role + stashed reason). One verifier per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 9): `Authenticator` keys them by `tenant.ID` — the request store's `settings.Store.Tenant()`, through an injected `TenantSource`; `tenant.Default` on the tenant-exempt routes — built from each tenant's `auth` block by `Reconfigure`, dropped by `Prune` once the tenant stops being served (rejected or removed), released by `Close`; the secrets (`Config`) are boot-level and shared. A JWKS key set is fetched off the boot and reload paths: until one has been stored the verifier is pending and a token-bearing request gets `503` + `Retry-After` from `api.refuseUnverifiable` (`auth.ErrVerifierPending`), never a `default_role` evaluation; refresh is library-managed (Eric, 2026-09-22), response capped at 1 MiB; the operator key's admin role is the request tenant's -- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, `..
.
.` for a namespace, each field escaped by `keyenc` where the key is built (a `Namespace` carries the raw table and scope, so no caller escapes) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version folded into every namespace and query key of one tenant, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read and `Set` files the fill under it, so a write landing mid-query orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) +- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for the caller's query key and its singleflight (escaped whole as the lead field of the stored key, `|.||…`), `..
.
.` for a namespace, each field escaped by `keyenc` where the key is built (a `Namespace` carries the raw table and scope, so no caller escapes) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version folded into every namespace and query key of one tenant, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read and `Set` files the fill under it, so a write landing mid-query orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) - **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability — boot is the validator, there is no dry run - **`dedupe/`** — `Deduplicator` interface → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) -- **`keyenc/`** — the one escaping composite keys are built from: `Escape` keeps `[A-Za-z0-9_-]` (exactly the tenant-id grammar, so a tenant id is its own escaped form) and writes every other byte as `%XX`, `Unescape` is `url.PathUnescape` (lenient: either hex case, and a byte left unescaped reads as itself, so a `%2D` an earlier build wrote still reads), `Join`/`AppendJoin` escape each field and put a separator between them (they panic on no fields, and on a separator the escaping could write or one outside ASCII) and `Split` reverses them. The package that builds a key takes raw names and escapes them itself, so no caller has to and no field reaches a key unescaped: NATS subjects (`Join`/`Split` after the verbatim tenant) and the cache's version-index keys (`cache.VersionManager`) use it; changing what it keeps orphans every stored key +- **`keyenc/`** — the one escaping composite keys are built from: `Escape` keeps `[A-Za-z0-9_-]` (exactly the tenant-id grammar, so a tenant id is its own escaped form) and writes every other byte as `%XX`, `Unescape` is `url.PathUnescape` (lenient: either hex case, and a byte left unescaped reads as itself, so a `%2D` an earlier build wrote still reads), `Join`/`AppendJoin` escape each field and put a separator between them (they panic on no fields, and on a separator the escaping could write or one outside ASCII) and `Split` reverses them. The package that builds a key takes raw names and escapes them itself, so no caller has to and no field reaches a key unescaped: NATS subjects (`Join`/`Split` after the verbatim tenant) and the cache's keys (`cache.VersionManager`) use it; changing what it keeps orphans every stored key - **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`, `deadletter.go`), whose subject tokens are escaped by the shared `internal/keyenc`; `internal/app` constructs it and hands everything else a `mq.Broker` - **`observability/`** — OpenTelemetry pipeline: `InitProvider` wires trace/metric/log providers via OTLP gRPC (each signal independently gated). A top-level `Prometheus` config block drives an optional `/metrics` scrape endpoint that runs independently of OTLP push — standalone (Alloy/Mimir scrape, no collector), alongside OTLP, or off. `NewLogger` produces a slog handler that fans out to stdout AND OTLP (stdout always 100%, OTLP sample-rate-aware). `TraceHandler` injects trace_id/span_id from active spans. `tracer.go` provides W3C trace context propagation over message headers (`InjectHeaders`/`ExtractHeaders` on a plain header map; `internal/mq` injects on every publish and extracts on the `Subscribe` path, so this package never sees a NATS type). - **`pipes/`** — Named query pipes: `NamedQuery` type + `BindParams` + `Source` (read per request; `settings.Store` in production, `Static(q...)` in tests) @@ -436,7 +436,7 @@ internal/config/ → Configuration structs + loader internal/dedupe/ → Optional deduplication (interface + embedded/distributed) internal/discovery/ → ClickHouse schema introspection + ingest validation internal/ingest/ → Batch buffer with DLQ + Active Sweeper (NATS message lifecycle) -internal/keyenc/ → One escaping for composite keys (NATS subject tokens, cache namespace tokens) +internal/keyenc/ → One escaping for composite keys (NATS subject tokens, cache keys) internal/mq/ → MQ boundary (the only NATS/JetStream importer: owned message/consumer/stream types + embedded server) internal/observability/ → OpenTelemetry pipeline (traces/metrics/logs providers, Prometheus exporter, slog fan-out, message-header trace propagation) internal/pipes/ → Named query pipes (types, parameter binding, Source) diff --git a/CHANGELOG.md b/CHANGELOG.md index 70dd7efd9..2c8dce0b6 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -48,7 +48,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **The docs site now consumes the *published* `@wavehouse/sdk`, not the workspace one** (`docs/package.json`, `pnpm-workspace.yaml`, `Makefile`, `scripts/classify-paths.sh`). The landing page's live demo streams against a separately-deployed backend on its own release cadence, but took its SDK from the tree — so this release's wire change would have reached the deployed site the moment it merged, while the backend still spoke the old envelope: the panel keeps reporting "live" while every frame is dropped for want of a schema announcement, with no `error` callback ([#568](https://github.com/Wave-RF/WaveHouse/issues/568)). `docs` now pins `^0.1.1` from the registry, which takes `0.1.x` patches and stops short of `0.2.0`, so moving the site onto the new wire is a deliberate bump lined up with tagging the SDK release rather than a side effect of merging. `tests/e2e/sdk` deliberately keeps `workspace:*`. Depending on our own package from the registry also made `minimumReleaseAge` apply to it for the first time, and the exclude list named only `@wave-rf/*` (the plugin scope), so a freshly tagged SDK would have been uninstallable by the docs site for seven days — `@wavehouse/*` is now exempt too. The docs build no longer needs `build-ts`. -- **The MQ boundary is sealed: only `internal/mq` imports NATS/JetStream** (`internal/mq/{mq,embedded,subject,purge}.go`, `internal/api/{dlq,stream,ingest,structured_query,router}.go`, `internal/ingest/{worker,sweeper}.go`, `internal/query/ident.go`, `internal/app/{app,wire}.go`, `internal/cache/version_manager.go`, `internal/observability/{tracer,metrics}.go`, `internal/testutil/mocks.go`, `.golangci.yml`, `docs/src/content/docs/{architecture,development,ingest-pipeline}.md`, `AGENTS.md`; part of [#583](https://github.com/Wave-RF/WaveHouse/issues/583), story 4): `internal/api`, `internal/ingest`, `internal/cache`, `internal/observability`, and `internal/testutil` each reached past the `Publisher`/`Subscriber` interfaces for raw NATS types — the `jetstream.JetStream` handle for the DLQ handler, the sweeper, and SSE gap-fill; a `*nats.Conn` for the ingest worker; `jetstream.Msg` in the worker's per-table batches; the `*server.Server` for the system gauges; `nats.Header` in the trace propagator; a never-wired `*nats.Conn` on the cache's `VersionManager` — so the tenant subject token (story 5) would have touched a dozen call sites in five packages. Every one of those now goes through an mq-owned type, and the boundary is semantic as well as an import rule — nothing outside `internal/mq` builds a subject, names a stream, or reasons in sequences. Events are addressed by `mq.Topic{Table, Scope}` with raw names (the `ingest.`/`dlq.` prefixes, the `>` wildcards, the stream names, and the subject-token encoder — formerly `query.SafeEncodeNATS`/`SafeDecodeNATS` — are private to `internal/mq/subject.go`; a delivered `Message` exposes `TopicKey()`/`Topic()` instead of a subject, keeping the delivered form so the per-message path and the DLQ prefix swap decode nothing), and the interfaces state intent: `mq.Headers` (which `PublishOpt`s such as `WithHeader` now shape, rather than a `*nats.Msg`), `mq.ConsumerManager` → `Consumer` from a `ConsumerConfig` (the worker's durable pull consumer, with the same ack-wait, max-ack-pending, and prefetch), `mq.Purger.PurgeAcked` (drop what is both acked by a consumer and stored before a cutoff — the Active Sweeper's ack-floor/binary-search arithmetic moved from `internal/ingest/sweeper.go` to `internal/mq/purge.go`, leaving the sweeper the schedule and the gap window), `mq.DeadLetterer.DeadLetter` and `mq.DeadLetterStats.DeadLetterCounts` (park a message under the topic it arrived on; count what is parked — the DLQ handler no longer reads stream state), `Publisher.Publish` reporting a full ingest stream as `mq.ErrQueueFull` (the API's 503 + `Retry-After` no longer matches on the broker's error text), `mq.Replayer.ReplaySince` (the SSE gap-fill's `DeliverByStartTime` consumer), `mq.Broker` composing all of it — `internal/app` holds that, not `*mq.EmbeddedNATS` — with `SetMaxBytes` (the `mq.max_bytes_gb` reload, moved out of `internal/app` together with its time bounds and rollback: the ingest stream and the DLQ at a tenth of it are resized as a pair, and `NewEmbedded` now creates both, replacing `api.EnsureDLQStream` and the standalone `Resize`), and `Stats` feeding `observability.RegisterSystemMetrics`, which takes a `func() (MQStats, error)` instead of the server. The raw accessors `JetStream()`, `NatsConn()`, and `GetServer()` are gone, the never-read `api.Dependencies.JS` and the never-set `cache.NewVersionManager` connection parameter with them, and the trace propagator is `InjectHeaders`/`ExtractHeaders` over a plain header map that `internal/mq` injects on every publish and extracts on the `Subscribe` path (the worker's consumer path never read a per-message context, before or after). Behavior-preserving: subjects, stream names, consumer settings, ack semantics, and the `X-DLQ-*` headers are what they were, and no envelope or endpoint shape changes. Seven differences, none on the wire today: the sweeper now warns only when the buffer consumer genuinely does not exist yet (`mq.ErrConsumerNotFound`) and errors on any other lookup failure, where it used to warn on both; the DLQ republish now goes through `DeadLetterer.DeadLetter`, which like every publish injects W3C trace headers when its context carries a span — the worker's flush context never does, so a DLQ copy's headers are unchanged in practice (the `X-DLQ-*` set and the envelope bytes are untouched either way); the ingest worker's delivery handoff now also watches its context, so a delivery still in flight when the worker stops returns instead of pinning the client's delivery goroutine on a full channel (the message is unacked and redelivered, as a dropped one always was); and SSE gap-fill distinguishes "caught up" (the client's no-messages or request-timeout answer) from a pull that fails outright (a closed connection, a deleted consumer) — the stream still falls through to live events either way, but the second is now logged at `WARN` where it used to end the replay silently as if it were the first; `GET /v1/ops/dlq/stats` returns the documented `500` when the broker cannot be read, where any stream-lookup failure used to read as an empty queue (a genuinely absent queue, `mq.ErrNoDeadLetterQueue`, still does); and the sweeper's binary search no longer discards the lower half of the stream when its midpoint holds no message, which could place the purge bound inside the replay window — a sequence with no message is now kept as a candidate bound (purging less, never more), and a lookup that fails outright aborts that sweep instead of being read as a missing sequence; and a consumer that dies underneath the ingest worker no longer stalls ingestion silently — the broker client reports a deleted consumer or a closed connection only through an asynchronous error callback that was never wired (before this PR either), so `mq.Consumer.Consume` now also returns a `failed` channel (`mq.ErrDeliveryEnded`, wrapping the broker's reason), the worker flushes and acks what it holds and reports the failure (as does a consumer that cannot start), and `internal/app` returns it from `Run`, stopping the process so the supervisor's restart recreates the consumer; passing conditions reported through the same callback are logged at `WARN` ([#587](https://github.com/Wave-RF/WaveHouse/issues/587)). The cache keeps its namespace-token encoder as the broker-neutral `query.SafeEncodeToken` (same output, so cache keys are unchanged). The boundary is enforced, not just documented — a `depguard` rule in `.golangci.yml` fails `make lint` on any `github.com/nats-io` import outside `internal/mq` (Key Design Decision #20). The shared test mocks follow: `testutil.MockMessage`, `MockPurger`, and `MockDeadLetterStats` (against the mq interfaces) replace `MockJetStreamMsg`, `MockJetStream`, `MockStream`, and `MockConsumer` (against the upstream ones), and `MockPublisher` now records the topic and headers of every publish and dead-letter parking. +- **The MQ boundary is sealed: only `internal/mq` imports NATS/JetStream** (`internal/mq/{mq,embedded,subject,purge}.go`, `internal/api/{dlq,stream,ingest,structured_query,router}.go`, `internal/ingest/{worker,sweeper}.go`, `internal/query/ident.go`, `internal/app/{app,wire}.go`, `internal/cache/version_manager.go`, `internal/observability/{tracer,metrics}.go`, `internal/testutil/mocks.go`, `.golangci.yml`, `docs/src/content/docs/{architecture,development,ingest-pipeline}.md`, `AGENTS.md`; part of [#583](https://github.com/Wave-RF/WaveHouse/issues/583), story 4): `internal/api`, `internal/ingest`, `internal/cache`, `internal/observability`, and `internal/testutil` each reached past the `Publisher`/`Subscriber` interfaces for raw NATS types — the `jetstream.JetStream` handle for the DLQ handler, the sweeper, and SSE gap-fill; a `*nats.Conn` for the ingest worker; `jetstream.Msg` in the worker's per-table batches; the `*server.Server` for the system gauges; `nats.Header` in the trace propagator; a never-wired `*nats.Conn` on the cache's `VersionManager` — so the tenant subject token (story 5) would have touched a dozen call sites in five packages. Every one of those now goes through an mq-owned type, and the boundary is semantic as well as an import rule — nothing outside `internal/mq` builds a subject, names a stream, or reasons in sequences. Events are addressed by `mq.Topic{Table, Scope}` with raw names (the `ingest.`/`dlq.` prefixes, the `>` wildcards, the stream names, and the subject-token encoder — formerly `query.SafeEncodeNATS`/`SafeDecodeNATS` — are private to `internal/mq/subject.go`; a delivered `Message` exposes `TopicKey()`/`Topic()` instead of a subject, keeping the delivered form so the per-message path and the DLQ prefix swap decode nothing), and the interfaces state intent: `mq.Headers` (which `PublishOpt`s such as `WithHeader` now shape, rather than a `*nats.Msg`), `mq.ConsumerManager` → `Consumer` from a `ConsumerConfig` (the worker's durable pull consumer, with the same ack-wait, max-ack-pending, and prefetch), `mq.Purger.PurgeAcked` (drop what is both acked by a consumer and stored before a cutoff — the Active Sweeper's ack-floor/binary-search arithmetic moved from `internal/ingest/sweeper.go` to `internal/mq/purge.go`, leaving the sweeper the schedule and the gap window), `mq.DeadLetterer.DeadLetter` and `mq.DeadLetterStats.DeadLetterCounts` (park a message under the topic it arrived on; count what is parked — the DLQ handler no longer reads stream state), `Publisher.Publish` reporting a full ingest stream as `mq.ErrQueueFull` (the API's 503 + `Retry-After` no longer matches on the broker's error text), `mq.Replayer.ReplaySince` (the SSE gap-fill's `DeliverByStartTime` consumer), `mq.Broker` composing all of it — `internal/app` holds that, not `*mq.EmbeddedNATS` — with `SetMaxBytes` (the `mq.max_bytes_gb` reload, moved out of `internal/app` together with its time bounds and rollback: the ingest stream and the DLQ at a tenth of it are resized as a pair, and `NewEmbedded` now creates both, replacing `api.EnsureDLQStream` and the standalone `Resize`), and `Stats` feeding `observability.RegisterSystemMetrics`, which takes a `func() (MQStats, error)` instead of the server. The raw accessors `JetStream()`, `NatsConn()`, and `GetServer()` are gone, the never-read `api.Dependencies.JS` and the never-set `cache.NewVersionManager` connection parameter with them, and the trace propagator is `InjectHeaders`/`ExtractHeaders` over a plain header map that `internal/mq` injects on every publish and extracts on the `Subscribe` path (the worker's consumer path never read a per-message context, before or after). Behavior-preserving: subjects, stream names, consumer settings, ack semantics, and the `X-DLQ-*` headers are what they were, and no envelope or endpoint shape changes. Seven differences, none on the wire today: the sweeper now warns only when the buffer consumer genuinely does not exist yet (`mq.ErrConsumerNotFound`) and errors on any other lookup failure, where it used to warn on both; the DLQ republish now goes through `DeadLetterer.DeadLetter`, which like every publish injects W3C trace headers when its context carries a span — the worker's flush context never does, so a DLQ copy's headers are unchanged in practice (the `X-DLQ-*` set and the envelope bytes are untouched either way); the ingest worker's delivery handoff now also watches its context, so a delivery still in flight when the worker stops returns instead of pinning the client's delivery goroutine on a full channel (the message is unacked and redelivered, as a dropped one always was); and SSE gap-fill distinguishes "caught up" (the client's no-messages or request-timeout answer) from a pull that fails outright (a closed connection, a deleted consumer) — the stream still falls through to live events either way, but the second is now logged at `WARN` where it used to end the replay silently as if it were the first; `GET /v1/ops/dlq/stats` returns the documented `500` when the broker cannot be read, where any stream-lookup failure used to read as an empty queue (a genuinely absent queue, `mq.ErrNoDeadLetterQueue`, still does); and the sweeper's binary search no longer discards the lower half of the stream when its midpoint holds no message, which could place the purge bound inside the replay window — a sequence with no message is now kept as a candidate bound (purging less, never more), and a lookup that fails outright aborts that sweep instead of being read as a missing sequence; and a consumer that dies underneath the ingest worker no longer stalls ingestion silently — the broker client reports a deleted consumer or a closed connection only through an asynchronous error callback that was never wired (before this PR either), so `mq.Consumer.Consume` now also returns a `failed` channel (`mq.ErrDeliveryEnded`, wrapping the broker's reason), the worker flushes and acks what it holds and reports the failure (as does a consumer that cannot start), and `internal/app` returns it from `Run`, stopping the process so the supervisor's restart recreates the consumer; passing conditions reported through the same callback are logged at `WARN` ([#587](https://github.com/Wave-RF/WaveHouse/issues/587)). The cache no longer borrows the subject encoder: it escapes the raw names in its own keys with `internal/keyenc`, which the broker's subjects use too. The boundary is enforced, not just documented — a `depguard` rule in `.golangci.yml` fails `make lint` on any `github.com/nats-io` import outside `internal/mq` (Key Design Decision #20). The shared test mocks follow: `testutil.MockMessage`, `MockPurger`, and `MockDeadLetterStats` (against the mq interfaces) replace `MockJetStreamMsg`, `MockJetStream`, `MockStream`, and `MockConsumer` (against the upstream ones), and `MockPublisher` now records the topic and headers of every publish and dead-letter parking. - **Policy format v2: `policies.json` is role-first, and the two operations are separate permission types** (BREAKING; `internal/policy/policy.go`, `internal/policy/rowfilter.go`, `internal/settings/validate.go`, `internal/query/builder.go`, `internal/api/{ingest,structured_query}.go`, `internal/stream/hub.go`, `clients/ts/src/{types,index}.ts`, `deployments/compose/settings/policies.json`, `docs/src/content/docs/{access-control.mdx,settings-directory.mdx,architecture.md}`, `AGENTS.md`, `tests/e2e/sdk/`): a table entry was keyed `tables.
.select.`; it is now keyed `tables.
..select`. A role appears once per table and its grant carries two optional blocks, so a role that could both read and write no longer has to be written out twice, and "this role has no insert grant" is a missing block rather than an absence you have to notice in a second map. The blocks are now distinct types rather than one struct whose halves were inert per operation: `select` takes `allow_columns`, `deny_columns`, `filter`, `allowed_aggregations`, `denied_aggregations` and the four `max_*` limits; `insert` takes `allow_columns`, `deny_columns`, `check`. Field names and semantics are unchanged — only the nesting moves — but a field on the wrong side that the old layout *accepted* — the four `max_*` limits and the two aggregation rules — is now a validation error instead of being silently ignored. `filter` under an `insert` grant and `check` under a `select` one are rejected too, but that is not new here — [#541](https://github.com/Wave-RF/WaveHouse/pull/541), also unreleased, added the runtime check; the split types now refuse them one layer earlier, as unknown keys at the strict decode. Upgrading from **0.1.0**, though, none of the eight were enforced: a 0.1.0 policy could carry an insert-side `filter` that resolved into a `WHERE` the insert path never read, as well as an ignored limit. Converting to the role-first layout drops both. Internally `ResolvedPermissions` splits the same way (`.Select` / `.Insert`), and `IsColumnAllowed` takes the side to consult, which is what stops the read allowlist from ever answering a write question or vice versa. **There is no automatic conversion** — the settings files are the source of truth and WaveHouse has no write path back to them — so `policies.json` must be converted by hand; run `wavehouse validate` before restarting. A document still in the old layout is reported as one clear finding naming the table and operation and pointing at [the migration note](https://wavehouse.dev/access-control#migrating-from-the-operation-first-layout), instead of the confusing strict-decode "unknown field" error it would otherwise produce (or, for an empty operation block, silently decoding as a role named `select` with no grants — which fails the undeclared-role check when `roles.json` does not declare a role named `select` — the usual case, since `checkRoleRefs` errors and moves on before reaching the "grant sets neither select nor insert" warning. If such a role *is* declared, you get that warning instead and the document adopts). One shape is refused differently: a grant keyed by a role named after the *other* operation (`tables.t.select.insert`) reads as a different grant under each layout — different role, different operation, or both — so it gets its own error asking you to rename the role rather than the migration pointer. A role named after its *own* operation (`tables.t.select.select`) means the same thing either way and is accepted. @@ -80,7 +80,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed -- **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/query/ident.go` (removed, + tests), `internal/app/wire.go`, `docs/src/content/docs/{architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` and pipe execution rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL. The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. The cache now escapes table and scope names itself: a `Namespace` carries them raw and the version index builds its keys with `internal/keyenc`, so neither the structured-query read nor the ingest worker's invalidation escapes them (`query.SafeEncodeToken` is gone) and no name reaches a key unescaped. The suite pins that a name holding a dot, a space or a `%` is read and bumped under one key, and that names which would run together unescaped (`a.0.b` against `a` with scope `b.0.`) stay two entries. The keys themselves are unchanged apart from the query key's sha, now escaped too; they live in the process, so nothing stored is orphaned. +- **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/query/ident.go` (removed, + tests), `internal/app/wire.go`, `docs/src/content/docs/{architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` and pipe execution rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL. The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. The cache now escapes table and scope names itself: a `Namespace` carries them raw and the version index builds its keys with `internal/keyenc`, so neither the structured-query read nor the ingest worker's invalidation escapes them (`query.SafeEncodeToken` is gone) and no name reaches a key unescaped. The suite pins that a name holding a dot, a space or a `%` is read and bumped under one key, and that names which would run together unescaped (`a.0.b` against `a` with scope `b.0.`) stay two entries. The keys themselves are unchanged apart from the caller's query key at the head of an entry's key, now escaped whole too; they live in the process, so nothing stored is orphaned. - **An explicit `false`, `0` or `""` in `config.yaml` is no longer replaced by the key's default** (`internal/config/config.go`, `internal/config/defaults_test.go` (new), `docs/src/content/docs/configuration.mdx`, `AGENTS.md`): [#631](https://github.com/Wave-RF/WaveHouse/issues/631). Defaults lived in cleanenv `env-default` tags, which cleanenv applies after the YAML decode to any field still at its zero value, so it could not tell a key the file set to its zero value from one the file left out. `otel.traces.enabled: false`, `otel.metrics.enabled: false` and `otel.logs.enabled: false` came back `true`; `otel.traces.sample_rate: 0` and `otel.logs.sample_rate: 0` came back `1.0`; `server.shutdown_timeout: 0` came back `10`; `cache.l1_max_cost: 0`, `prometheus.path: ""` and `data_dir: ""` came back as their defaults; `server.port: 0` came back `8080`. All of it was silent. Defaults now live in one Go function, `defaults()`, which `Load` starts from before decoding the file and then applying `WH_*` variables, so the order is env > YAML > default and a key the file sets always wins. **Behaviour change if your file relied on the bug:** a zero you wrote now takes effect. A file that says `sample_rate: 0` now exports no traces (or no DEBUG/INFO logs), where it silently exported everything; a signal set `enabled: false` is now off; `shutdown_timeout: 0` now skips the drain. `cache.l1_max_cost: 0`, `server.port: 0`, and `data_dir: ""` now refuse boot (`cache init: MaxCost can't be zero`, `server.port 0 out of range`, `data_dir (WH_DATA_DIR) is required`) instead of running on the default; an empty `prometheus.path` refuses boot when `prometheus.enabled` is true. Delete the key to get the default back. Env vars are unchanged: they already honoured an explicit zero. New tests load through `config.Load` for every affected key (a YAML zero is kept, an absent key gets the default, env wins in both directions), refuse an `env-default` tag on any field, and pin each documented default in `configuration.mdx` to `defaults()`. - **An embedded queue store that cannot be created fails boot at once, naming the cause** (`internal/mq/embedded.go` (+ tests)): part of [#617](https://github.com/Wave-RF/WaveHouse/issues/617). A regular file at `/nats`, or a `nats` directory that could not be created there, failed JetStream in the background, so boot waited out the server's 5s readiness check and reported only `nats server not ready`. `NewEmbedded` now creates the directory first (at `0700`, as the server does) and refuses boot with the mkdir error. An existing but unwritable `nats` directory still takes the old path. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 971b94974..84fb608f9 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -61,7 +61,7 @@ internal/ ├── dedupe/ Optional deduplication (Pebble) ├── discovery/ ClickHouse schema introspection and validation ├── ingest/ Batch buffering, DLQ, and Active Sweeper -├── keyenc/ The one escaping composite keys are built from (NATS subject tokens, cache version-index keys) +├── keyenc/ The one escaping composite keys are built from (NATS subject tokens, cache keys) ├── mq/ MQ boundary: the only NATS/JetStream importer (owned message/consumer/stream types + embedded server) ├── observability/ OpenTelemetry pipeline (traces/metrics/logs + Prometheus exposition) ├── pipes/ Named query pipes (NamedQuery type, parameter binding, Source) @@ -110,9 +110,9 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ ### `cache/` — Query Cache -- **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on, each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope by their raw names, which the cache escapes where it builds a key: one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)). `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. +- **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on — one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)) — each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope by their raw names, which the cache escapes where it builds a key. `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. -- **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: a namespace key is `..
.
.`, its fields joined with `keyenc.Join` so a dot in a table or scope name is escaped rather than read as a separator, and a query key, `|.||…` with the sha escaped, folds in the tenant's version and each dependency's namespace key and namespace version, so bumping a table (a scopeless write) or one scope — scope is reserved and empty today, so every write is the whole-table bump — orphans every dependent entry without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results, and the flat directory's single tenant simply carries the `0` prefix. `BumpTenant` (behind `InvalidateTenant`) advances the tenant version that leads every namespace key of one tenant, orphaning its every namespace and every cached query in one step (the tenant version is folded into every query key, so a pipe result with no dependencies is orphaned too) — a table no bump ever keyed included, which an enumeration of the index would miss — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. +- **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: a namespace key is `..
.
.`, its fields joined with `keyenc.Join` so a dot in a table or scope name is escaped rather than read as a separator, and an entry's key, `|.||…` with the caller's query key escaped whole (its `:` become `%3A`), folds in the tenant's version and each dependency's namespace key and namespace version, so bumping a table (a scopeless write) or one scope — scope is reserved and empty today, so every write is the whole-table bump — orphans every dependent entry without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results, and the flat directory's single tenant simply carries the `0` prefix. `BumpTenant` (behind `InvalidateTenant`) advances the tenant version that leads every namespace key of one tenant, orphaning its every namespace and every cached query in one step (the tenant version is folded into every query key, so a pipe result with no dependencies is orphaned too) — a table no bump ever keyed included, which an enumeration of the index would miss — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. ### `config/` — Configuration @@ -204,7 +204,7 @@ The hot-reloadable half of configuration: a directory of four JSON files (`confi ### `keyenc/` — Key Escaping -- **keyenc.go** — The one escaping composite keys are built from, so a name can never be mistaken for a separator: `Escape` keeps ASCII letters, digits, `_` and `-` — exactly the tenant-id grammar, so a tenant id is its own escaped form — and writes every other byte as `%XX` (uppercase hex); `Unescape` is `url.PathUnescape`, which decodes `%XX` in either case and takes any other byte as itself, so a `%2D` for `-` that an earlier build wrote still reads. `Join`/`AppendJoin` escape each field and put a separator between them, panicking on no fields and on a separator the escaping could write or one outside ASCII, and `Split` reverses them. The package that builds a key takes raw names and escapes them itself, so no caller has to: NATS subject tokens (`internal/mq`) and the cache's version-index keys (`internal/cache`) both use it. Keys built from it are stored, so changing what it keeps orphans them. +- **keyenc.go** — The one escaping composite keys are built from, so a name can never be mistaken for a separator: `Escape` keeps ASCII letters, digits, `_` and `-` — exactly the tenant-id grammar, so a tenant id is its own escaped form — and writes every other byte as `%XX` (uppercase hex); `Unescape` is `url.PathUnescape`, which decodes `%XX` in either case and takes any other byte as itself, so a `%2D` for `-` that an earlier build wrote still reads. `Join`/`AppendJoin` escape each field and put a separator between them, panicking on no fields and on a separator the escaping could write or one outside ASCII, and `Split` reverses them. The package that builds a key takes raw names and escapes them itself, so no caller has to: NATS subject tokens (`internal/mq`) and the cache's keys (`internal/cache`) both use it. Keys built from it are stored, so changing what it keeps orphans them. ## Data Flows diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index a25165bba..291bb54c7 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -389,7 +389,7 @@ The folder name is the tenant id, and each folder is a complete settings directo **The admin routes take the operator key only.** `/v1/ops/*` reaches every tenant, so over a nested directory no tenant's admin role opens it: the [operator key](/api#authentication) alone does, and a token carrying an admin role gets `403`. Boot a nested directory without `auth.operator_key` and no caller can reach these routes at all, which leaves `SIGHUP` as the only reload; the server warns about it at boot. `GET /v1/ops/pipes`, `GET /v1/ops/pipes/{name}`, `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the same `?tenant=`, and address tenant `0` without it; `GET /v1/ops/dlq/stats` takes it too, and reads a rejected or removed tenant's dead-letter queue like a served one's, since the queue is kept; a tenant that has none is a `404`. On the routes that take it the parameter is parsed strictly — a query string that does not parse, an empty or repeated `tenant`, or a malformed id is a `400`, never a silent read of the default tenant or, on the reload route, a reload of every tenant. The SDK sends it as the [`tenant` option](/sdk/admin#settings--whsettings). -**What a tenant's folder decides.** A request is evaluated against its own tenant's `policies.json` and `pipes.json` (ingest, structured queries, pipes), its `query.*` keys, its `cors.allowed_origins`, and its `dedupe` block: whether its records are deduplicated, by which id, against the tenant's own store, which that folder's `dedupe.enabled` opens and closes on reload exactly as [the single-tenant one](/settings-directory#deduplication) does (every tenant's store is a share of the one Pebble instance at `/pebble`, each key led by its tenant), so `wavehouse_ingest_dedupe_disabled_total` ticks only across a tenant's own reload, whatever the other tenants' switches say. A tenant's seen ids are its own: the same event id is first seen under each tenant that sends it. Its `auth` block is its own too: each tenant's folder wires that tenant's token verifier (`jwks_url`, `role_claim`), built when the folder is adopted and rebuilt when its wiring changes, so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its provider's key set. Under another tenant's header a token is treated as invalid, and the request falls back to that tenant's `default_role` like any other unverifiable token, possibly after a rate-limited key refetch (see [Authentication](/settings-directory#authentication)). Keep `X-Tenant-ID` pinned at the proxy so a token is never presented under the wrong tenant. Tenants can still accept each other's tokens: those that leave `jwks_url` empty share the boot HMAC secret when `auth.jwt_secret` is set, so a token verifies under any of them (with no secret they validate no token at all), and those whose `jwks_url` names the same key set accept each other's tokens; isolate them by provider, or scope rows by a signed claim ([row-level security](/access-control#row-level-security)). A tenant whose `jwks_url` has not been fetched yet answers `503` with `Retry-After` to its token-bearing requests alone. A tenant that stops being served — its folder rejected or removed — loses its verifier and the JWKS refresh with it, and gets a fresh one when its folder is adopted again. The HMAC secret and the operator key stay boot config, shared by every tenant; the operator key is stamped with the request tenant's `admin_role`. A tenant's `clickhouse` and `schema` blocks are its own as well: each tenant reads and writes its own ClickHouse — one native pool per distinct address, database, user, password and `tls` tuple, shared by the tenants naming it, under the process-wide [connection ceiling](/settings-directory#clickhouse) — and discovers its own tables from its own database on its own `schema.refresh_interval`. Its message queue is its own as well: its events are queued on a stream of their own, capped at its own `mq.max_bytes_gb` — at that budget its ingest answers `503` while every other tenant's keeps publishing — beside a dead-letter stream of its own at a tenth of it, and the history that gap-fill replays from it is kept for its own `stream.gap_window_minutes`. Nothing checks what the tenants' budgets add up to against the disk, so size them together ([Message Queue](/settings-directory#message-queue)). An event is published on its tenant's subject (`ingest.{tenant}.{table}`), so a `GET /v1/stream` connection is authorized by its own tenant's `policies.json` and receives its own tenant's rows alone, the ingest worker inserts a row into its own tenant's ClickHouse, a failed row is parked under its own tenant's `dlq.enabled` and subject (`dlq.{tenant}.{table}`), and two tenants' tables of one name never share a batch. The query cache is one pool, but its entries are keyed by tenant: identical `POST /v1/query` and pipe requests from two tenants are two entries and two queries to ClickHouse, and a tenant is never served another's cached rows. An insert invalidates the table's cached results under every tenant on the same ClickHouse address and database as the tenant it was ingested for, whatever their user or `tls` block, since they read the same tables; a tenant on no pool — its folder rejected or removed, or no pool could be opened for it, such as by the ceiling — is out of that fan-out while it is, and has its cached `POST /v1/query` results dropped the moment it is back on one, so a repaired or restored folder never serves query rows cached before the inserts it missed, and so does a tenant whose folder moves it to another address or database, whose cached rows came from other tables. Both drop the tenant's cached pipe results too, but no insert invalidates one, since a pipe names no table: between those, a cached pipe result stays until its TTL expires. One setting weighs every tenant: the SSE keepalive, where the wheel runs at the shortest `stream.keepalive_interval` among the tenants being served, with that tenant's `stream.keepalive_buckets`. +**What a tenant's folder decides.** A request is evaluated against its own tenant's `policies.json` and `pipes.json` (ingest, structured queries, pipes), its `query.*` keys, its `cors.allowed_origins`, and its `dedupe` block: whether its records are deduplicated, by which id, against the tenant's own store, which that folder's `dedupe.enabled` opens and closes on reload exactly as [the single-tenant one](/settings-directory#deduplication) does (every tenant's store is a share of the one Pebble instance at `/pebble`, each key led by its tenant), so `wavehouse_ingest_dedupe_disabled_total` ticks only across a tenant's own reload, whatever the other tenants' switches say. A tenant's seen ids are its own: the same event id is first seen under each tenant that sends it. Its `auth` block is its own too: each tenant's folder wires that tenant's token verifier (`jwks_url`, `role_claim`), built when the folder is adopted and rebuilt when its wiring changes, so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its provider's key set. Under another tenant's header a token is treated as invalid, and the request falls back to that tenant's `default_role` like any other unverifiable token, possibly after a rate-limited key refetch (see [Authentication](/settings-directory#authentication)). Keep `X-Tenant-ID` pinned at the proxy so a token is never presented under the wrong tenant. Tenants can still accept each other's tokens: those that leave `jwks_url` empty share the boot HMAC secret when `auth.jwt_secret` is set, so a token verifies under any of them (with no secret they validate no token at all), and those whose `jwks_url` names the same key set accept each other's tokens; isolate them by provider, or scope rows by a signed claim ([row-level security](/access-control#row-level-security)). A tenant whose `jwks_url` has not been fetched yet answers `503` with `Retry-After` to its token-bearing requests alone. A tenant that stops being served — its folder rejected or removed — loses its verifier and the JWKS refresh with it, and gets a fresh one when its folder is adopted again. The HMAC secret and the operator key stay boot config, shared by every tenant; the operator key is stamped with the request tenant's `admin_role`. A tenant's `clickhouse` and `schema` blocks are its own as well: each tenant reads and writes its own ClickHouse — one native pool per distinct address, database, user, password and `tls` tuple, shared by the tenants naming it, under the process-wide [connection ceiling](/settings-directory#clickhouse) — and discovers its own tables from its own database on its own `schema.refresh_interval`. Its message queue is its own as well: its events are queued on a stream of their own, capped at its own `mq.max_bytes_gb` — at that budget its ingest answers `503` while every other tenant's keeps publishing — beside a dead-letter stream of its own at a tenth of it, and the history that gap-fill replays from it is kept for its own `stream.gap_window_minutes`. Nothing checks what the tenants' budgets add up to against the disk, so size them together ([Message Queue](/settings-directory#message-queue)). An event is published on its tenant's subject (`ingest.{tenant}.{table}`), so a `GET /v1/stream` connection is authorized by its own tenant's `policies.json` and receives its own tenant's rows alone, the ingest worker inserts a row into its own tenant's ClickHouse, a failed row is parked under its own tenant's `dlq.enabled` and subject (`dlq.{tenant}.{table}`), and two tenants' tables of one name never share a batch. The query cache is one pool, but its entries are keyed by tenant: identical `POST /v1/query` and pipe requests from two tenants are two entries and two queries to ClickHouse, and a tenant is never served another's cached rows. An insert invalidates the table's cached results under every tenant on the same ClickHouse address and database as the tenant it was ingested for, whatever their user or `tls` block, since they read the same tables; a tenant on no pool — its folder rejected or removed, or no pool could be opened for it, such as by the ceiling — is out of that fan-out while it is, and has its cached `POST /v1/query` results dropped the moment it is back on one, so a repaired or restored folder never serves query rows cached before the inserts it missed, and so does a tenant whose folder moves it to another address or database, whose cached rows came from other tables. Both also drop the tenant's cached pipe results; apart from them a cached pipe result stays until its TTL expires, since a pipe names no table and no insert invalidates it. One setting weighs every tenant: the SSE keepalive, where the wheel runs at the shortest `stream.keepalive_interval` among the tenants being served, with that tenant's `stream.keepalive_buckets`. **What a lost tenant `0` costs.** A `0` folder that a reload rejects or removes stops tenant `0` being served like any other, and what becomes of the shared settings depends on how they are read. Tenant `0` leaves its ClickHouse pool (closed only once no served tenant names its tuple), and its schema registry and verifier are released with the folder, like any other tenant's; the `/v1/ops/*` routes, which resolve no tenant, verify against it, so a token there reads as invalid (`401`) rather than merely non-admin (`403`) until tenant `0` is served again — the operator key, which never consults a verifier, is unaffected. CORS does not stay either: the responses that read tenant `0`'s list — the tenant-exempt routes, the refusals, a preflight naming no tenant — carry no CORS headers until the folder is served again, while every other tenant's routes keep their own list. Tenant `0`'s own dedupe store closes, as any rejected or removed tenant's does, its seen ids kept for the folder that restores it. What is read per event follows the event's tenant, so tenant `0`'s events are the ones affected: with no ClickHouse to insert into, its rows fail and are parked on the DLQ whatever its switch said, and its open `GET /v1/stream` connections are ended, as any tenant's are when it stops being served — the other tenants' events are untouched. A nested directory that has never served a tenant `0` — no `0` folder, or one rejected at boot — serves every other tenant from its own ClickHouse. Outside `/v1/ops/*`, a `/v1` request that sends no `X-Tenant-ID` resolves to tenant `0`, so with no `0` folder it answers `404 unknown tenant: 0` (`503` with a rejected one) — the SDK's `/v1/health` reachability ping included. diff --git a/docs/src/content/docs/development.md b/docs/src/content/docs/development.md index 16b65a74a..66efe2f90 100644 --- a/docs/src/content/docs/development.md +++ b/docs/src/content/docs/development.md @@ -461,7 +461,7 @@ WaveHouse/ │ ├── dedupe/ # Optional deduplication (Pebble) │ ├── discovery/ # ClickHouse schema introspection + validation │ ├── ingest/ # Batch buffering + DLQ + Active Sweeper -│ ├── keyenc/ # One escaping for composite keys (NATS subject tokens, cache namespace tokens) +│ ├── keyenc/ # One escaping for composite keys (NATS subject tokens, cache keys) │ ├── mq/ # MQ boundary: the only NATS/JetStream importer │ ├── observability/ # OpenTelemetry pipeline (traces/metrics/logs + Prometheus) │ ├── pipes/ # Named query pipes (types + parameter binding) diff --git a/internal/app/app_test.go b/internal/app/app_test.go index 98e50e1c0..bf3338b96 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -655,8 +655,8 @@ func TestSharedTables_InvalidatesTheTenantsSharingTheTables(t *testing.T) { // A tenant back on a pool after an absence — its folder rejected, then // repaired; removed, then restored — was out of the fan-out while away, so -// the wiring orphans its table-keyed cache as it comes back; a tenant that stayed -// is never touched, and a reload that changes nothing bumps nobody. +// the wiring orphans its cache as it comes back; a tenant that stayed is +// never touched, and a reload that changes nothing bumps nobody. func TestReload_ReadmittedTenantCacheIsOrphaned(t *testing.T) { root := writeNestedSettings(t, map[string]map[string]any{"acme": nil, "globex": nil}) a := newApp(t, testConfig(t, root), Options{}) diff --git a/internal/app/wire.go b/internal/app/wire.go index 140a14e78..9af2c4d3a 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -189,9 +189,10 @@ func dlqFor(tenants *settings.Registry) func(tenant.ID, string) bool { // rejected, removed, or one no pool could be opened for, such as by the // connection ceiling — is out of the fan-out, and its cache is orphaned // when it gets one (wireClickHouse, Cache.InvalidateTenant), so a folder -// repaired or restored inside a TTL never serves pre-insert rows; a pipe -// result names no table, so no insert invalidates it and between those it -// stays until its TTL expires (#343). +// repaired or restored inside a TTL never serves pre-insert rows. That also +// drops the tenant's cached pipe results; apart from it a pipe result stays +// until its TTL expires, since a pipe names no table and no insert +// invalidates it (#343). type sharedTables struct { cache.Cache sharing func(tenant.ID) []tenant.ID diff --git a/internal/keyenc/keyenc.go b/internal/keyenc/keyenc.go index 0e3434bc6..46e1f6f0a 100644 --- a/internal/keyenc/keyenc.go +++ b/internal/keyenc/keyenc.go @@ -1,6 +1,6 @@ // Package keyenc is the one escaping composite WaveHouse keys are built -// from: NATS subject tokens and cache namespace tokens. A field keeps ASCII -// letters, digits, '_' and '-' as they are and writes every other byte as %XX +// from: NATS subject tokens and cache keys. A field keeps ASCII letters, +// digits, '_' and '-' as they are and writes every other byte as %XX // (uppercase hex), so no separator, wildcard, whitespace, brace or non-ASCII // byte ever appears in it unescaped, and any table name ClickHouse accepts // encodes. The bytes it keeps are exactly a tenant id's (tenant.Parse), so a From 1e70d64506fd6f847776a55aee215b8b61974ad8 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 20:10:07 -0400 Subject: [PATCH 26/79] docs(cache): the Redis token-key layout is a protocol between builds Every process sharing a server reads and bumps the table and scope token keys for itself, so two builds that lay them out differently (a layout change, or a change to what keyenc keeps) split them: one build's bumps miss the entries the other filed, which are served until their TTL for the whole rolling deploy. The token-key comment says such a change needs a compatibility step or a documented flush; a valueKey change only orphans values. architecture.md and AGENTS.md name the shared backend's Redis keys among keyenc's users and state that rolling-upgrade consequence. development.md lists the shared backend in the cache tree line and says internal/cache's integration tests start their own Redis, Valkey, Dragonfly and Redis Cluster containers; the tech-stack row reads "Shared Cache". Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01B1tJWUp6oaDoH1usLwGtLF --- AGENTS.md | 2 +- docs/src/content/docs/architecture.md | 4 ++-- docs/src/content/docs/development.md | 4 ++-- internal/cache/redis_codec.go | 9 +++++++++ 4 files changed, 14 insertions(+), 5 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 692e1ba57..0f1dd85ad 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -38,7 +38,7 @@ Nineteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`dedupe/`** — `Deduplicator` interface → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) -- **`keyenc/`** — the one escaping composite keys are built from: `Escape` keeps `[A-Za-z0-9_-]` (exactly the tenant-id grammar, so a tenant id is its own escaped form) and writes every other byte as `%XX`, `Unescape` is `url.PathUnescape` (lenient: either hex case, and a byte left unescaped reads as itself, so a `%2D` an earlier build wrote still reads), `Join`/`AppendJoin` escape each field and put a separator between them (they panic on no fields, and on a separator the escaping could write or one outside ASCII) and `Split` reverses them. The package that builds a key takes raw names and escapes them itself, so no caller has to and no field reaches a key unescaped: NATS subjects (`Join`/`Split` after the verbatim tenant) and the cache's keys (`cache.VersionManager`) use it; changing what it keeps orphans every stored key +- **`keyenc/`** — the one escaping composite keys are built from: `Escape` keeps `[A-Za-z0-9_-]` (exactly the tenant-id grammar, so a tenant id is its own escaped form) and writes every other byte as `%XX`, `Unescape` is `url.PathUnescape` (lenient: either hex case, and a byte left unescaped reads as itself, so a `%2D` an earlier build wrote still reads), `Join`/`AppendJoin` escape each field and put a separator between them (they panic on no fields, and on a separator the escaping could write or one outside ASCII) and `Split` reverses them. The package that builds a key takes raw names and escapes them itself, so no caller has to and no field reaches a key unescaped: NATS subjects (`Join`/`Split` after the verbatim tenant) and the cache's keys — the version index and the shared backend's Redis keys (`internal/cache`) — use it; changing what it keeps orphans every stored key, and on the shared backend, whose keys every process builds for itself, splits them between builds for the length of a rolling upgrade (a bump one build makes misses the entries the other filed, served until their TTL) - **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`, `deadletter.go`), whose subject tokens are escaped by the shared `internal/keyenc`; `internal/app` constructs it and hands everything else a `mq.Broker` - **`observability/`** — OpenTelemetry pipeline: `InitProvider` wires trace/metric/log providers via OTLP gRPC (each signal independently gated). A top-level `Prometheus` config block drives an optional `/metrics` scrape endpoint that runs independently of OTLP push — standalone (Alloy/Mimir scrape, no collector), alongside OTLP, or off. `NewLogger` produces a slog handler that fans out to stdout AND OTLP (stdout always 100%, OTLP sample-rate-aware). `TraceHandler` injects trace_id/span_id from active spans. `tracer.go` provides W3C trace context propagation over message headers (`InjectHeaders`/`ExtractHeaders` on a plain header map; `internal/mq` injects on every publish and extracts on the `Subscribe` path, so this package never sees a NATS type). - **`pipes/`** — Named query pipes: `NamedQuery` type + `BindParams` + `Source` (read per request; `settings.Store` in production, `Static(q...)` in tests) diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index c2a77b1c7..e18ace9c1 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -209,7 +209,7 @@ The hot-reloadable half of configuration: a directory of four JSON files (`confi ### `keyenc/` — Key Escaping -- **keyenc.go** — The one escaping composite keys are built from, so a name can never be mistaken for a separator: `Escape` keeps ASCII letters, digits, `_` and `-` — exactly the tenant-id grammar, so a tenant id is its own escaped form — and writes every other byte as `%XX` (uppercase hex); `Unescape` is `url.PathUnescape`, which decodes `%XX` in either case and takes any other byte as itself, so a `%2D` for `-` that an earlier build wrote still reads. `Join`/`AppendJoin` escape each field and put a separator between them, panicking on no fields and on a separator the escaping could write or one outside ASCII, and `Split` reverses them. The package that builds a key takes raw names and escapes them itself, so no caller has to: NATS subject tokens (`internal/mq`) and the cache's keys (`internal/cache`) both use it. Keys built from it are stored, so changing what it keeps orphans them. +- **keyenc.go** — The one escaping composite keys are built from, so a name can never be mistaken for a separator: `Escape` keeps ASCII letters, digits, `_` and `-` — exactly the tenant-id grammar, so a tenant id is its own escaped form — and writes every other byte as `%XX` (uppercase hex); `Unescape` is `url.PathUnescape`, which decodes `%XX` in either case and takes any other byte as itself, so a `%2D` for `-` that an earlier build wrote still reads. `Join`/`AppendJoin` escape each field and put a separator between them, panicking on no fields and on a separator the escaping could write or one outside ASCII, and `Split` reverses them. The package that builds a key takes raw names and escapes them itself, so no caller has to: NATS subject tokens (`internal/mq`) and the cache's keys — the version index and the shared backend's Redis keys (`internal/cache`) — use it. Keys built from it are stored, so changing what it keeps orphans them — and on the shared backend, whose keys every process builds for itself, it splits them for the length of a rolling upgrade: a bump one build makes does not reach the entries the other build filed, which are served until their TTL. ## Data Flows @@ -373,7 +373,7 @@ Client GET /v1/stream | Analytics DB | ClickHouse | Primary data store + schema source of truth | | Message Queue | NATS + JetStream | Durable event streaming | | L1 Cache | Ristretto v2 | In-process memory cache | -| Shared cache | [rueidis](https://github.com/redis/rueidis) | Redis-compatible client for the shared backend (not yet selectable) | +| Shared Cache | [rueidis](https://github.com/redis/rueidis) | Redis-compatible client for the shared backend (not yet selectable) | | Embedded KV | Pebble | Optional deduplication | | Config | cleanenv | YAML + env var config loading | | Release | GoReleaser | Cross-platform binary builds | diff --git a/docs/src/content/docs/development.md b/docs/src/content/docs/development.md index a62476ee7..272ab8cd1 100644 --- a/docs/src/content/docs/development.md +++ b/docs/src/content/docs/development.md @@ -345,7 +345,7 @@ Each test target writes `covdata` to `tmp/coverage//data/`, renders a tex | E2E tests (SDK) | `tests/e2e/sdk/*.test.ts` | Yes | `make test-e2e` | - **Unit tests** live beside the code they test (e.g., `internal/discovery/discovery_test.go`). They use mocks or embedded NATS (in-process, no Docker needed). -- **Integration tests** use the `//go:build integration` build tag. `TestMain` starts one ClickHouse testcontainer and boots the production wiring against it through `app.New` (embedded NATS, ingest worker, sweeper, hub, the API server on a random loopback port); tests reach it via `env(t)` and create their own tables. DLQ tests use `assert.Eventually` with a 30-second timeout for the 5-second ingest worker batch window. +- **Integration tests** use the `//go:build integration` build tag. In `tests/integration`, `TestMain` starts one ClickHouse testcontainer and boots the production wiring against it through `app.New` (embedded NATS, ingest worker, sweeper, hub, the API server on a random loopback port); tests reach it via `env(t)` and create their own tables. DLQ tests use `assert.Eventually` with a 30-second timeout for the 5-second ingest worker batch window. `internal/cache`'s integration tests start their own containers instead — Redis, Valkey, Dragonfly and a one-node Redis Cluster — for the shared backend. Shared test utilities live in `internal/testutil/`. The packages log through `slog.Default()`, so tests reach log output through `internal/testutil/logtest`: `logtest.Silence()` in a package's `TestMain` discards it, and `logtest.Capture(t, level)` routes it to a buffer for a test that asserts on log lines — such a test must not call `t.Parallel()`, because the default logger is process-wide. @@ -454,7 +454,7 @@ WaveHouse/ │ ├── api/ # HTTP handlers, router, middleware │ ├── app/ # Process wiring (build every component, run under one errgroup, release in reverse) │ ├── auth/ # JWT/JWKS authentication middleware -│ ├── cache/ # Query cache: Ristretto L1 + the tenant-led version index +│ ├── cache/ # Query cache: Ristretto L1 + the tenant-led version index; the Redis-compatible shared backend │ ├── chconn/ # ClickHouse pools, one per connection tuple (reconciled on settings reload) │ ├── chsql/ # Shared ClickHouse SQL helpers (quoting + bind-safety) │ ├── config/ # YAML + env var configuration diff --git a/internal/cache/redis_codec.go b/internal/cache/redis_codec.go index 1ce16f0cb..e9ec64963 100644 --- a/internal/cache/redis_codec.go +++ b/internal/cache/redis_codec.go @@ -53,6 +53,15 @@ func tenantTokenKey(prefix string, id tenant.ID) string { // tableTokenKey and scopeTokenKey take raw names and escape them after the // fixed prefix (keyenc), so no ':' in a name reads as the separator. +// +// Their layout is a protocol between builds: every process on the server +// reads and bumps these keys for itself, so two builds that lay them out +// differently — a change here, or to what keyenc keeps — split them, and one +// build's bumps miss the entries the other filed, which are served until +// their TTL for the whole rolling deploy. Such a change needs a +// compatibility step (bump both layouts through the transition) or a +// documented flush. The tenant token, placed verbatim, stays shared; a +// valueKey change only orphans values, which is safe to roll. func tableTokenKey(prefix string, id tenant.ID, table string) string { return string(keyenc.AppendJoin([]byte(prefix+":{"+string(id)+"}:B:"), ':', table)) } From dbd368185894a20a673d8ddccc50b0da5237788d Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 20:19:51 -0400 Subject: [PATCH 27/79] docs(cache): InvalidateTenant drops pipe results; a pipe passes no deps The #610 changelog entry still said InvalidateTenant leaves pipe results alone, which this PR makes false; its "structured-query results" now read "cached results". QueryKey's comment said a pipe passes several namespaces (none yet, #343), and AGENTS.md asked every cache.Cache implementation, not every backend, to run cachetest.Run. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01B1tJWUp6oaDoH1usLwGtLF --- AGENTS.md | 2 +- CHANGELOG.md | 4 ++-- internal/cache/version_manager.go | 4 ++-- 3 files changed, 5 insertions(+), 5 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 1c167da27..2451f802b 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -128,7 +128,7 @@ Tooling notes (the non-obvious bits `make help` won't tell you): - **Shared mocks in `internal/testutil/`**: Use `MockPublisher` (records `Publish` and `DeadLetter`), `MockCache`, `MockDeduplicator`, `MockSubscriber`, `MockMessage`, `MockPurger`, `MockDeadLetterStats` instead of creating ad-hoc mocks. See `testutil/mocks.go`. - **JWT helpers**: Use `testutil.MakeJWT(t, claims)` and `testutil.MakeExpiredJWT(t, claims)` for auth tests. See `testutil/jwt.go`. - **Schema helpers**: Use `testutil.NewTestSchemaRegistry(t, tables)` for schema-aware tests — it builds the registry through the real discovery path (`Refresh` against a mock ClickHouse connection), so timestamp specs are precomputed like production. -- **Cache backends**: every `cache.Cache` implementation runs `cachetest.Run` (`internal/testutil/cachetest`), the backend-agnostic conformance suite; a behavior the contract promises goes there, not in one backend's tests. +- **Cache backends**: every `cache.Cache` backend runs `cachetest.Run` (`internal/testutil/cachetest`), the backend-agnostic conformance suite; a behavior the contract promises goes there, not in one backend's tests. - **Policy helpers**: Use `policy.Static(p)` for a fixed `policy.Source` in tests. - **Pipes helpers**: Use `pipes.Static(queries...)` for a fixed `pipes.Source` in tests. - **Response assertions**: Use `testutil.AssertJSONResponse(t, rec, status, expected)` and `testutil.AssertJSONContains(t, rec, status, substring)`. diff --git a/CHANGELOG.md b/CHANGELOG.md index 2c8dce0b6..a6556d5a7 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -11,7 +11,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. -- **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. +- **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (since #614 this drops the tenant's cached pipe results too; no insert invalidates a pipe result, which names no table). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. - **ClickHouse TLS, HTTP-interface headers, pool sizes and a connection ceiling** (`internal/settings/{settings,validate,store}.go` + seed, `internal/chconn/chconn.go`, `internal/ingest/worker.go`, `internal/api/query.go`, `internal/config/config.go`, `internal/app/wire.go`, `deployments/compose/settings/config.json`, `deployments/compose/standalone.yaml`, `config.yaml`, `docs/src/content/docs/{settings-directory,configuration,reverse-proxy}.mdx`, `docs/src/content/docs/{architecture,deployment}.md`): the tenant-agnostic first slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). `config.json`'s `clickhouse` block gains a `tls` block (`enabled`, `ca_file`, `cert_file`, `key_file`, `insecure_skip_verify`, `server_name`), a `headers` map for the HTTP interface, and `max_open_conns` / `max_idle_conns`. **Every key is required, so an existing `config.json` must add them**; the seed values (`tls.enabled: false` with the other `tls` keys empty, `headers: {}`, `10` / `5`) change nothing. `tls.enabled` switches the native hop to TLS, `http_scheme` stays the HTTP hop's switch, and the material applies to whichever hop uses TLS: the driver gets the TLS config and the pool sizes, the ingest worker and the raw-SQL proxy get the TLS config and the headers, set ahead of their own so the credentials win (naming `X-ClickHouse-User`, `X-ClickHouse-Key` or `Authorization` is a validation error, and so are two spellings of one name). Validation checks shape only — the paths are not opened, so `wavehouse validate` runs anywhere — warns when `insecure_skip_verify` is on or when only one of the two hops is on TLS (each carries the credentials in the clear without it), and a certificate file that cannot be read or parsed refuses boot or leaves a reload's connection unchanged; the files are read when the connection is built, so a file replaced in place needs a restart. Boot config gains the optional `clickhouse.max_total_conns` (`WH_CH_MAX_TOTAL_CONNS`, `0` = no ceiling): a settings pool above it refuses boot, naming both numbers in the error, and a reload that raises the pool above it is refused and logged (the reload still reports adopted), leaving the connection as it was. @@ -83,7 +83,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/query/ident.go` (removed, + tests), `internal/app/wire.go`, `docs/src/content/docs/{architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` and pipe execution rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL. The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. The cache now escapes table and scope names itself: a `Namespace` carries them raw and the version index builds its keys with `internal/keyenc`, so neither the structured-query read nor the ingest worker's invalidation escapes them (`query.SafeEncodeToken` is gone) and no name reaches a key unescaped. The suite pins that a name holding a dot, a space or a `%` is read and bumped under one key, and that names which would run together unescaped (`a.0.b` against `a` with scope `b.0.`) stay two entries. The keys themselves are unchanged apart from the caller's query key at the head of an entry's key, now escaped whole too; they live in the process, so nothing stored is orphaned. - **An explicit `false`, `0` or `""` in `config.yaml` is no longer replaced by the key's default** (`internal/config/config.go`, `internal/config/defaults_test.go` (new), `docs/src/content/docs/configuration.mdx`, `AGENTS.md`): [#631](https://github.com/Wave-RF/WaveHouse/issues/631). Defaults lived in cleanenv `env-default` tags, which cleanenv applies after the YAML decode to any field still at its zero value, so it could not tell a key the file set to its zero value from one the file left out. `otel.traces.enabled: false`, `otel.metrics.enabled: false` and `otel.logs.enabled: false` came back `true`; `otel.traces.sample_rate: 0` and `otel.logs.sample_rate: 0` came back `1.0`; `server.shutdown_timeout: 0` came back `10`; `cache.l1_max_cost: 0`, `prometheus.path: ""` and `data_dir: ""` came back as their defaults; `server.port: 0` came back `8080`. All of it was silent. Defaults now live in one Go function, `defaults()`, which `Load` starts from before decoding the file and then applying `WH_*` variables, so the order is env > YAML > default and a key the file sets always wins. **Behaviour change if your file relied on the bug:** a zero you wrote now takes effect. A file that says `sample_rate: 0` now exports no traces (or no DEBUG/INFO logs), where it silently exported everything; a signal set `enabled: false` is now off; `shutdown_timeout: 0` now skips the drain. `cache.l1_max_cost: 0`, `server.port: 0`, and `data_dir: ""` now refuse boot (`cache init: MaxCost can't be zero`, `server.port 0 out of range`, `data_dir (WH_DATA_DIR) is required`) instead of running on the default; an empty `prometheus.path` refuses boot when `prometheus.enabled` is true. Delete the key to get the default back. Env vars are unchanged: they already honoured an explicit zero. New tests load through `config.Load` for every affected key (a YAML zero is kept, an absent key gets the default, env wins in both directions), refuse an `env-default` tag on any field, and pin each documented default in `configuration.mdx` to `defaults()`. - **An embedded queue store that cannot be created fails boot at once, naming the cause** (`internal/mq/embedded.go` (+ tests)): part of [#617](https://github.com/Wave-RF/WaveHouse/issues/617). A regular file at `/nats`, or a `nats` directory that could not be created there, failed JetStream in the background, so boot waited out the server's 5s readiness check and reported only `nats server not ready`. `NewEmbedded` now creates the directory first (at `0700`, as the server does) and refuses boot with the mkdir error. An existing but unwritable `nats` directory still takes the old path. -- **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. +- **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its cached results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. - **SSE gap-fill re-reads the policy per replayed row** (`internal/stream/hub.go`): `ReplayProjector` captured the policy once when the replay began, so a policy adopted mid-fill — a revoked grant, say — applied only after the fill ended. It now reads it per event, as `Broadcast` does on the live path. diff --git a/internal/cache/version_manager.go b/internal/cache/version_manager.go index 48cf17c84..f6a763315 100644 --- a/internal/cache/version_manager.go +++ b/internal/cache/version_manager.go @@ -76,8 +76,8 @@ func (vm *VersionManager) NamespaceKey(ns Namespace) string { // version and every dependency's namespace key AND its namespace version, so // a bump of the tenant or of any dependency misses the key — a result with no // deps (a pipe) is orphaned by BumpTenant too. A structured query passes one -// Namespace; a pipe passes several. Deps are sorted so their order never -// changes the key. The key nests two levels: the escaped sha and the +// Namespace; a pipe passes none yet (#343). Deps are sorted so their order +// never changes the key. The key nests two levels: the escaped sha and the // '.'-joined tenant and dependency segments, separated by '|', which no // escaped field or '.' join ever holds. func (vm *VersionManager) QueryKey(id tenant.ID, sha string, deps []Namespace) string { From e6a824a4d747bdea658d77c57ac0ae8f8457c922 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 20:23:09 -0400 Subject: [PATCH 28/79] docs(cache): a token-key layout change needs both layouts read and bumped The compatibility step the token-key comment named did not close the split: an old build cannot be changed, so it keeps bumping only old keys, and a new build that bumps both but files under the new layout alone still misses every write the old build handles. The comment now says the new build must read and bump both layouts, folding the old tokens into what it files, until no old build is left; or the upgrade must never run two builds against the server at once. The VersionTTL floor's error no longer claims every value under 2s rounds to EX 0: EX has one-second resolution, and only below about 1.1s can the jittered TTL truncate to EX 0. The 2s floor stays. CONTRIBUTING.md, development.md and AGENTS.md say integration tests may live beside a package tested against its own external server, and that the suite also starts Redis, Valkey, Dragonfly and a one-node Redis Cluster. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01B1tJWUp6oaDoH1usLwGtLF --- AGENTS.md | 4 ++-- CONTRIBUTING.md | 2 +- docs/src/content/docs/development.md | 2 +- internal/cache/redis.go | 2 +- internal/cache/redis_codec.go | 10 ++++++---- 5 files changed, 11 insertions(+), 9 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index a7f4bb306..9ba3cdd75 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -151,7 +151,7 @@ If `make ci` passes locally, your commit has crossed the same gates CI will run ### Running `make ci` (for agents) -`make ci` is **self-contained**: the integration suite (`tests/integration/`) and the E2E orchestrator (`scripts/orchestrator/`) each boot ClickHouse via **testcontainers on random host ports**. The only prerequisite is a running **Docker daemon** — do **not** `make deps-up` or start ClickHouse first (`deps-up` is for `make dev` only). +`make ci` is **self-contained**: the integration suite (`tests/integration/`) and the E2E orchestrator (`scripts/orchestrator/`) each boot ClickHouse via **testcontainers on random host ports**, and the shared cache backend's integration tests (`internal/cache/`) start their own Redis, Valkey, Dragonfly and one-node Redis Cluster containers the same way. The only prerequisite is a running **Docker daemon** — do **not** `make deps-up` or start ClickHouse first (`deps-up` is for `make dev` only). Run it via the **background Bash tool** (`run_in_background: true`) and wait for the completion notification; the harness re-invokes you on exit, so polling the log with `tail` only burns context: @@ -447,7 +447,7 @@ internal/stream/ → SSE fan-out (event Hub: project once per role, Subsc internal/tenant/ → Tenant id (type, grammar, reserved default, request header name) internal/testutil/ → Shared test helpers (mocks, JWT + schema helpers; logtest/ captures or silences the default logger; cachetest/ is the conformance suite every cache.Cache backend runs) tests/ → Integration & E2E tests -tests/integration/ → Go integration tests (//go:build integration; ClickHouse testcontainer) +tests/integration/ → Go integration tests (//go:build integration; ClickHouse testcontainer). A package tested against its own external server keeps them beside it: internal/cache/redis_integration_test.go (Redis, Valkey, Dragonfly, Redis Cluster testcontainers) tests/e2e/ → E2E test stack (scripts/orchestrator boots a ClickHouse testcontainer + the wavehouse-cov binary) tests/e2e/fixtures/ → Idempotent ClickHouse DDL scripts for test tables tests/e2e/sdk/ → E2E integration tests via TypeScript SDK (Vitest) diff --git a/CONTRIBUTING.md b/CONTRIBUTING.md index c02dfa1d9..70bbb00ba 100644 --- a/CONTRIBUTING.md +++ b/CONTRIBUTING.md @@ -39,7 +39,7 @@ Open a [feature request issue](https://github.com/Wave-RF/WaveHouse/issues/new?t The pre-push hook (installed by `make tools`) blocks a push until the tree has been validated locally: a code change needs `make ci`, a docs/prose-only change needs only `make verify` (the same split CI makes). `make lint` / `make test` / `make build` are fast inner-loop subsets. -2. Write tests for new functionality. Unit tests go alongside the code in `internal/`. Integration tests go in `tests/` with the `//go:build integration` tag. +2. Write tests for new functionality. Unit tests go alongside the code in `internal/`. Integration tests carry the `//go:build integration` tag and go in `tests/integration/`, or beside the package when they test one package against its own external server (e.g. `internal/cache/redis_integration_test.go`), with the package added to the `test-integration` target. 3. Update documentation if your change affects: - API endpoints → update `docs/src/content/docs/api.md` diff --git a/docs/src/content/docs/development.md b/docs/src/content/docs/development.md index 272ab8cd1..df84fb8ee 100644 --- a/docs/src/content/docs/development.md +++ b/docs/src/content/docs/development.md @@ -16,7 +16,7 @@ You need these on your `PATH` before any `make` recipe will work end-to-end: | **Go** | 1.26+ (matches `go.mod`) | Compiles `cmd/wavehouse`; also runs the pinned `tool` deps (`gotestsum`, `gofumpt`, `goimports`, `govulncheck`, `deadcode`, `gsa`, `goda`) via `go tool` | [go.dev/dl](https://go.dev/dl/) | | **GNU Make** | **4.0+** | The Makefile uses `--output-sync=target` (Make 4 only) and bash-pinned recipes. macOS ships with BSD Make 3.81, which **will not work** | macOS: `brew install make` then use `gmake` or put `$(brew --prefix make)/libexec/gnubin` on your PATH. Linux: usually already installed | | **bash** | 4+ recommended | Recipes are pinned to `bash`; the helper scripts under `scripts/` use `set -euo pipefail` and bash arrays | macOS default is bash 3.2 (works for current recipes, but `brew install bash` is safer); Linux distros ship 4+ | -| **Docker** *(or Podman)* | Engine 20.10+ with the Compose **v2** plugin (`docker compose`, no hyphen) | Compose stacks under `deployments/compose/`; the E2E and integration suites boot ClickHouse via testcontainers (no compose file) | [Docker Desktop](https://docs.docker.com/get-docker/), [colima](https://github.com/abiosoft/colima), or [Podman](https://podman.io) with `podman-compose` / the `podman compose` plugin. The testcontainers Go library also honors `DOCKER_HOST` for rootless Podman setups | +| **Docker** *(or Podman)* | Engine 20.10+ with the Compose **v2** plugin (`docker compose`, no hyphen) | Compose stacks under `deployments/compose/`; the E2E and integration suites boot ClickHouse via testcontainers (no compose file), and the integration suite also starts Redis, Valkey, Dragonfly (pulled from `docker.dragonflydb.io`) and a one-node Redis Cluster for the shared cache backend | [Docker Desktop](https://docs.docker.com/get-docker/), [colima](https://github.com/abiosoft/colima), or [Podman](https://podman.io) with `podman-compose` / the `podman compose` plugin. The testcontainers Go library also honors `DOCKER_HOST` for rootless Podman setups | | **Node.js** | 22 LTS — pinned via `.nvmrc` at the repo root | Runtime for pnpm and the Vitest suites. Pinned to match CI (`setup-node` uses 22) and to avoid Node-major surprises; older Vitest versions in this repo were known to crash on Node 26 with a V8 heap-allocation abort | [nodejs.org](https://nodejs.org/) or `nvm use` / `fnm use` / `volta` (all read `.nvmrc`) | | **pnpm** | 11.21+ (pinned via `packageManager` in the root `package.json`) | Package manager for the TypeScript SDK, E2E test harness, and docs site (managed as a single pnpm workspace from the repo root); `make build-ts`, `make test-ts`, `make test-e2e`, `make build-docs`, `make dev-docs`, `make preview-docs` all shell out to `pnpm` | `corepack enable && corepack prepare pnpm@11.21.0 --activate` (recommended), or `npm i -g pnpm` | | **git** + **curl** | any recent | `git` for source + version metadata in builds; `curl` is used by the Makefile to fetch the pinned `golangci-lint` binary into `.bin/` | usually preinstalled | diff --git a/internal/cache/redis.go b/internal/cache/redis.go index f9ce1a0fe..01c6a6c41 100644 --- a/internal/cache/redis.go +++ b/internal/cache/redis.go @@ -123,7 +123,7 @@ func (c RedisConfig) withDefaults() (RedisConfig, error) { } } if c.VersionTTL != 0 && c.VersionTTL < 2*time.Second { - return c, fmt.Errorf("cache: redis version ttl %s is under 2s: jittered, it would round to EX 0", c.VersionTTL) + return c, fmt.Errorf("cache: redis version ttl %s is under 2s: EX has one-second resolution, and jittered below ~1.1s it could be EX 0", c.VersionTTL) } c.KeyPrefix = cmpOr(c.KeyPrefix, DefaultRedisKeyPrefix) c.Timeout = cmpOr(c.Timeout, DefaultRedisTimeout) diff --git a/internal/cache/redis_codec.go b/internal/cache/redis_codec.go index e9ec64963..96cddb452 100644 --- a/internal/cache/redis_codec.go +++ b/internal/cache/redis_codec.go @@ -58,10 +58,12 @@ func tenantTokenKey(prefix string, id tenant.ID) string { // reads and bumps these keys for itself, so two builds that lay them out // differently — a change here, or to what keyenc keeps — split them, and one // build's bumps miss the entries the other filed, which are served until -// their TTL for the whole rolling deploy. Such a change needs a -// compatibility step (bump both layouts through the transition) or a -// documented flush. The tenant token, placed verbatim, stays shared; a -// valueKey change only orphans values, which is safe to roll. +// their TTL for the whole rolling deploy. Such a change needs the new build +// to read and bump both layouts (fold the old tokens into what it files) +// until no old build is left, a later build dropping the old; or an upgrade +// that never runs two builds against the server at once. The tenant token, +// placed verbatim, stays shared; a valueKey change only orphans values, +// which is safe to roll. func tableTokenKey(prefix string, id tenant.ID, table string) string { return string(keyenc.AppendJoin([]byte(prefix+":{"+string(id)+"}:B:"), ':', table)) } From 674e0e20df452398631611a677e31520fffa5e62 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 20:26:50 -0400 Subject: [PATCH 29/79] docs(changelog): #382 never reached pipes; entry keys gained the tenant version A pipe's key folded no version before this change, so the mid-query re-filing could not happen to it; and an entry's key now carries the tenant's version as well as the escaped query key. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01B1tJWUp6oaDoH1usLwGtLF --- CHANGELOG.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index a6556d5a7..dd546ee2f 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -80,7 +80,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed -- **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/query/ident.go` (removed, + tests), `internal/app/wire.go`, `docs/src/content/docs/{architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` and pipe execution rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL. The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. The cache now escapes table and scope names itself: a `Namespace` carries them raw and the version index builds its keys with `internal/keyenc`, so neither the structured-query read nor the ingest worker's invalidation escapes them (`query.SafeEncodeToken` is gone) and no name reaches a key unescaped. The suite pins that a name holding a dot, a space or a `%` is read and bumped under one key, and that names which would run together unescaped (`a.0.b` against `a` with scope `b.0.`) stay two entries. The keys themselves are unchanged apart from the caller's query key at the head of an entry's key, now escaped whole too; they live in the process, so nothing stored is orphaned. +- **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/query/ident.go` (removed, + tests), `internal/app/wire.go`, `docs/src/content/docs/{architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL (a pipe result's key folded no version, so pipes were unaffected; with the tenant's version in every key they now take the same snapshot). The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. The cache now escapes table and scope names itself: a `Namespace` carries them raw and the version index builds its keys with `internal/keyenc`, so neither the structured-query read nor the ingest worker's invalidation escapes them (`query.SafeEncodeToken` is gone) and no name reaches a key unescaped. The suite pins that a name holding a dot, a space or a `%` is read and bumped under one key, and that names which would run together unescaped (`a.0.b` against `a` with scope `b.0.`) stay two entries. Namespace keys are unchanged; an entry's key now carries the tenant's version and escapes the caller's query key whole. All of them live in the process, so nothing stored is orphaned. - **An explicit `false`, `0` or `""` in `config.yaml` is no longer replaced by the key's default** (`internal/config/config.go`, `internal/config/defaults_test.go` (new), `docs/src/content/docs/configuration.mdx`, `AGENTS.md`): [#631](https://github.com/Wave-RF/WaveHouse/issues/631). Defaults lived in cleanenv `env-default` tags, which cleanenv applies after the YAML decode to any field still at its zero value, so it could not tell a key the file set to its zero value from one the file left out. `otel.traces.enabled: false`, `otel.metrics.enabled: false` and `otel.logs.enabled: false` came back `true`; `otel.traces.sample_rate: 0` and `otel.logs.sample_rate: 0` came back `1.0`; `server.shutdown_timeout: 0` came back `10`; `cache.l1_max_cost: 0`, `prometheus.path: ""` and `data_dir: ""` came back as their defaults; `server.port: 0` came back `8080`. All of it was silent. Defaults now live in one Go function, `defaults()`, which `Load` starts from before decoding the file and then applying `WH_*` variables, so the order is env > YAML > default and a key the file sets always wins. **Behaviour change if your file relied on the bug:** a zero you wrote now takes effect. A file that says `sample_rate: 0` now exports no traces (or no DEBUG/INFO logs), where it silently exported everything; a signal set `enabled: false` is now off; `shutdown_timeout: 0` now skips the drain. `cache.l1_max_cost: 0`, `server.port: 0`, and `data_dir: ""` now refuse boot (`cache init: MaxCost can't be zero`, `server.port 0 out of range`, `data_dir (WH_DATA_DIR) is required`) instead of running on the default; an empty `prometheus.path` refuses boot when `prometheus.enabled` is true. Delete the key to get the default back. Env vars are unchanged: they already honoured an explicit zero. New tests load through `config.Load` for every affected key (a YAML zero is kept, an absent key gets the default, env wins in both directions), refuse an `env-default` tag on any field, and pin each documented default in `configuration.mdx` to `defaults()`. - **An embedded queue store that cannot be created fails boot at once, naming the cause** (`internal/mq/embedded.go` (+ tests)): part of [#617](https://github.com/Wave-RF/WaveHouse/issues/617). A regular file at `/nats`, or a `nats` directory that could not be created there, failed JetStream in the background, so boot waited out the server's 5s readiness check and reported only `nats server not ready`. `NewEmbedded` now creates the directory first (at `0700`, as the server does) and refuses boot with the mkdir error. An existing but unwritable `nats` directory still takes the old path. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its cached results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. From ee51e8401cac9e4e8669a6924dcd6b1931e5445a Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 04:50:24 -0400 Subject: [PATCH 30/79] test(app): assert the wired cache is a pruner before swapping it MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit TestReload_PrunesCacheIndexToServedTenants replaced a.cache with a recorder by hand, so the run-time a.cache.(pruner) check against the cache wireCache actually builds was never exercised: a change that stopped wiring a pruning cache left the test green while pruning silently went dead. Assert the wired cache satisfies pruner before the swap. Also reword a stale assertion message ("a rejection releases; it orphans nothing yet") that no longer describes real wiring now that a rejection drops the tenant's cache index via Prune, and correct the CHANGELOG's claim that the Redis-compatible backend (#613) already bounds its versions with a TTL — that backend does not exist yet. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- CHANGELOG.md | 2 +- internal/app/app_test.go | 5 ++++- 2 files changed, 5 insertions(+), 2 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 2515caa5a..674c2df03 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -79,7 +79,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed - **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/app/wire.go`, `docs/src/content/docs/{architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` and pipe execution rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL. The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. -- **The cache's version index no longer grows with every bump, and forgets a tenant no longer served** (`internal/cache/{local,version_manager}.go` (+ tests), `internal/app/wire.go` (+ tests), `docs/src/content/docs/architecture.md`, `AGENTS.md`): fixes [#262](https://github.com/Wave-RF/WaveHouse/issues/262) for the in-process cache, part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The index nested each table under its tenant's version and each scope under its table's, and never pruned, so every tenant invalidation left the tenant's whole index behind, and it grew with every tenant ever served. It now holds one version per tenant, per (tenant, table) and per (tenant, table, scope), bumped in place. A tenant invalidation drops the tenant's index and hands its next key a generation unique within the process, so nothing cached before it can match again, and a table bump drops the table's scope versions. After each settings reload the index of every tenant no longer served, removed or rejected, is dropped the same way; its cached results are orphaned with it, as they already were when such a tenant came back on a pool. No change to what is cached or served. The Redis-compatible backend (#613) bounds its versions with a TTL instead. +- **The cache's version index no longer grows with every bump, and forgets a tenant no longer served** (`internal/cache/{local,version_manager}.go` (+ tests), `internal/app/wire.go` (+ tests), `docs/src/content/docs/architecture.md`, `AGENTS.md`): fixes [#262](https://github.com/Wave-RF/WaveHouse/issues/262) for the in-process cache, part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The index nested each table under its tenant's version and each scope under its table's, and never pruned, so every tenant invalidation left the tenant's whole index behind, and it grew with every tenant ever served. It now holds one version per tenant, per (tenant, table) and per (tenant, table, scope), bumped in place. A tenant invalidation drops the tenant's index and hands its next key a generation unique within the process, so nothing cached before it can match again, and a table bump drops the table's scope versions. After each settings reload the index of every tenant no longer served, removed or rejected, is dropped the same way; its cached results are orphaned with it, as they already were when such a tenant came back on a pool. No change to what is cached or served. The Redis-compatible backend (#613) will bound its versions with a TTL instead. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/internal/app/app_test.go b/internal/app/app_test.go index 84c670703..d21f376a0 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -671,7 +671,7 @@ func TestReload_ReadmittedTenantCacheIsOrphaned(t *testing.T) { rewriteSettings(t, filepath.Join(root, "globex"), invalidQuery) a.tenants.Reload("test") - assert.Empty(t, mock.GetTenants(), "a rejection releases; it orphans nothing yet") + assert.Empty(t, mock.GetTenants(), "a rejection calls no InvalidateTenant (Prune drops its index)") rewriteSettings(t, filepath.Join(root, "globex"), nil) _, adopted = a.tenants.Reload("test") require.True(t, adopted) @@ -712,6 +712,9 @@ func (p *pruneRecorder) last() map[tenant.ID]bool { func TestReload_PrunesCacheIndexToServedTenants(t *testing.T) { root := writeNestedSettings(t, map[string]map[string]any{"acme": nil, "globex": nil}) a := newApp(t, testConfig(t, root), Options{}) + _, ok := a.cache.(pruner) + require.True(t, ok, "the wired cache prunes") + rec := &pruneRecorder{} a.cache = rec From 55eb1035bf9e761de7c6103e3230332ae6a934eb Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 04:54:54 -0400 Subject: [PATCH 31/79] fix(cache): take the snapshot before choosing the tenant's pool Both cached handlers resolved the tenant's connection before their cache Lookup. A reload that repoints the tenant (Pools.Reconcile, then InvalidateTenant) between the two left a request on the old pool with a snapshot of the new tenant version, so its fill filed the old database's rows as fresh. The handlers now look up first and resolve the pool after, still answering 503 for a tenant on no pool before anything is served, a hit included. TestCachedRoutes_ReloadAsThePoolIsTakenOrphansTheFill pins both halves. Two conformance cases could not fail: no Lookup reads the key a zero or foreign-dependency snapshot would land under. The foreign-dependency case now asserts the refused Lookup returns the zero Snapshot, and the zero snapshot case counts entries through a new optional Options.Entries hook, which LocalCache supplies. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- AGENTS.md | 2 +- CHANGELOG.md | 2 +- docs/src/content/docs/api.md | 4 +- docs/src/content/docs/architecture.md | 4 +- internal/api/cache_tenant_test.go | 54 +++++++++++++++++++++++- internal/api/pipes.go | 39 +++++++++-------- internal/api/structured_query.go | 39 +++++++++-------- internal/api/tenant_clickhouse_test.go | 5 ++- internal/cache/cache.go | 14 +++--- internal/cache/export_test.go | 12 ++++++ internal/cache/local_test.go | 5 ++- internal/ingest/worker_test.go | 12 +++--- internal/testutil/cachetest/cachetest.go | 28 ++++++++---- 13 files changed, 152 insertions(+), 68 deletions(-) create mode 100644 internal/cache/export_test.go diff --git a/AGENTS.md b/AGENTS.md index dcce37b92..60474bc35 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -31,7 +31,7 @@ Twenty internal packages under `internal/` (plus `internal/testutil/` for shared - **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers; `ch_errors.go` (`writeCHError`) is the one mapping from a failed ClickHouse query to status, `code` and `retryable` - **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` handing the sweeper each tenant's own gap window (a rejected tenant's as its folder last had it, unbounded for one rejected since boot) and the `mq.max_bytes_gb` reconcile each served tenant's byte budget, and `defaultPolicy` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, and the same hook's `Hub.Prune` ends the open streams of a tenant no longer served), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `New` wires only what the process's `roles` need (discovery, dedupe, auth verifiers, the hub bridge and keepalive per API process; the ingest worker per ingest process; the sweeper under its lease through `elected`); a process without `api` serves `api.NewOpsRouter` — probes, `/version`, metrics, and the settings reload behind the operator key alone. `cmd/wavehouse` and `tests/integration` both boot through it - **`auth/`** — JWT auth middleware: HMAC **or** JWKS verification with `alg` pinned to the active verifier, role extraction from a configurable claim path; always runs, never rejects (bad token → empty role + stashed reason). One verifier per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 9): `Authenticator` keys them by `tenant.ID` — the request store's `settings.Store.Tenant()`, through an injected `TenantSource`; `tenant.Default` on the tenant-exempt routes — built from each tenant's `auth` block by `Reconfigure`, dropped by `Prune` once the tenant stops being served (rejected or removed), released by `Close`; the secrets (`Config`) are boot-level and shared. A JWKS key set is fetched off the boot and reload paths: until one has been stored the verifier is pending and a token-bearing request gets `503` + `Retry-After` from `api.refuseUnverifiable` (`auth.ErrVerifierPending`), never a `default_role` evaluation; refresh is library-managed (Eric, 2026-09-22), response capped at 1 MiB; the operator key's admin role is the request tenant's -- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for the caller's query key and its singleflight (escaped whole as the lead field of the stored key, `|.||…`), `..
.
.` for a namespace, each field escaped by `keyenc` where the key is built (a `Namespace` carries the raw table and scope, so no caller escapes) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version folded into every namespace and query key of one tenant, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read and `Set` files the fill under it, so a write landing mid-query orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) +- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for the caller's query key and its singleflight (escaped whole as the lead field of the stored key, `|.||…`), `..
.
.` for a namespace, each field escaped by `keyenc` where the key is built (a `Namespace` carries the raw table and scope, so no caller escapes) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version folded into every namespace and query key of one tenant, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read, taken before the handler chooses any input a bump invalidates — the tenant's connection included — and `Set` files the fill under it, so a write landing mid-query, or a reload repointing the tenant after the request took its connection, orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config. `Classify` (`errclass.go`) says what a failed ClickHouse request means for the request — `Unavailable`, `Denied`, `Rejected` (any unlisted exception code: the server read it and refused it), or `Unknown` (no code, no recognizable transport failure) — over the driver's error types and the HTTP interface's `HTTPError`; the ingest worker and the query handlers (`api/ch_errors.go` `writeCHError`, [#403](https://github.com/Wave-RF/WaveHouse/issues/403), [#271](https://github.com/Wave-RF/WaveHouse/issues/271)) both use it - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) - **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability when a selected backend keeps state there (`NeedsDataDir`); `backends.go` holds each layer's `.backend` (only the in-process value today); `config.go` holds `roles` (`Has(Role)`) and `instance_id`, and `Validate` refuses a role split the backends cannot serve (any split over the embedded MQ; `api` without `ingest`, or the reverse, over a local cache) — boot is the validator, there is no dry run diff --git a/CHANGELOG.md b/CHANGELOG.md index cbd483ec1..e8ccf372c 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -88,7 +88,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **A failed ClickHouse query answers by what went wrong, not a flat `500`/`502`** (`internal/api/ch_errors.go` (new, + tests), `internal/api/{errors,query,structured_query,pipes,schema,ch_settings}.go`, `internal/chconn/errclass.go` (`HTTPStatus` exported), `clients/ts/src/errors.ts` (+ tests), `tests/integration/query_errors_test.go` (new), `tests/integration/query_limits_test.go`, `internal/app/app_test.go`, `tests/e2e/sdk/{admin,query}.test.ts`, `AGENTS.md`, `docs/src/content/docs/{api,architecture}.md`, `docs/src/content/docs/{access-control,configuration}.mdx`, `docs/src/content/docs/sdk/{reference.md,index.mdx}`): fixes [#403](https://github.com/Wave-RF/WaveHouse/issues/403) and [#271](https://github.com/Wave-RF/WaveHouse/issues/271), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). ClickHouse answers a syntax error, a missing grant and an overloaded server alike with HTTP `500`, so `/v1/ops/query` turned a bad statement into a `502` and `/v1/query` and pipes into a `500` the SDK retried. All three now class the failure with `chconn.Classify` through one helper, `writeCHError`: a statement ClickHouse refused is `400 clickhouse.rejected`; a query over a rows/bytes limit, the role's own memory cap, or its time cap where that is no longer than `query_timeout` is `400 clickhouse.limit_exceeded`; `ACCESS_DENIED` is `403 clickhouse.access_denied`; credentials, user or database refused, or a redirect or `4xx` with no exception code from whatever fronts ClickHouse, is `502 clickhouse.misconfigured`; ClickHouse down, unreachable or overloaded is `503 clickhouse.unavailable` with `Retry-After: 5`; a failure with no verdict stays `500` (`502` on the proxy) as `clickhouse.unknown`. The error envelope gains `code` and `retryable` next to `error` on these responses — additive. A role with `max_execution_time` now queries with no context deadline and a cancel two seconds past the cap instead: clickhouse-go overwrote the cap's `max_execution_time` with deadline+5s for any deadline over 1s, so an overrun came back as a bare deadline, indistinguishable from waiting for a pooled connection; ClickHouse now enforces the cap itself and reports `TIMEOUT_EXCEEDED`. `POST /v1/ops/schema/refresh` against an unreachable ClickHouse is a `503` with `Retry-After` instead of a `500`. **SDK:** `WaveHouseError.code` and `retryable` now take the server's `code`/`retryable` when the body has them (`HTTP_` and "5xx retries" otherwise), so a rejected query is `clickhouse.rejected` rather than `HTTP_500`, and is not retried. - **An unavailable ClickHouse is retried with backoff instead of dead-lettering every row** (`internal/chconn/errclass.go` (new, + tests), `internal/ingest/{worker,backoff}.go` (`backoff.go` new, + tests), `internal/mq/{mq,embedded}.go`, `internal/testutil/mocks.go`, `tests/integration/ingest_outage_test.go` (new), `AGENTS.md`, `README.md`, `docs/src/content/docs/{ingest-pipeline,architecture,api,deployment,why-wavehouse}.md`, `docs/src/content/docs/{settings-directory,index,access-control}.mdx`): workstream A of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). A failed batch insert used to go through row-by-row isolation whatever the failure, so a ClickHouse that was down, overloaded or read-only failed every row twice and parked the whole batch on the DLQ. `chconn.Classify` now classes the failure first — `Rejected` (any ClickHouse exception code outside the availability and credential lists: the server read the row and refused it), `Unavailable` (connection refused/reset, timeouts, `TOO_MANY_SIMULTANEOUS_QUERIES`, `SERVER_OVERLOADED`, `MEMORY_LIMIT_EXCEEDED`, `TOO_MANY_PARTS`, `READONLY`, `TABLE_IS_READ_ONLY`, `KEEPER_EXCEPTION`, …), `Denied` (`AUTHENTICATION_FAILED`, `ACCESS_DENIED`, …) or `Unknown` (no code, no recognizable transport failure). Only `Rejected` is isolated and dead-lettered as before, and a multi-row batch refused with `TOO_MANY_PARTS` or `MEMORY_LIMIT_EXCEEDED` is split row by row first (`chconn.Splittable`), because a batch spanning too many partitions or too much memory can fail where each of its rows inserts; every other class hands the batch back to the queue with a delayed nak (`mq.Message.NakWithDelay`, new) under a jittered 1 s → 30 s backoff shared by every table on the same ClickHouse pool (a failure of one table — read-only, too many parts or mutations, a grant missing on it, `chconn.TableScoped` — backs off that table alone), which turns rows away without a request while it runs and probes once per window, and ClickHouse going away mid-isolation stops isolation and retries the rows it had not settled. Counted by the new `wavehouse_ingest_retries_total{table, reason}`; logged at `WARN` when an outage starts and at most every 30 s during it. A long outage now shows as a growing ingest stream and, at `mq.max_bytes_gb`, ingest `503`s — not as a full DLQ; a lasting failure of one table holds back its tenant's other tables once its waiting rows reach `maxAckPending`. Retried rows come back out of arrival order, which matters only to a `ReplacingMergeTree` without a version column or a `CollapsingMergeTree`. - **Schema discovery's retry loop jitters its backoff** (`internal/discovery/discovery.go` (+ tests), `internal/app/wire.go`, `internal/api/errors.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`): `RetryRefresh` slept exactly `2s * 2^n` capped at 60s, so instances retrying against one recovering ClickHouse fired in lockstep, every 60s on the same second. Each sleep is now drawn uniformly from below the backoff (full jitter), spreading the retries over the whole window and halving the mean wait — so a failing tenant's retries, their log lines and `wavehouse_schema_refresh_failures_total` come about twice as often ([#141](https://github.com/Wave-RF/WaveHouse/issues/141)). -- **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/query/ident.go` (removed, + tests), `internal/app/wire.go`, `docs/src/content/docs/{architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL (a pipe result's key folded no version, so pipes were unaffected; with the tenant's version in every key they now take the same snapshot). The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. The cache now escapes table and scope names itself: a `Namespace` carries them raw and the version index builds its keys with `internal/keyenc`, so neither the structured-query read nor the ingest worker's invalidation escapes them (`query.SafeEncodeToken` is gone) and no name reaches a key unescaped. The suite pins that a name holding a dot, a space or a `%` is read and bumped under one key, and that names which would run together unescaped (`a.0.b` against `a` with scope `b.0.`) stay two entries. Namespace keys are unchanged; an entry's key now carries the tenant's version and escapes the caller's query key whole. All of them live in the process, so nothing stored is orphaned. +- **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/query/ident.go` (removed, + tests), `internal/app/wire.go`, `docs/src/content/docs/{api,architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL (a pipe result's key folded no version, so pipes were unaffected; with the tenant's version in every key they now take the same snapshot). The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. The snapshot is taken before any input a bump invalidates is chosen, the tenant's ClickHouse connection included: both handlers now look up before they resolve the tenant's pool, so a reload that repoints the tenant after a request took the old pool orphans that request's fill instead of filing the old database's rows as fresh under the new tenant version. A tenant on no pool is still a `503` before anything is served, a cached result included. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. The cache now escapes table and scope names itself: a `Namespace` carries them raw and the version index builds its keys with `internal/keyenc`, so neither the structured-query read nor the ingest worker's invalidation escapes them (`query.SafeEncodeToken` is gone) and no name reaches a key unescaped. The suite pins that a name holding a dot, a space or a `%` is read and bumped under one key, and that names which would run together unescaped (`a.0.b` against `a` with scope `b.0.`) stay two entries. Namespace keys are unchanged; an entry's key now carries the tenant's version and escapes the caller's query key whole. All of them live in the process, so nothing stored is orphaned. - **An explicit `false`, `0` or `""` in `config.yaml` is no longer replaced by the key's default** (`internal/config/config.go`, `internal/config/defaults_test.go` (new), `docs/src/content/docs/configuration.mdx`, `AGENTS.md`): [#631](https://github.com/Wave-RF/WaveHouse/issues/631). Defaults lived in cleanenv `env-default` tags, which cleanenv applies after the YAML decode to any field still at its zero value, so it could not tell a key the file set to its zero value from one the file left out. `otel.traces.enabled: false`, `otel.metrics.enabled: false` and `otel.logs.enabled: false` came back `true`; `otel.traces.sample_rate: 0` and `otel.logs.sample_rate: 0` came back `1.0`; `server.shutdown_timeout: 0` came back `10`; `cache.l1_max_cost: 0`, `prometheus.path: ""` and `data_dir: ""` came back as their defaults; `server.port: 0` came back `8080`. All of it was silent. Defaults now live in one Go function, `defaults()`, which `Load` starts from before decoding the file and then applying `WH_*` variables, so the order is env > YAML > default and a key the file sets always wins. **Behaviour change if your file relied on the bug:** a zero you wrote now takes effect. A file that says `sample_rate: 0` now exports no traces (or no DEBUG/INFO logs), where it silently exported everything; a signal set `enabled: false` is now off; `shutdown_timeout: 0` now skips the drain. `cache.l1_max_cost: 0`, `server.port: 0`, and `data_dir: ""` now refuse boot (`cache init: MaxCost can't be zero`, `server.port 0 out of range`, `data_dir (WH_DATA_DIR) is required`) instead of running on the default; an empty `prometheus.path` refuses boot when `prometheus.enabled` is true. Delete the key to get the default back. Env vars are unchanged: they already honoured an explicit zero. New tests load through `config.Load` for every affected key (a YAML zero is kept, an absent key gets the default, env wins in both directions), refuse an `env-default` tag on any field, and pin each documented default in `configuration.mdx` to `defaults()`. - **An embedded queue store that cannot be created fails boot at once, naming the cause** (`internal/mq/embedded.go` (+ tests)): part of [#617](https://github.com/Wave-RF/WaveHouse/issues/617). A regular file at `/nats`, or a `nats` directory that could not be created there, failed JetStream in the background, so boot waited out the server's 5s readiness check and reported only `nats server not ready`. `NewEmbedded` now creates the directory first (at `0700`, as the server does) and refuses boot with the mkdir error. An existing but unwritable `nats` directory still takes the old path. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its cached results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index a32e2198c..3f95cac89 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -578,7 +578,7 @@ The inbound request body is capped at 1 MiB; a body over the cap is rejected wit | 400 / 403 / 502 / 503 | `{"error":"clickhouse query: …","code":"clickhouse.…","retryable":…}` | ClickHouse failed the query: a column dropped since the schema was discovered (`400 clickhouse.rejected`), the role's `max_rows_to_read`/`max_memory_usage` cap, or a `max_execution_time` no longer than `clickhouse.query_timeout` (`400 clickhouse.limit_exceeded`), ClickHouse down (`503 clickhouse.unavailable`, `Retry-After: 5`), … — see [ClickHouse errors on the query paths](#clickhouse-errors-on-the-query-paths) | | 500 | `{"error":"…","code":"clickhouse.unknown","retryable":true}` | A failure with no verdict | | 503 | `{"error":"schema not loaded yet"}` | The tenant's first schema discovery has not succeeded yet, so whether the table exists is not known; `Retry-After: 5` | -| 503 | `{"error":"no ClickHouse connection is open for this tenant"}` | The tenant is on no ClickHouse pool — [no pool could be opened for it](/settings-directory#clickhouse), such as one the connection ceiling refused — so the query cannot run; decided ahead of the cache, so nothing cached before is served either; `Retry-After: 30`, a settings reload retries the pool | +| 503 | `{"error":"no ClickHouse connection is open for this tenant"}` | The tenant is on no ClickHouse pool — [no pool could be opened for it](/settings-directory#clickhouse), such as one the connection ceiling refused — so the query cannot run; decided before anything is served, so nothing cached before is served either; `Retry-After: 30`, a settings reload retries the pool | | 503 | `{"error":"token verifier not ready: the tenant's JWKS has not been fetched yet"}` | A token was supplied, with no valid operator key, while the tenant's JWKS has not been fetched yet; refused before any policy runs, with a `Retry-After: 30` header — see [Authentication](#authentication) | --- @@ -609,7 +609,7 @@ The POST parameter body is capped at 1 MiB; a body over the cap is rejected with | Status | Body | Cause | | ------ | ---- | ----- | | 404 | `{"error":"pipe not found"}` | Pipe name not registered | -| 503 | `{"error":"no ClickHouse connection is open for this tenant"}` | The tenant is on no ClickHouse pool — [no pool could be opened for it](/settings-directory#clickhouse), such as one the connection ceiling refused; decided ahead of the cache; `Retry-After: 30` | +| 503 | `{"error":"no ClickHouse connection is open for this tenant"}` | The tenant is on no ClickHouse pool — [no pool could be opened for it](/settings-directory#clickhouse), such as one the connection ceiling refused; decided before anything is served, a cached result included; `Retry-After: 30` | | 403 | `{"error":"forbidden"}` | Role not in pipe's `allowed_roles` (and not the admin role). Fails closed: a request with no role (no token, or a JWT missing `auth.role_claim`) is denied unless a `default_role` resolves it into the list; a pipe with no `allowed_roles` denies everyone but the admin role. | | 400 | `{"error":"missing required parameter: x"}` | Required parameter not supplied | | 400 | `{"error":"parameter \"x\": unsupported parameter type object"}` | A non-scalar value with no SQL literal form — a JSON object, whether supplied directly or nested as an array element. A JSON **array** is valid and renders as an `IN`-style `(…)` list. | diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 42b8ead34..6a4f83ff8 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -86,7 +86,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request - **ingest.go** — Accepts `POST /v1/ingest?table={table}` in three body shapes: one flat JSON object, a JSON array of them, or NDJSON. The **required** `Content-Type` chooses the format *family* — `application/json` versus the four NDJSON spellings — and within the JSON family the body's first non-whitespace byte picks array versus single object; the bytes never choose the family. Anything that is not exactly one readable media type is a `415`, decided before the body is read: the header is parsed per RFC 9110 §8.3, and because `Content-Type` is a singleton field, repeated header lines must all resolve to the same format and a value carrying a comma is refused unless the value as a whole parses as one media type — a comma inside a *quoted* parameter value is data, so `application/json; a=", application/x-ndjson; b="` is accepted. It then reads the whole (`MaxBytesReader`-capped) body into a pooled buffer and runs the per-format record readers over those bytes, so the `413` lands before any record is processed and peak memory per request is O(body) rather than O(record). Then it validates each record against the discovered schema, optional dedup, and publishes each row through `mq.Publisher` on `mq.Topic{Tenant, Table, Scope}` (the request's tenant, read off its resolved store — `store.Tenant()` — and raw names; the subject it becomes is `internal/mq`'s; a full queue comes back as `mq.ErrQueueFull`, which is the `503` + `Retry-After: 30`, and a broker that cannot be reached or does not answer in time as `mq.ErrUnavailable`, the `503` + `Retry-After: 5`). When dedup is on, a row missing the configured `id_field` can't be deduped: it is logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total` (labeled by `table`), then published un-deduped — or rejected when `dedupe.require_id` is set ([#219](https://github.com/Wave-RF/WaveHouse/issues/219)). - **query.go** — Proxies raw SQL for `POST /v1/ops/query` straight to the `?tenant=`'s ClickHouse HTTP interface (`chconn.Pools.Target` by the resolved store's tenant; the zero target — no pool — is a `503` with `Retry-After`). **Not cached** — sets `Cache-Control: no-store` so every request hits ClickHouse; DateTime is rendered ISO-8601 via `date_time_output_format=iso` (the Go-side type conversion lives in the structured-query / pipes path, not here). - **stream.go** — Real-time streaming via SSE. Callers select a table with the `?table=` query parameter. Each connection registers one `Subscriber` (the `stream/` package) with both the event `Hub` (under its `(topic, role)`) and the shared keepalive wheel, then drains both from a single byte-pump — so idle streams keep emitting `:` keepalive comments (surviving reverse-proxy idle timeouts) while live events arrive already projected and serialized. Per-event projection/serialization happens **once per role** in the `Hub`, not once per subscriber ([#294](https://github.com/Wave-RF/WaveHouse/issues/294)); the handler also snapshots the connection's JWT claims onto the `Subscriber`, which the `Hub` evaluates per subscriber when the role carries a row-level `filter` ([#319](https://github.com/Wave-RF/WaveHouse/issues/319)). Gap-fill replay (`mq.Replayer.ReplaySince` on the connection's `mq.Topic` — a `DeliverByStartTime` consumer inside `internal/mq`) stays per-connection (low-volume, one-time on connect). A stream ends, a gap-fill in progress included, when the server begins shutting down (`Closing`) or its `Subscriber` is evicted because its tenant is no longer served (`Hub.Prune`); one admitted just before the reload that stopped serving its tenant, and registered just after the prune, is ended right after it registers (`Served`). -- **schema.go** — Schema discovery API of one tenant, the `?tenant=` (`opsStore`): list all schemas, get one table, trigger refresh. `lookupSchema`, shared with the ingest and structured-query handlers, is the one reading of a `SchemaRegistry.Lookup` miss: `503` with `Retry-After` before the tenant's first discovery (`ErrNotLoaded`, or no registry built yet), `404` for a table the discovered schema lacks; the list answers the same `503` rather than `[]`. A refresh of a tenant on no pool (`discovery.ErrNoConnection`) is a `503` with `Retry-After` too. The handlers hold `RegistrySource`, `func(*settings.Store) *discovery.SchemaRegistry`, and the query paths a `func(*settings.Store) driver.Conn` beside it — each resolves the request's tenant per call, and a nil connection (a tenant no pool could be opened for, such as by the connection ceiling) is a `503` ahead of the cache, so nothing cached before is served. +- **schema.go** — Schema discovery API of one tenant, the `?tenant=` (`opsStore`): list all schemas, get one table, trigger refresh. `lookupSchema`, shared with the ingest and structured-query handlers, is the one reading of a `SchemaRegistry.Lookup` miss: `503` with `Retry-After` before the tenant's first discovery (`ErrNotLoaded`, or no registry built yet), `404` for a table the discovered schema lacks; the list answers the same `503` rather than `[]`. A refresh of a tenant on no pool (`discovery.ErrNoConnection`) is a `503` with `Retry-After` too. The handlers hold `RegistrySource`, `func(*settings.Store) *discovery.SchemaRegistry`, and the query paths a `func(*settings.Store) driver.Conn` beside it — each resolves the request's tenant per call, and a nil connection (a tenant no pool could be opened for, such as by the connection ceiling) is a `503` before anything is served, a cached result included. The cached paths resolve it after their cache `Lookup`, so the snapshot predates the connection (see `cache.go` below). - **dlq.go** — DLQ stats endpoint (`GET /v1/ops/dlq/stats`): asks `mq.DeadLetterStats.DeadLetterCounts` for one tenant's per-table parked counts (optionally one table) and its total — the tenant `?tenant=` names, read strictly by `opsTenant`, tenant `0` without it. The tenant is looked up in the MQ, not the settings registry, so a rejected or removed tenant's parked rows are read like a served one's; a tenant with no dead-letter queue (`mq.ErrNoDeadLetterQueue`) is a 404, and any other failure to read it a 500. The queue itself is `internal/mq`'s. - **health.go** — Liveness (`/livez`), readiness (`/readyz`), and a content-free `Online` ping (`/v1/health`, the SDK's public liveness check); `/healthz` is a permanent alias of `/livez`, and `/health`/`/ready` are deprecated aliases. All three consult an optional `BootState` so they can return 503 while boot-time schema discovery is still failing in the retry loop (see `internal/app`; over a nested directory, while no tenant's has succeeded); once `BootState.Set(nil)` fires, `/livez` returns 200 and stays there. `/readyz` additionally runs a `Ping` each call — `chconn.Pools.Ping` in production: every open pool at once, ready at the first answer, every pool's error joined when none answers; `/v1/health` deliberately does not. @@ -112,7 +112,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ ### `cache/` — Query Cache -- **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on — one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)) — each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope by their raw names, which the cache escapes where it builds a key. `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. +- **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on — one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)) — each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope by their raw names, which the cache escapes where it builds a key. `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The snapshot is taken before any input a bump invalidates is chosen, the tenant's connection included: a reload that repoints the tenant runs `Pools.Reconcile` and then `InvalidateTenant`, so a request that took the old pool files the old database's rows under a version the bump has already orphaned. The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. - **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: a namespace key is `..
.
.`, its fields joined with `keyenc.Join` so a dot in a table or scope name is escaped rather than read as a separator, and an entry's key, `|.||…` with the caller's query key escaped whole (its `:` become `%3A`), folds in the tenant's version and each dependency's namespace key and namespace version, so bumping a table (a scopeless write) or one scope — scope is reserved and empty today, so every write is the whole-table bump — orphans every dependent entry without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results, and the flat directory's single tenant simply carries the `0` prefix. `BumpTenant` (behind `InvalidateTenant`) advances the tenant version that leads every namespace key of one tenant, orphaning its every namespace and every cached query in one step (the tenant version is folded into every query key, so a pipe result with no dependencies is orphaned too) — a table no bump ever keyed included, which an enumeration of the index would miss — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. diff --git a/internal/api/cache_tenant_test.go b/internal/api/cache_tenant_test.go index 742b55f25..38ea67267 100644 --- a/internal/api/cache_tenant_test.go +++ b/internal/api/cache_tenant_test.go @@ -64,6 +64,13 @@ var cachedRoutes = []struct{ name, path, body string }{ // to conn and c. Every request resolves to the viewer role, which may read // clicks.page and run top_pages. func cachedRouter(t *testing.T, tenants *settings.Registry, conn driver.Conn, c cache.Cache) http.Handler { + t.Helper() + return cachedRouterOver(t, tenants, fixedConn(conn), c) +} + +// cachedRouterOver is cachedRouter with the tenant's connection chosen per +// request by connFor. +func cachedRouterOver(t *testing.T, tenants *settings.Registry, connFor func(*settings.Store) driver.Conn, c cache.Cache) http.Handler { t.Helper() reg := testRegistry(t) viewer := staticPolicy(&policy.Policy{ @@ -74,8 +81,8 @@ func cachedRouter(t *testing.T, tenants *settings.Registry, conn driver.Conn, c return NewRouter(Dependencies{ Tenants: tenants, Ingest: NewIngestHandler(fixedRegistry(reg), &testutil.MockPublisher{}), - StructuredQuery: NewStructuredQueryHandler(fixedConn(conn), c, fixedRegistry(reg), viewer, func(*settings.Store) int { return 60 }, timeout, nil), - Pipes: NewPipesHandler(staticPipes(&pipes.NamedQuery{Name: "top_pages", SQL: "SELECT 1", AllowedRoles: []string{"viewer"}}), viewer, fixedConn(conn), c, timeout), + StructuredQuery: NewStructuredQueryHandler(connFor, c, fixedRegistry(reg), viewer, func(*settings.Store) int { return 60 }, timeout, nil), + Pipes: NewPipesHandler(staticPipes(&pipes.NamedQuery{Name: "top_pages", SQL: "SELECT 1", AllowedRoles: []string{"viewer"}}), viewer, connFor, c, timeout), Query: &QueryHandler{}, SSE: NewStreamHandler(stream.NewHub(nil, nil, nil), nil), Health: &HealthHandler{}, @@ -255,6 +262,49 @@ func TestCachedRoutes_BumpDuringQueryOrphansTheFill(t *testing.T) { } } +// The snapshot is taken before the tenant's pool is chosen. A reload that +// repoints the tenant — Pools.Reconcile, then InvalidateTenant — landing +// between the two leaves the request on the old pool: its fill, read from +// the old database, is orphaned by the bump rather than filed as fresh under +// the new tenant version. And a tenant on no pool is a 503 even when its +// Lookup hit (#583 story 6). +func TestCachedRoutes_ReloadAsThePoolIsTakenOrphansTheFill(t *testing.T) { + for _, route := range cachedRoutes { + t.Run(route.name, func(t *testing.T) { + l1, err := cache.NewLocal(1 << 20) + require.NoError(t, err) + t.Cleanup(func() { _ = l1.Close() }) + conn := &countingConn{} + var taken atomic.Int32 + var noPool atomic.Bool + connFor := func(*settings.Store) driver.Conn { + if noPool.Load() { + return nil + } + if taken.Add(1) == 1 { + require.NoError(t, l1.InvalidateTenant(t.Context(), tenant.Default)) + } + return conn + } + router := cachedRouterOver(t, testTenants(), connFor, l1) + xcache := func() string { + w := serveAs(t, router, route.path, route.body, "") + require.Equal(t, http.StatusOK, w.Code, "body: %s", w.Body.String()) + l1.Wait() + return w.Header().Get("X-Cache") + } + assert.Equal(t, "MISS", xcache()) + assert.Equal(t, "MISS", xcache(), "the fill of a query on the pool a reload replaced is orphaned") + assert.Equal(t, "HIT", xcache()) + assert.Equal(t, int32(2), conn.queries.Load()) + + noPool.Store(true) + w := serveAs(t, router, route.path, route.body, "") + assertUnavailable(t, w, noConnectionMessage, retryAfterPool) + }) + } +} + // A structured query files its result under the table as the request names // it, raw — the namespace the ingest worker bumps after an insert into that // table (ingest's TestFlushTable_BumpsWhatTheReadFiles) — and the cache diff --git a/internal/api/pipes.go b/internal/api/pipes.go index c85785470..81dd9d710 100644 --- a/internal/api/pipes.go +++ b/internal/api/pipes.go @@ -152,32 +152,35 @@ func (h *PipesHandler) Execute(w http.ResponseWriter, r *http.Request) { return } - // The tenant's pool, ahead of the cache: a tenant on none — its tuple - // could not be opened, such as by the connection ceiling — fails - // closed rather than serve what it cached before (#583 story 6). - conn := connOf(h.CHConn, store) - if conn == nil { - writeUnavailable(w, noConnectionMessage, retryAfterPool) - return - } - // Cache. A pipe can read several tables, but the current pipe impl doesn't // expose its table/scope dependencies, so we pass no deps: the result folds // the tenant's version alone, so InvalidateTenant orphans it but no insert - // does (TTL-bound until #343). The snapshot is of the versions before the - // query runs, so a bump landing mid-query orphans the fill (#382). + // does (TTL-bound until #343). The snapshot is of the versions before + // anything the query reads is chosen, so a bump landing after — mid-query + // (#382), or a reload repointing the tenant once its pool below is taken — + // orphans the fill. // TODO: once pipes expose their tables/scopes, pass them as deps here so writes // invalidate cached pipe results. cacheKey := queryCacheKey(store.Tenant(), sql, params) + var entry cache.Entry var snap cache.Snapshot if h.Cache != nil { - var entry cache.Entry - if entry, snap, _ = h.Cache.Lookup(r.Context(), store.Tenant(), cacheKey, nil); entry.Value != nil { - w.Header().Set("Content-Type", "application/json") - w.Header().Set("X-Cache", "HIT") - _, _ = w.Write(entry.Value) - return - } + entry, snap, _ = h.Cache.Lookup(r.Context(), store.Tenant(), cacheKey, nil) + } + + // The tenant's pool, ahead of serving a hit: a tenant on none — its + // tuple could not be opened, such as by the connection ceiling — fails + // closed rather than serve what it cached before (#583 story 6). + conn := connOf(h.CHConn, store) + if conn == nil { + writeUnavailable(w, noConnectionMessage, retryAfterPool) + return + } + if entry.Value != nil { + w.Header().Set("Content-Type", "application/json") + w.Header().Set("X-Cache", "HIT") + _, _ = w.Write(entry.Value) + return } // Execute with singleflight. diff --git a/internal/api/structured_query.go b/internal/api/structured_query.go index 3c10d29ed..9b192a0b3 100644 --- a/internal/api/structured_query.go +++ b/internal/api/structured_query.go @@ -159,15 +159,6 @@ func (h *StructuredQueryHandler) Handle(w http.ResponseWriter, r *http.Request) return } - // The tenant's pool, ahead of the cache: a tenant on none — its tuple - // could not be opened, such as by the connection ceiling — fails - // closed rather than serve what it cached before (#583 story 6). - conn := connOf(h.CHConn, store) - if conn == nil { - writeUnavailable(w, noConnectionMessage, retryAfterPool) - return - } - // Cache key, led by the tenant the store was resolved for (#583 story 8); // the singleflight key too. cacheKey := queryCacheKey(store.Tenant(), result.SQL, result.Params) @@ -179,17 +170,29 @@ func (h *StructuredQueryHandler) Handle(w http.ResponseWriter, r *http.Request) // worker's invalidation passes them; the cache escapes both sides alike. deps := []cache.Namespace{{Tenant: store.Tenant(), Table: table, Scope: scope}} - // Try cache. The snapshot is of the versions before the query runs, so a - // write landing mid-query orphans the fill (#382). + // The snapshot is of the versions before anything the query reads is + // chosen, so a bump landing after — an insert mid-query (#382), or a + // reload repointing the tenant once its pool below is taken — orphans the + // fill. + var entry cache.Entry var snap cache.Snapshot if h.Cache != nil { - var entry cache.Entry - if entry, snap, _ = h.Cache.Lookup(r.Context(), store.Tenant(), cacheKey, deps); entry.Value != nil { - w.Header().Set("Content-Type", "application/json") - w.Header().Set("X-Cache", "HIT") - _, _ = w.Write(entry.Value) - return - } + entry, snap, _ = h.Cache.Lookup(r.Context(), store.Tenant(), cacheKey, deps) + } + + // The tenant's pool, ahead of serving a hit: a tenant on none — its + // tuple could not be opened, such as by the connection ceiling — fails + // closed rather than serve what it cached before (#583 story 6). + conn := connOf(h.CHConn, store) + if conn == nil { + writeUnavailable(w, noConnectionMessage, retryAfterPool) + return + } + if entry.Value != nil { + w.Header().Set("Content-Type", "application/json") + w.Header().Set("X-Cache", "HIT") + _, _ = w.Write(entry.Value) + return } // Bare Select reads: this handler resolved the grant for "select" (above), diff --git a/internal/api/tenant_clickhouse_test.go b/internal/api/tenant_clickhouse_test.go index d28dabd34..faab375ca 100644 --- a/internal/api/tenant_clickhouse_test.go +++ b/internal/api/tenant_clickhouse_test.go @@ -109,8 +109,9 @@ func selectAllQuery() query.StructuredQuery { return query.StructuredQuery{Selec // A tenant on no pool — its tuple could not be opened, such as by the // connection ceiling — fails closed on every route that reaches its -// ClickHouse: a 503 with Retry-After ahead of the cache, so nothing it -// cached before is served either, and on the refresh, which cannot run. +// ClickHouse: a 503 with Retry-After before anything is served, so nothing +// it cached before is served either (TestCachedRoutes_ReloadAsThePoolIsTakenOrphansTheFill +// pins that with a hit), and on the refresh, which cannot run. func TestClickHouseRoutes_NoPoolIs503(t *testing.T) { t.Parallel() reg := testRegistry(t) diff --git a/internal/cache/cache.go b/internal/cache/cache.go index c9e66a96b..745adb58a 100644 --- a/internal/cache/cache.go +++ b/internal/cache/cache.go @@ -14,10 +14,12 @@ type Entry struct { TTL time.Duration // remaining } -// Snapshot is the dependency versions a Lookup observed. Set files a result -// under the snapshot taken before its query ran, so a bump that lands while -// the query runs orphans the fill rather than re-homing pre-write rows under -// the post-bump versions (#382). The zero Snapshot makes Set a no-op. +// Snapshot is the dependency versions a Lookup observed. A caller takes it +// before choosing any input a bump invalidates — the tenant's connection as +// well as the rows its query reads — and Set files the result under it, so a +// bump that lands after the Lookup orphans the fill rather than re-homing +// what was read before the bump under the post-bump versions (#382). The +// zero Snapshot makes Set a no-op. type Snapshot struct { key string // the backend's key for the entry at the observed versions } @@ -42,8 +44,8 @@ type Cache interface { // ErrForeignDependency. An error is a miss with a zero Snapshot. Lookup(ctx context.Context, id tenant.ID, sha string, deps []Namespace) (Entry, Snapshot, error) - // Set stores value under snap, the Snapshot a Lookup returned before the - // value was computed. It returns an error only when the backend failed; + // Set stores value under snap, the Snapshot a Lookup returned before any + // input of the value was chosen. It returns an error only when the backend failed; // a value the cache declines to keep — too large, refused admission, a // non-positive ttl, or a zero snap — is not an error. Set(ctx context.Context, snap Snapshot, value []byte, ttl time.Duration) error diff --git a/internal/cache/export_test.go b/internal/cache/export_test.go new file mode 100644 index 000000000..5e45e3ef6 --- /dev/null +++ b/internal/cache/export_test.go @@ -0,0 +1,12 @@ +package cache + +// Len counts the unexpired entries l holds, for the conformance suite's +// Options.Entries: no Lookup reads the key a zero snapshot would land under. +func (l *LocalCache) Len() int { + n := 0 + l.cache.IterValues(func([]byte) bool { + n++ + return false + }) + return n +} diff --git a/internal/cache/local_test.go b/internal/cache/local_test.go index 2fdb3237e..5a62ebda6 100644 --- a/internal/cache/local_test.go +++ b/internal/cache/local_test.go @@ -21,5 +21,8 @@ func newLocal(t *testing.T) cache.Cache { func TestLocalCache_Conformance(t *testing.T) { t.Parallel() - cachetest.Run(t, newLocal, cachetest.Options{MaxValueBytes: localMaxCost}) + cachetest.Run(t, newLocal, cachetest.Options{ + MaxValueBytes: localMaxCost, + Entries: func(c cache.Cache) int { return c.(*cache.LocalCache).Len() }, + }) } diff --git a/internal/ingest/worker_test.go b/internal/ingest/worker_test.go index 3a6ed4afa..354292ab7 100644 --- a/internal/ingest/worker_test.go +++ b/internal/ingest/worker_test.go @@ -598,12 +598,12 @@ func TestInvalidate_ReachesOneTenantsEntries(t *testing.T) { assert.Equal(t, []byte("globex rows"), e.Value, "globex's entry survives acme's insert") } -// From the envelope the /v1/ingest producer publishes to the real cache: an -// insert into a table whose name holds a dot or a space — scoped or not — -// orphans the whole-table result a structured query on that table filed, -// under the raw name the request carries (the namespace internal/api's -// TestStructuredQuery_RawTableNameMeetsTheInsertsBump reads through), and -// leaves another table's. +// From an envelope carrying the raw table and scope (makeEnvelope) to the +// real cache: an insert into a table whose name holds a dot or a space — +// scoped or not — orphans the whole-table result a structured query on that +// table filed, under the raw name the request carries (the namespace +// internal/api's TestStructuredQuery_RawTableNameMeetsTheInsertsBump reads +// through), and leaves another table's. func TestFlushTable_BumpsWhatTheReadFiles(t *testing.T) { t.Parallel() ok := &testutil.MockRoundTripper{Fn: func(*http.Request) (*http.Response, error) { diff --git a/internal/testutil/cachetest/cachetest.go b/internal/testutil/cachetest/cachetest.go index acc24a134..405e42a9e 100644 --- a/internal/testutil/cachetest/cachetest.go +++ b/internal/testutil/cachetest/cachetest.go @@ -26,6 +26,11 @@ type Options struct { // NewPair returns two instances over one shared store, as two processes // see it; nil skips the cross-instance cases. NewPair func(t *testing.T) (a, b cache.Cache) + + // Entries counts the entries a cache from the factory holds. No Lookup + // reads the key a zero Snapshot would land under, so only a count shows + // that one stored nothing; nil skips that case. + Entries func(c cache.Cache) int } // Run runs the suite, each case on a fresh cache from newCache. @@ -40,7 +45,7 @@ func Run(t *testing.T, newCache func(t *testing.T) cache.Cache, opts Options) { {"overwrite", testOverwrite}, {"ttl expiry", testTTLExpiry}, {"non-positive ttl stores nothing", testNonPositiveTTL}, - {"zero snapshot stores nothing", testZeroSnapshot}, + {"zero snapshot stores nothing", func(t *testing.T, c cache.Cache) { testZeroSnapshot(t, c, opts.Entries) }}, {"deps order does not matter", testDepsOrder}, {"deps are part of the key", testDepsKeyed}, {"tenant isolation", testTenantIsolation}, @@ -169,10 +174,17 @@ func testNonPositiveTTL(t *testing.T, c cache.Cache) { } } -func testZeroSnapshot(t *testing.T, c cache.Cache) { +// A zero Snapshot — what a failed Lookup returns — stores nothing, and is not +// an error. The fill after it proves the count sees what Set stores. +func testZeroSnapshot(t *testing.T, c cache.Cache, entries func(cache.Cache) int) { require.NoError(t, c.Set(context.Background(), cache.Snapshot{}, []byte("rows"), ttl)) settle(c) - assertMiss(t, c, acme, "") + if entries == nil { + t.Skip("the backend has no Options.Entries") + } + assert.Zero(t, entries(c)) + fill(t, c, acme, "q", "rows") + assert.Equal(t, 1, entries(c)) } func testDepsOrder(t *testing.T, c cache.Cache) { @@ -210,16 +222,14 @@ func testTenantIsolation(t *testing.T, c cache.Cache) { } // Every version an entry is filed under is its own tenant's, so a Lookup -// naming another tenant's namespace is refused, and its snapshot stores -// nothing. +// naming another tenant's namespace is refused with a zero snapshot, which +// stores nothing (testZeroSnapshot): a handler Sets whatever a Lookup +// returned, error or not. func testForeignDependency(t *testing.T, c cache.Cache) { e, snap, err := c.Lookup(context.Background(), acme, "q", []cache.Namespace{ns(globex, "events", "")}) require.ErrorIs(t, err, cache.ErrForeignDependency) assert.Nil(t, e.Value) - require.NoError(t, c.Set(context.Background(), snap, []byte("rows"), ttl)) - settle(c) - assertMiss(t, c, acme, "q", ns(acme, "events", "")) - assertMiss(t, c, globex, "q", ns(globex, "events", "")) + assert.Zero(t, snap, "a refused Lookup returns the zero Snapshot") } // A scoped bump orphans that scope and the whole-table view; a scopeless From 19720fb3d9703694ef64f931f15cbffaf6ba4a53 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 04:58:18 -0400 Subject: [PATCH 32/79] fix(api): classify writes the way ClickHouse lexes them isMutation missed a write hidden behind a backslash-escaped quote (in '...', "..." or backticks), a nested block comment, or leading whitespace outside space, tab, CR and LF (\v, \f, a no-break space, a byte-order mark and the other Unicode spaces ClickHouse skips). The missed write went through Query, which ran it and then failed the call, so a client retrying the 500 wrote again. The same gaps could make a read look like a write. The scanners now share one whitespace, comment and quote skipper that follows ClickHouse's lexer: a backslash or doubled quote escapes the next byte, block comments nest, and the whitespace set is the one ClickHouse 26.6 accepts, each character checked against a live server. The CTE-name lookahead uses the same skipper. Docs: pipes.mdx says to write each placeholder bare, since a quoted one lets a string value's own quotes close the template's; the write-pipe section lists the verbs the classifier recognizes; "only write path" now reads "only way to define or change a pipe". Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- CHANGELOG.md | 2 + docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/pipes.mdx | 10 +- internal/api/clickhouse_exec.go | 226 ++++++++++++++------------ internal/api/clickhouse_exec_test.go | 47 ++++++ internal/pipes/pipes.go | 4 +- 6 files changed, 181 insertions(+), 110 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 976e847ab..e9a7f7f33 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -79,6 +79,8 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed - **A pipe that writes runs on every call instead of being answered from the cache** (`internal/api/pipes.go` (+ tests), `docs/src/content/docs/{pipes.mdx,api.md,architecture.md}`, `AGENTS.md`): fixes [#386](https://github.com/Wave-RF/WaveHouse/issues/386). `/v1/pipes/{name}` sent a write's SQL to ClickHouse through `Exec`, but still cached the `[]` it returned and coalesced identical calls in flight, so a repeat within the TTL answered `200` without executing and concurrent identical calls became one write — silently dropped writes, and with a shared cache ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)) on every instance. A pipe whose bound SQL `isMutation` classifies as a write — the same classifier that picks `Exec` — now skips the cache lookup, the fill and singleflight, and answers `X-Cache: BYPASS` with `Cache-Control: no-store`, so an HTTP cache in front of a `GET` cannot drop the write either. Classification stays automatic rather than a declared pipe property, so an operator cannot forget to mark one, and costs no ClickHouse round trip. Read pipes are unchanged. Not in this fix: a write pipe still does not invalidate cached reads of the table it writes ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). +- **The write classifier skips whitespace, comments and quoted text the way ClickHouse's lexer does** (`internal/api/clickhouse_exec.go` (+ tests)): `isMutation` picks `Exec` for a write, and since [#386](https://github.com/Wave-RF/WaveHouse/issues/386) keeps a write pipe out of the cache. It missed a write behind a backslash-escaped quote (`'it\'s'`, and the same inside `"…"` and `` `…` ``), behind a nested block comment (`/* a /* b */ SELECT */ INSERT …`), or behind leading whitespace other than space, tab, CR and LF: `\v`, `\f`, a no-break space, a byte-order mark, and the other Unicode spaces ClickHouse skips. A missed write went through `Query`, which ran it and then failed the call with a `500`; the TypeScript SDK retries that, so one call could write three times. The same gaps could make a read look like a write, which runs through `Exec` and answers `[]`. +- **The pipes page no longer says a parameter can never break out of its literal** (`docs/src/content/docs/pipes.mdx`, `internal/pipes/pipes.go`): that holds only for a placeholder written bare. A string value brings its own quotes, so inside a quoted placeholder they close the template's: the body `{"id": " OR 1=1 OR id = "}` turns `WHERE id = '{{id}}'` into `WHERE id = '' OR 1=1 OR id = ''`, which matches every row. The page now says to write each placeholder bare, never inside quotes. - **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/app/wire.go`, `docs/src/content/docs/{architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` and pipe execution rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL. The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its structured-query results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. - **The `?token=` strip no longer repairs a query string that does not parse** (`internal/auth/auth.go`): `bearerToken` removed a query-string token by parsing the query, deleting `token`, and re-encoding what was left — and `url.ParseQuery` skips a pair it cannot read, so the re-encoding erased that pair. A handler that parses the query strictly in order to refuse a malformed one would then see a clean query: `GET /v1/ops/pipes?tenant=acme;x=1&token=…` would have answered `200` with the default tenant's pipes. The token is read exactly as before and a query that parses is rewritten exactly as before; a query that does not parse now loses its token pairs and nothing else, byte for byte. Pinned through `api.NewRouter` with the real authenticator, since a handler-level test never runs the middleware that rewrote the URL. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 3aff5cc6e..0f2473bb3 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -78,7 +78,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request - **router.go** — Route definitions. Public: `/livez`, `/readyz`, and the content-free `/v1/health` SDK ping (plus the permanent `/healthz` alias and the deprecated `/health`, `/ready` aliases). Policy-gated: `/v1/ingest?table={table}`, `/v1/query?table={table}` (structured), `/v1/pipes/{name}` (named pipes), `/v1/stream`. Admin-only (`RequireAdmin` — role == `policy.admin_role`, or a request bearing the operator key's operator bit, which passes even under a nil policy; over a nested settings directory `NewRouter` mounts the gate with no policy at all, whatever `Dependencies.PolicySource` was wired, so the operator key alone passes): `/v1/ops/schema/*`, `/v1/ops/dlq/stats`, `GET /v1/ops/pipes[/{name}]`, `/v1/ops/settings/reload`, `/v1/ops/query` (raw SQL — same gate as the rest of `/v1/ops/*`). - **auth middleware** — the JWT/JWKS authentication middleware is its own package, [`auth/`](#auth--authentication); the router runs it on every `/v1/*` route. - **tenant.go** — `TenantMW` resolves the request's tenant ahead of the auth middleware on every `/v1` route outside `/v1/ops/*`: the [`X-Tenant-ID`](/deployment#multi-tenant-deployments) header (absent means `tenant.Default`), validated by `tenant.Parse` (`400`), looked up in the `settings.Registry` (`resolveStore`: `404` for an id it does not hold, a bare `503` for a tenant whose folder was rejected — the findings stay out of a body answered before authentication), and the resolved `*settings.Store` stored in the request context (`WithStore` / `StoreFromContext` — here rather than in `tenant/`, because `settings` names `tenant.ID`). A handler reads the store once and passes it down as an argument — the per-tenant getters it holds take it as a parameter (`(*settings.Store).Policy`, `.DedupeFor`, `.DefaultMaxRows`, … in production) — and nothing below a handler reads the context; a tenant route reached without a resolved store answers `500` rather than fall back to a tenant. The probes, `/version`, the metrics path, and `/v1/ops/*` are tenant-exempt; an ops route that addresses one tenant — the admin pipe reads, the schema routes, the raw-SQL proxy, the settings reload and the DLQ stats — names it in `?tenant=` (`opsTenant`, and `opsStore` over it for the routes that need the tenant's store; the DLQ stats need none, since the MQ holds the queue), parsed strictly so that a query `url.ParseQuery` would half-read is a `400` rather than a read of the default tenant, which is what it means when absent on the reads (on the reload, absent is the whole directory). -- **pipes.go** — Named query pipe handlers: admin listing (`GET /v1/ops/pipes[/{name}]`, read per request from its `pipes.Source`) and execution with parameter binding. A read is cached and coalesced; a write — bound SQL that `isMutation` (`clickhouse_exec.go`) classifies as one — bypasses both and runs every call. `pipes.json` is the only write path. +- **pipes.go** — Named query pipe handlers: admin listing (`GET /v1/ops/pipes[/{name}]`, read per request from its `pipes.Source`) and execution with parameter binding. A read is cached and coalesced; a write — bound SQL that `isMutation` (`clickhouse_exec.go`) classifies as one — bypasses both and runs every call. `pipes.json` is the only way to define or change a pipe. - **structured_query.go** — Handler for `POST /v1/query?table={table}`: validates query AST, enforces permissions, builds and executes SQL. - **ingest.go** — Accepts `POST /v1/ingest?table={table}` in three body shapes: one flat JSON object, a JSON array of them, or NDJSON. The **required** `Content-Type` chooses the format *family* — `application/json` versus the four NDJSON spellings — and within the JSON family the body's first non-whitespace byte picks array versus single object; the bytes never choose the family. Anything that is not exactly one readable media type is a `415`, decided before the body is read: the header is parsed per RFC 9110 §8.3, and because `Content-Type` is a singleton field, repeated header lines must all resolve to the same format and a value carrying a comma is refused unless the value as a whole parses as one media type — a comma inside a *quoted* parameter value is data, so `application/json; a=", application/x-ndjson; b="` is accepted. It then reads the whole (`MaxBytesReader`-capped) body into a pooled buffer and runs the per-format record readers over those bytes, so the `413` lands before any record is processed and peak memory per request is O(body) rather than O(record). Then it validates each record against the discovered schema, optional dedup, and publishes each row through `mq.Publisher` on `mq.Topic{Tenant, Table, Scope}` (the request's tenant, read off its resolved store — `store.Tenant()` — and raw names; the subject it becomes is `internal/mq`'s; a full queue comes back as `mq.ErrQueueFull`, which is the `503` + `Retry-After`). When dedup is on, a row missing the configured `id_field` can't be deduped: it is logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total` (labeled by `table`), then published un-deduped — or rejected when `dedupe.require_id` is set ([#219](https://github.com/Wave-RF/WaveHouse/issues/219)). - **query.go** — Proxies raw SQL for `POST /v1/ops/query` straight to the `?tenant=`'s ClickHouse HTTP interface (`chconn.Pools.Target` by the resolved store's tenant; the zero target — no pool — is a `503` with `Retry-After`). **Not cached** — sets `Cache-Control: no-store` so every request hits ClickHouse; DateTime is rendered ISO-8601 via `date_time_output_format=iso` (the Go-side type conversion lives in the structured-query / pipes path, not here). diff --git a/docs/src/content/docs/pipes.mdx b/docs/src/content/docs/pipes.mdx index d7101edf1..e8a6ea768 100644 --- a/docs/src/content/docs/pipes.mdx +++ b/docs/src/content/docs/pipes.mdx @@ -78,7 +78,7 @@ Independently, every `required` declared parameter must be supplied regardless o ### How a value becomes SQL -Bound values are **inlined directly into the SQL string** (not sent as positional driver parameters — that lets a parameter sit anywhere ClickHouse allows a literal, including `LIMIT`). Inlining is type-aware and escaped: +Bound values are **inlined directly into the SQL string** (not sent as positional driver parameters — that lets a parameter sit anywhere ClickHouse allows a literal, including `LIMIT`). Inlining is type-aware and escaped, and a string brings its own quotes, so **write each placeholder bare, never inside quotes** — `WHERE id = {{id}}`, not `WHERE id = '{{id}}'`: | Supplied value | Rendered as | Note | | -------------- | ----------- | ---- | @@ -89,7 +89,7 @@ Bound values are **inlined directly into the SQL string** (not sent as positiona | array | `('a', 'b')` | parenthesized list of escaped elements — for `IN` clauses (see below) | | null | `NULL` | | -Every leaf value is escaped the same way — including each element of an array — so a parameter value can't break out of its literal and inject SQL. The SQL *structure* still comes only from the operator-authored template. Values with no safe scalar form are **rejected** with a `400`: a JSON object, and an empty array (which would render as the invalid `IN ()`). +Every leaf value is escaped the same way — including each element of an array — so a value bound to a bare placeholder can't break out of its literal and inject SQL, and the SQL *structure* comes only from the operator-authored template. A quoted placeholder gives that up: the value's own quotes close the template's, so in `WHERE id = '{{id}}'` the body `{"id": " OR 1=1 OR id = "}` binds to `WHERE id = '' OR 1=1 OR id = ''`, which matches every row. Values with no safe scalar form are **rejected** with a `400`: a JSON object, and an empty array (which would render as the invalid `IN ()`). #### Array parameters and `IN` lists @@ -122,7 +122,7 @@ This is the *only* authorization check on the execute path — pipes deliberatel ## Creating and managing pipes -Pipes are defined in the settings directory's `pipes.json` — a `{"pipes": [...]}` list of the definitions above — and the files are the only write path: edit the file (standalone: on the host; on WaveHouse Cloud the control plane writes it) and the running server re-validates and adopts it on file change, `SIGHUP`, or `POST /v1/ops/settings/reload`, so a create, update, or delete applies without a restart. Validation (`wavehouse validate`, boot, and every reload) rejects an unknown key, a duplicate or empty name, empty SQL, an unknown parameter `type`, and an `allowed_roles` entry not declared in `roles.json`; a rejected reload keeps the previous pipes in effect. See [Settings Directory — `pipes.json`](/settings-directory#pipesjson) for the full rules. +Pipes are defined in the settings directory's `pipes.json` — a `{"pipes": [...]}` list of the definitions above — and the files are the only way to define or change one: edit the file (standalone: on the host; on WaveHouse Cloud the control plane writes it) and the running server re-validates and adopts it on file change, `SIGHUP`, or `POST /v1/ops/settings/reload`, so a create, update, or delete applies without a restart. Validation (`wavehouse validate`, boot, and every reload) rejects an unknown key, a duplicate or empty name, empty SQL, an unknown parameter `type`, and an `allowed_roles` entry not declared in `roles.json`; a rejected reload keeps the previous pipes in effect. See [Settings Directory — `pipes.json`](/settings-directory#pipesjson) for the full rules. ```json { @@ -180,9 +180,9 @@ The response is a JSON array of rows. Results flow through the shared in-process ### Pipes that write -A pipe's SQL may be a write — `INSERT`, `ALTER … DELETE`, `CREATE`, and the rest of ClickHouse's statements that return no rows, including a `WITH … INSERT`. Such a pipe runs on **every** call: it never reads or fills the cache and is never coalesced with an identical call in flight, so ten identical calls are ten writes. The response is `[]` with `X-Cache: BYPASS` and `Cache-Control: no-store`, so an HTTP cache in front of a `GET` does not answer a repeat either. WaveHouse classifies the statement from its leading keyword (after any `WITH` list), with the same classifier that sends it to ClickHouse as a write, so no pipe property marks it. +A pipe's SQL may be a write: a statement led by a write verb WaveHouse recognizes — `INSERT`, `UPDATE`, `DELETE`, `ALTER` (so `ALTER … DELETE`), `CREATE`, `DROP`, `TRUNCATE`, `RENAME`, `EXCHANGE`, `REPLACE`, `OPTIMIZE`, `ATTACH`, `DETACH`, `GRANT`, `REVOKE`, `KILL`, `SET`, `USE` or `SYSTEM` — directly or after a `WITH` list, as in `WITH … INSERT`. Such a pipe runs on **every** call: it never reads or fills the cache and is never coalesced with an identical call in flight, so ten identical calls are ten writes. The response is `[]` with `X-Cache: BYPASS` and `Cache-Control: no-store`, so an HTTP cache in front of a `GET` does not answer a repeat either. WaveHouse classifies the statement from its leading keyword (after any `WITH` list), with the same classifier that sends it to ClickHouse as a write, so no pipe property marks it; a statement led by any other keyword runs as a read. -`allowed_roles` is a write pipe's only gate: the [policy engine](/access-control)'s insert rules do not apply to it, so any role you list — including a [`default_role`](/access-control#default_role--public-unauthenticated-access) that anonymous callers resolve to — can run the write. The operator fixes the statement and its predicate when authoring the pipe; callers supply only literal values. +`allowed_roles` is a write pipe's only gate: the [policy engine](/access-control)'s insert rules do not apply to it, so any role you list — including a [`default_role`](/access-control#default_role--public-unauthenticated-access) that anonymous callers resolve to — can run the write. The operator fixes the statement and its predicate when authoring the pipe; callers supply only literal values, provided every placeholder is written bare ([how a value becomes SQL](#how-a-value-becomes-sql)). Two things a write pipe does not do yet: it does not invalidate cached reads of the table it writes — a structured query or read pipe over that table can serve pre-write rows until its TTL ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)) — and its rows do not reach [`/v1/stream`](/api#get-v1stream--server-sent-events-stream) subscribers, which only the [ingest pipeline](/ingest-pipeline) feeds ([#362](https://github.com/Wave-RF/WaveHouse/issues/362)). For writes that should be seen at once, use [`POST /v1/ingest`](/api#post-v1ingesttabletable--ingest-data). diff --git a/internal/api/clickhouse_exec.go b/internal/api/clickhouse_exec.go index 6a296fc70..52a6228d6 100644 --- a/internal/api/clickhouse_exec.go +++ b/internal/api/clickhouse_exec.go @@ -6,6 +6,7 @@ import ( "reflect" "strings" "time" + "unicode/utf8" "github.com/ClickHouse/clickhouse-go/v2/lib/driver" "github.com/Wave-RF/WaveHouse/internal/settings" @@ -115,15 +116,14 @@ var mutationVerbs = map[string]struct{}{ // isMutation reports whether sql's leading statement is a non-SELECT — i.e. // one that returns no result set and must go through Exec, not Query. -// Leading whitespace and SQL line/block comments are skipped, then the first -// alphabetic token is matched case-insensitively against mutationVerbs. A -// leading WITH clause (CTE) routes through a paren-aware scan because -// ClickHouse accepts `WITH cte AS (...) INSERT INTO t SELECT * FROM cte` as -// equivalent to `INSERT INTO t WITH cte AS (...) SELECT * FROM cte` (see -// https://clickhouse.com/docs/sql-reference/statements/insert-into). Without -// the skip, the WITH form would classify as a read, route through Query, -// silently succeed, and return `[]` — the same silent-success class that -// motivated the original cache-bypass guard. +// Leading whitespace and comments are skipped as ClickHouse's lexer skips +// them, then the first alphabetic token is matched case-insensitively against +// mutationVerbs. A leading WITH clause (CTE) routes through a paren-aware scan +// because ClickHouse accepts `WITH cte AS (...) INSERT INTO t SELECT * FROM +// cte` as equivalent to `INSERT INTO t WITH cte AS (...) SELECT * FROM cte` +// (see https://clickhouse.com/docs/sql-reference/statements/insert-into). +// A write classified as a read goes through Query, which runs it and then +// fails the call, so a client that retries the error writes again. func isMutation(sql string) bool { s := stripLeadingSQLComments(sql) end := 0 @@ -163,9 +163,9 @@ var nonMutationVerbs = map[string]struct{}{ } // containsMutationVerbAtTopLevel scans s for the statement-introducing -// keyword at paren-depth 0, stepping over SQL string literals (`'…'` with -// `”` escape), quoted identifiers (`"…"` and “ `…` “), parenthesized CTE -// subqueries, and SQL comments. The CTE list contains ordinary identifiers +// keyword at paren-depth 0, stepping over string literals and quoted +// identifiers (skipQuoted), parenthesized CTE subqueries, and comments +// (skipComment). The CTE list contains ordinary identifiers // (CTE names, table/database names) that must not be matched as mutation // verbs — `system` would otherwise pattern-match `SYSTEM` and route a // `WITH … SELECT * FROM system.tables` read through `Exec` (silent empty- @@ -201,28 +201,8 @@ func containsMutationVerbAtTopLevel(s string) bool { depth-- } i++ - case c == '\'': - i++ - for i < len(s) { - if s[i] == '\'' { - if i+1 < len(s) && s[i+1] == '\'' { - i += 2 - continue - } - i++ - break - } - i++ - } - case c == '"' || c == '`': - q := c - i++ - for i < len(s) && s[i] != q { - i++ - } - if i < len(s) { - i++ - } + case c == '\'' || c == '"' || c == '`': + i = skipQuoted(s, i) case (c >= 'A' && c <= 'Z') || (c >= 'a' && c <= 'z'): start := i for i < len(s) { @@ -256,19 +236,8 @@ func containsMutationVerbAtTopLevel(s string) bool { return true } } - case c == '-' && i+1 < len(s) && s[i+1] == '-', c == '#': - for i < len(s) && s[i] != '\n' { - i++ - } - case c == '/' && i+1 < len(s) && s[i+1] == '*': - i += 2 - for i+1 < len(s) { - if s[i] == '*' && s[i+1] == '/' { - i += 2 - break - } - i++ - } + case c == '-' && i+1 < len(s) && s[i+1] == '-', c == '#', c == '/' && i+1 < len(s) && s[i+1] == '*': + i = skipComment(s, i) default: i++ } @@ -279,73 +248,124 @@ func containsMutationVerbAtTopLevel(s string) bool { // isCTENameLookahead returns true if the next non-whitespace, non-comment // token at or after pos is `AS` (case-insensitive, word-boundary terminated) // or `(` — signaling that whatever identifier just ended at pos is a CTE -// definition name (with optional column list before AS). Walks past space / -// tab / newline / `--` line comments / `#` line comments / `/* … */` block -// comments. Returns false on EOF or any other token. +// definition name (with optional column list before AS). Returns false on EOF +// or any other token. func isCTENameLookahead(s string, pos int) bool { - i := pos + i := skipSpaceAndComments(s, pos) + if i >= len(s) { + return false + } + if s[i] == '(' { + return true + } + end := i + for end < len(s) { + c := s[end] + if (c < 'A' || c > 'Z') && (c < 'a' || c > 'z') && (c < '0' || c > '9') && c != '_' { + break + } + end++ + } + return strings.EqualFold(s[i:end], "AS") +} + +// stripLeadingSQLComments trims whitespace and comments from the front of +// sql, the way ClickHouse's lexer skips them before the first token. +func stripLeadingSQLComments(sql string) string { + return sql[skipSpaceAndComments(sql, 0):] +} + +// skipSpaceAndComments returns the index of the first byte at or after i that +// is neither whitespace nor inside a comment. +func skipSpaceAndComments(s string, i int) int { for i < len(s) { - c := s[i] - switch { - case c == ' ' || c == '\t' || c == '\r' || c == '\n': - i++ - case c == '-' && i+1 < len(s) && s[i+1] == '-', c == '#': - for i < len(s) && s[i] != '\n' { - i++ - } - case c == '/' && i+1 < len(s) && s[i+1] == '*': - i += 2 - for i+1 < len(s) { - if s[i] == '*' && s[i+1] == '/' { - i += 2 - break - } - i++ - } - case c == '(': - return true - case (c >= 'A' && c <= 'Z') || (c >= 'a' && c <= 'z'): - end := i - for end < len(s) { - c2 := s[end] - if (c2 < 'A' || c2 > 'Z') && (c2 < 'a' || c2 > 'z') && (c2 < '0' || c2 > '9') && c2 != '_' { - break + if n := sqlSpaceLen(s, i); n > 0 { + i += n + continue + } + j := skipComment(s, i) + if j == i { + return i + } + i = j + } + return i +} + +// sqlSpaceLen is the byte length of the whitespace character at s[i], or 0. +// The set is ClickHouse's lexer's: ASCII space, \t \n \v \f \r, and the +// Unicode spaces it skips so that SQL pasted from a word processor parses. A +// leading one the classifier did not skip would hide the verb behind it. +func sqlSpaceLen(s string, i int) int { + switch s[i] { + case ' ', '\t', '\n', '\v', '\f', '\r': + return 1 + } + if s[i] < utf8.RuneSelf { + return 0 + } + r, n := utf8.DecodeRuneInString(s[i:]) + switch { + case r == 0x85, r == 0xA0, r == 0x180E, r >= 0x2000 && r <= 0x200D, + r == 0x2028, r == 0x2029, r == 0x202F, r == 0x205F, r == 0x2060, + r == 0x3000, r == 0xFEFF: + return n + } + return 0 +} + +// skipComment returns the index just past the comment starting at s[i], or i +// if none starts there: `--` and MySQL-compat `#` to end of line, `/* … */` +// nesting as ClickHouse's do. An unclosed block comment runs to the end, as +// it does for ClickHouse, which then rejects the statement. +func skipComment(s string, i int) int { + switch { + case strings.HasPrefix(s[i:], "--"), s[i] == '#': + if j := strings.IndexByte(s[i:], '\n'); j >= 0 { + return i + j + 1 + } + return len(s) + case strings.HasPrefix(s[i:], "/*"): + depth := 0 + for j := i; j+1 < len(s); { + switch { + case s[j] == '/' && s[j+1] == '*': + depth++ + j += 2 + case s[j] == '*' && s[j+1] == '/': + depth-- + j += 2 + if depth == 0 { + return j } - end++ + default: + j++ } - return strings.EqualFold(s[i:end], "AS") - default: - return false } + return len(s) } - return false + return i } -// stripLeadingSQLComments trims whitespace plus line comments (`-- …` and -// MySQL-compat `# …`, both accepted by ClickHouse) and `/* block */` -// comments from the front of sql, returning the remainder with no leading -// whitespace. Unclosed block comments swallow the rest of the string — -// matches what ClickHouse itself would do at parse time. -func stripLeadingSQLComments(sql string) string { - s := strings.TrimLeft(sql, " \t\r\n") - for { - switch { - case strings.HasPrefix(s, "--"), strings.HasPrefix(s, "#"): - if i := strings.IndexByte(s, '\n'); i >= 0 { - s = strings.TrimLeft(s[i+1:], " \t\r\n") - } else { - return "" - } - case strings.HasPrefix(s, "/*"): - if i := strings.Index(s[2:], "*/"); i >= 0 { - s = strings.TrimLeft(s[2+i+2:], " \t\r\n") - } else { - return "" +// skipQuoted returns the index just past the string literal or quoted +// identifier opening at s[i] (`'`, `"` or backtick). As in ClickHouse's lexer, +// a doubled quote or a backslash escapes the next byte; an unclosed one runs +// to the end. +func skipQuoted(s string, i int) int { + q := s[i] + for i++; i < len(s); i++ { + switch s[i] { + case '\\': + i++ + case q: + if i+1 < len(s) && s[i+1] == q { + i++ + continue } - default: - return s + return i + 1 } } + return len(s) } // transformRow converts ClickHouse-specific types to JSON-friendly values. diff --git a/internal/api/clickhouse_exec_test.go b/internal/api/clickhouse_exec_test.go index 947413b70..76441b6cf 100644 --- a/internal/api/clickhouse_exec_test.go +++ b/internal/api/clickhouse_exec_test.go @@ -110,6 +110,33 @@ func TestIsMutation(t *testing.T) { {"with parenthesized SELECT then system table (CTE-lookahead ordering regression)", "WITH x AS (SELECT 1) SELECT (1) FROM system.tables", false}, {"with tuple-shape SELECT then system table", "WITH x AS (SELECT 1) SELECT (a, b) FROM system.parts", false}, + // A backslash escapes the next byte inside all three quote kinds, so an + // escaped quote does not end the literal or identifier. + {"with backslash-escaped quote in literal then insert", `WITH m AS (SELECT 'it\'s' AS s) INSERT INTO t SELECT s FROM m`, true}, + {"with backslash-escaped quote in literal then select", `WITH m AS (SELECT 'a\'b' AS s) SELECT 'x) INSERT' FROM m`, false}, + {"with backslash-escaped double quote then insert", `WITH m AS (SELECT 'x' AS "a\"(b") INSERT INTO t SELECT * FROM m`, true}, + {"with backslash-escaped double quote then select", `WITH m AS (SELECT 1 AS "a\"b") SELECT 2 AS "x) INSERT" FROM m`, false}, + {"with backslash-escaped backtick then insert", "WITH m AS (SELECT 'x' AS `a\\`(b`) INSERT INTO t SELECT * FROM m", true}, + {"with backslash-escaped backtick then select", "WITH m AS (SELECT 1 AS `a\\`b`) SELECT 2 AS `x) INSERT` FROM m", false}, + + // ClickHouse's lexer skips \v, \f and Unicode spaces as whitespace + // (TestIsMutation_ClickHouseWhitespace covers the whole set). + {"leading form feed then insert", "\fINSERT INTO t VALUES (1)", true}, + {"leading vertical tab then insert", "\vINSERT INTO t VALUES (1)", true}, + {"leading NBSP then insert", "\u00a0INSERT INTO t VALUES (1)", true}, + {"leading BOM then insert", "\ufeffINSERT INTO t VALUES (1)", true}, + {"leading NBSP then select", "\u00a0SELECT 1", false}, + {"with CTE alias named set before form feed AS (read)", "WITH set\fAS (SELECT 1) SELECT * FROM set", false}, + {"with CTE alias named set before NBSP AS (read)", "WITH set\u00a0AS (SELECT 1) SELECT * FROM set", false}, + + // ClickHouse block comments nest. + {"nested block comment hiding select then insert", "/* a /* b */ SELECT */ INSERT INTO t VALUES (1)", true}, + {"nested block comment hiding insert then select", "/* a /* b */ INSERT */ SELECT 1", false}, + {"unclosed nested block comment", "/* a /* b */ INSERT INTO t VALUES (1)", false}, + {"with nested block comment hiding select then insert", "WITH x AS (SELECT 1) /* a /* b */ SELECT */ INSERT INTO t SELECT * FROM x", true}, + {"with nested block comment hiding insert then select", "WITH x AS (SELECT 1) /* a /* b */ INSERT */ SELECT * FROM x", false}, + {"with nested block comment before CTE AS (read)", "WITH set /* a /* b */ c */ AS (SELECT 1) SELECT * FROM set", false}, + {"empty", "", false}, {"comment only", "-- just a comment", false}, {"unclosed block comment", "/* never closed", false}, @@ -122,6 +149,26 @@ func TestIsMutation(t *testing.T) { } } +// TestIsMutation_ClickHouseWhitespace pins every character ClickHouse 26.6's +// lexer accepts as whitespace, each checked against a live server: ahead of a +// write it must not hide the verb, and ahead of a read it must not make one. +func TestIsMutation_ClickHouseWhitespace(t *testing.T) { + t.Parallel() + spaces := []rune{' ', '\t', '\n', '\v', '\f', '\r', 0x85, 0xA0, 0x180E, 0x2028, 0x2029, 0x202F, 0x205F, 0x2060, 0x3000, 0xFEFF} + for r := rune(0x2000); r <= 0x200D; r++ { + spaces = append(spaces, r) + } + for _, r := range spaces { + ws := string(r) + assert.True(t, isMutation(ws+"INSERT INTO t VALUES (1)"), "U+%04X before INSERT", r) + assert.False(t, isMutation(ws+"SELECT 1"), "U+%04X before SELECT", r) + assert.True(t, isMutation("WITH x AS (SELECT 1)"+ws+"INSERT INTO t SELECT * FROM x"), "U+%04X before a WITH's INSERT", r) + assert.False(t, isMutation("WITH set"+ws+"AS (SELECT 1) SELECT * FROM set"), "U+%04X between a CTE name and AS", r) + } + // Not whitespace to ClickHouse (it rejects the statement), so not skipped. + assert.False(t, isMutation("\u1680INSERT INTO t VALUES (1)")) +} + func TestExecuteCHQuery_MutationRoutesToExec(t *testing.T) { t.Parallel() // Mutations route through driver.Exec because clickhouse-go's diff --git a/internal/pipes/pipes.go b/internal/pipes/pipes.go index e146a2012..28e8b44db 100644 --- a/internal/pipes/pipes.go +++ b/internal/pipes/pipes.go @@ -143,7 +143,9 @@ func BindParams(q *NamedQuery, supplied map[string]any) (string, []any, error) { // comma-separated list of recursively formatted elements — the `(v1, v2, …)` // shape ClickHouse expects on the right of `IN`, matching how the // structured-query builder renders an IN clause. Because every scalar leaf is -// escaped, no value — or array element — can break out of its literal. +// escaped, no value — or array element — can break out of its literal, as long +// as the template writes the placeholder bare: inside quotes (`'{{id}}'`) the +// value's own quotes close the template's. // // Values with no scalar SQL representation are refused rather than emitted as // Go's `%v` text: a JSON object has no meaning here, and an empty array would From 9a4888a4959b8523f8c7907947928a2384068b66 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 04:58:48 -0400 Subject: [PATCH 33/79] fix(config): cache.redis defaults in defaults(); 0 never compresses Adopt #632's convention for the cache.redis block: its seven non-zero defaults move from env-default tags into defaults(), each gets a zero case, and the configuration.mdx rows use *(empty)* for an empty default. The doc-defaults test learns to read a duration and an empty list. With the file's zeros now kept, compress_min_bytes needs no -1: 0 means never compress, the backend's own meaning, from the file and the environment alike, and wire.go passes it through unchanged. A negative value refuses boot. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/configuration.mdx | 20 ++++++++++---------- internal/app/app_test.go | 15 +++++++++------ internal/app/wire.go | 6 +----- internal/config/cache_redis.go | 23 +++++++++++------------ internal/config/cache_redis_test.go | 19 ++++++++----------- internal/config/config.go | 14 +++++++++++--- internal/config/defaults_test.go | 14 ++++++++++++++ 9 files changed, 66 insertions(+), 49 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 6f742cb0c..7d0e6ea2a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone`, `cluster`, `sentinel`), `sentinel_master`, `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `-1` never compresses, since the loader reads a `0` in the file as unset) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. A malformed block — no address, an address without a port, an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. +- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone`, `cluster`, `sentinel`), `sentinel_master`, `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. A malformed block — no address, an address without a port, an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. - **A Redis-compatible shared cache backend** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row open a circuit breaker that skips the server until a probe succeeds, and deferred invalidations are retried until they land, collapsing to one tenant-wide bump per tenant past 100,000 keys. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. `cache.backend: redis` selects it (the entry above). Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression and server-stops-answering cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index fbe6324f4..2aba0526f 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -126,7 +126,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — the implementation each layer runs on, the process's `roles`, resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `cache.redis.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). - **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load` when `Config.NeedsDataDir` says a selected backend keeps state there, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. - **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` in `Validate`, between `validateRoles` and `validateTopology`), which refuses a value not on the list and names the ones that are. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns what a valid configuration is still likely to get wrong — the combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), a `cache.redis` block that is not read, certificate verification turned off — which `app.New` logs at `WARN`. -- **cache_redis.go** — `CacheRedisConfig`, the `cache.redis` sub-block, and its checks: an address, each `host:port`, a known mode, a sentinel's master name, `db` 0 in cluster mode, positive timeouts and sizes, a `version_ttl` of at least 2 s, and `compress_min_bytes` positive or `-1` (never: cleanenv reads a `0` in the file as unset and applies the default, so `0` cannot mean off). `CacheRedisTLS.Config` builds the `tls.Config`, reading the files; `Validate` calls it so an unreadable file refuses boot, and `internal/app` calls it again to build the connection. A TLS key set while `tls.enabled` is off is an error rather than a plaintext connection. +- **cache_redis.go** — `CacheRedisConfig`, the `cache.redis` sub-block, and its checks: an address, each `host:port`, a known mode, a sentinel's master name, `db` 0 in cluster mode, positive timeouts and sizes, a `version_ttl` of at least 2 s, and a `compress_min_bytes` that is not negative (`0` never compresses); its defaults are in `defaults()` with the rest. `CacheRedisTLS.Config` builds the `tls.Config`, reading the files; `Validate` calls it so an unreadable file refuses boot, and `internal/app` calls it again to build the connection. A TLS key set while `tls.enabled` is off is an error rather than a plaintext connection. - **config.go**, roles — `roles` (`[]Role`: `api`, `ingest`, `sweeper`; `AllRoles` by default; `Has(Role)`) picks which components `internal/app` wires, and `instance_id` names the process (`-<8 hex>` when empty, resolved in `Load`; today only logged at boot, and a distributed coordinator will record it as a lease's holder). `validateRoles` refuses an empty list, an empty entry, an unknown or a repeated role; `validateTopology` refuses a role set the backends cannot serve: any split over the embedded MQ, and a process with exactly one of `api` and `ingest` over a local cache. `NeedsDataDir` counts Pebble only for a process running `api`, and `Warnings` is empty without `api`, since only that role opens a cache it reads or a dedupe store. - **strict.go** — `rejectUnknownKeys`, the YAML half: re-reads the file as a generic tree and walks it against the struct's `yaml` tags, listing every key the struct doesn't declare. cleanenv itself is lenient by design, which is exactly wrong for boot config once keys have moved to the settings directory. - **persistence.go** — `WarnIfFreshDataDir` logs the startup `WARN` when `data_dir` is missing or empty (on a redeploy, the sign that the volume didn't persist); `LogStorageInitError` attaches the UID-65532 `permissionHint` to a NATS or Pebble open failure that looks like a permission denial — the same hint string `CheckDataDir` uses. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 91aa72d6b..0b5348549 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -150,23 +150,23 @@ The `redis` backend's settings, read only when `cache.backend` is `redis`. It is | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | -| `cache.redis.addrs` | `WH_CACHE_REDIS_ADDRS` | — | **Required** with `backend: redis`. `host:port` of the server; several are a cluster's seed nodes or the sentinels. Comma-separated in the env var. | +| `cache.redis.addrs` | `WH_CACHE_REDIS_ADDRS` | *(empty)* | **Required** with `backend: redis`. `host:port` of the server; several are a cluster's seed nodes or the sentinels. Comma-separated in the env var. | | `cache.redis.mode` | `WH_CACHE_REDIS_MODE` | `standalone` | `standalone`, `cluster` or `sentinel`. | -| `cache.redis.sentinel_master` | `WH_CACHE_REDIS_SENTINEL_MASTER` | — | The master set name. Required with `mode: sentinel`. | -| `cache.redis.username` | `WH_CACHE_REDIS_USERNAME` | — | ACL user. Empty uses the server's `default` user. | -| `cache.redis.password` | `WH_CACHE_REDIS_PASSWORD` | — | A secret: set it through the environment (or your secret store's env injection), not in a tracked `config.yaml`. | +| `cache.redis.sentinel_master` | `WH_CACHE_REDIS_SENTINEL_MASTER` | *(empty)* | The master set name. Required with `mode: sentinel`. | +| `cache.redis.username` | `WH_CACHE_REDIS_USERNAME` | *(empty)* | ACL user. Empty uses the server's `default` user. | +| `cache.redis.password` | `WH_CACHE_REDIS_PASSWORD` | *(empty)* | A secret: set it through the environment (or your secret store's env injection), not in a tracked `config.yaml`. | | `cache.redis.db` | `WH_CACHE_REDIS_DB` | `0` | Database number (`SELECT`). Standalone and sentinel only: a cluster has only database `0`, and any other value refuses boot. | | `cache.redis.tls.enabled` | `WH_CACHE_REDIS_TLS_ENABLED` | `false` | Connect over TLS, verifying the server against the system roots or `ca_file`. Any other `tls` key set while this is off refuses boot, rather than connecting in plaintext. | -| `cache.redis.tls.ca_file` | `WH_CACHE_REDIS_TLS_CA_FILE` | — | PEM file of the authorities to trust instead of the system roots. | -| `cache.redis.tls.cert_file` | `WH_CACHE_REDIS_TLS_CERT_FILE` | — | Client certificate (PEM) for mutual TLS. Set together with `key_file`. | -| `cache.redis.tls.key_file` | `WH_CACHE_REDIS_TLS_KEY_FILE` | — | The client certificate's private key (PEM). | -| `cache.redis.tls.server_name` | `WH_CACHE_REDIS_TLS_SERVER_NAME` | — | Name to verify the server's certificate against, when it differs from the address. | +| `cache.redis.tls.ca_file` | `WH_CACHE_REDIS_TLS_CA_FILE` | *(empty)* | PEM file of the authorities to trust instead of the system roots. | +| `cache.redis.tls.cert_file` | `WH_CACHE_REDIS_TLS_CERT_FILE` | *(empty)* | Client certificate (PEM) for mutual TLS. Set together with `key_file`. | +| `cache.redis.tls.key_file` | `WH_CACHE_REDIS_TLS_KEY_FILE` | *(empty)* | The client certificate's private key (PEM). | +| `cache.redis.tls.server_name` | `WH_CACHE_REDIS_TLS_SERVER_NAME` | *(empty)* | Name to verify the server's certificate against, when it differs from the address. | | `cache.redis.tls.insecure_skip_verify` | `WH_CACHE_REDIS_TLS_INSECURE_SKIP_VERIFY` | `false` | Accept any server certificate. Logged at `WARN` at boot: whoever can intercept the connection can read and replace cached results. | | `cache.redis.key_prefix` | `WH_CACHE_REDIS_KEY_PREFIX` | `wh` | Leads every key, so several deployments can share one server, provided you trust each as much as the others: any of them can overwrite what the rest serve. No `{` or `}`. | | `cache.redis.timeout` | `WH_CACHE_REDIS_TIMEOUT` | `100ms` | Per operation. A lookup or fill that takes longer is a miss or a skipped fill, never a failed query. | | `cache.redis.dial_timeout` | `WH_CACHE_REDIS_DIAL_TIMEOUT` | `1s` | Per connection attempt. | | `cache.redis.max_value_bytes` | `WH_CACHE_REDIS_MAX_VALUE_BYTES` | `1048576` | Largest result stored, after compression (1 MiB). A larger one is returned to the caller but not cached. | -| `cache.redis.compress_min_bytes` | `WH_CACHE_REDIS_COMPRESS_MIN_BYTES` | `1024` | Results at least this large are zstd-compressed when that makes them smaller. `-1` never compresses. `0` is not "off": in the YAML file it reads as unset and takes the default, and in the env var it refuses boot. | +| `cache.redis.compress_min_bytes` | `WH_CACHE_REDIS_COMPRESS_MIN_BYTES` | `1024` | Results at least this large are zstd-compressed when that makes them smaller. `0` never compresses. | | `cache.redis.version_ttl` | `WH_CACHE_REDIS_VERSION_TTL` | `168h` | How long a table's or tenant's version token outlives its last write, so the tokens of dropped tables and removed tenants eventually expire. At least `2s`. An expired token only causes misses. | **When the server is unreachable or misbehaves, the cache is bypassed; queries are not.** A failure or a timeout makes the lookup a miss and the fill a no-op, and five in a row open a circuit breaker that skips the server entirely until a probe, every 5 s, gets an answer. Queries then go straight to ClickHouse, still coalesced per instance by `singleflight`. An invalidation the server did not take is kept and retried until it lands. `/readyz` does not depend on the cache. @@ -284,7 +284,7 @@ cache: timeout: 100ms dial_timeout: 1s max_value_bytes: 1048576 - compress_min_bytes: 1024 # -1 = never + compress_min_bytes: 1024 # 0 = never version_ttl: 168h dedupe: diff --git a/internal/app/app_test.go b/internal/app/app_test.go index 2ca4a205c..204c13711 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -796,9 +796,10 @@ func TestNew_RedisCacheRefusesAnUnreadableTLSFile(t *testing.T) { require.ErrorContains(t, err, "cache init: cache.redis.tls.ca_file") } -// The boot config's defaults are the backend's, and -1 is the backend's -// "never compress". Driven from Load, not a literal, so a default changed on -// one side only fails here. +// The boot config's defaults are the backend's, and a compress_min_bytes of +// 0 reaches the backend as its "never compress" rather than its default. +// Driven from Load, not a literal, so a default changed on one side only +// fails here. func TestRedisConfig_FromLoadedDefaults(t *testing.T) { t.Setenv("WH_SETTINGS_DIR", t.TempDir()) t.Setenv("WH_CACHE_BACKEND", "redis") @@ -815,11 +816,13 @@ func TestRedisConfig_FromLoadedDefaults(t *testing.T) { CompressMinBytes: cache.DefaultRedisCompressMinBytes, VersionTTL: cache.DefaultRedisVersionTTL, }, got) - loaded.Cache.Redis.CompressMinBytes = -1 - loaded.Cache.Redis.Mode = config.RedisCluster + t.Setenv("WH_CACHE_REDIS_COMPRESS_MIN_BYTES", "0") + t.Setenv("WH_CACHE_REDIS_MODE", "cluster") + loaded, err = config.Load(filepath.Join(t.TempDir(), "none.yaml")) + require.NoError(t, err) got, err = redisConfig(loaded.Cache.Redis) require.NoError(t, err) - assert.Zero(t, got.CompressMinBytes) + assert.Zero(t, got.CompressMinBytes, "the backend's never, not its default") assert.Equal(t, cache.RedisCluster, got.Mode) assert.Equal(t, cache.RedisSentinel, config.RedisSentinel) } diff --git a/internal/app/wire.go b/internal/app/wire.go index 60c9e4a9e..ea23606d0 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -664,10 +664,6 @@ func redisConfig(r config.CacheRedisConfig) (cache.RedisConfig, error) { if err != nil { return cache.RedisConfig{}, err } - compressMin := r.CompressMinBytes - if compressMin < 0 { - compressMin = 0 // the backend's "never" - } return cache.RedisConfig{ Addrs: r.Addrs, Mode: r.Mode, @@ -680,7 +676,7 @@ func redisConfig(r config.CacheRedisConfig) (cache.RedisConfig, error) { Timeout: r.Timeout, DialTimeout: r.DialTimeout, MaxValueBytes: r.MaxValueBytes, - CompressMinBytes: compressMin, + CompressMinBytes: r.CompressMinBytes, VersionTTL: r.VersionTTL, }, nil } diff --git a/internal/config/cache_redis.go b/internal/config/cache_redis.go index 3ae3114ea..997d19a3a 100644 --- a/internal/config/cache_redis.go +++ b/internal/config/cache_redis.go @@ -25,24 +25,23 @@ type CacheRedisConfig struct { // Addrs are host:port pairs: the server, or seeds for a cluster, or the // sentinels. Addrs []string `yaml:"addrs" env:"WH_CACHE_REDIS_ADDRS"` - Mode string `yaml:"mode" env:"WH_CACHE_REDIS_MODE" env-default:"standalone"` + Mode string `yaml:"mode" env:"WH_CACHE_REDIS_MODE"` SentinelMaster string `yaml:"sentinel_master" env:"WH_CACHE_REDIS_SENTINEL_MASTER"` Username string `yaml:"username" env:"WH_CACHE_REDIS_USERNAME"` Password string `yaml:"password" env:"WH_CACHE_REDIS_PASSWORD"` DB int `yaml:"db" env:"WH_CACHE_REDIS_DB"` TLS CacheRedisTLS `yaml:"tls"` // KeyPrefix leads every key, so deployments can share one server. - KeyPrefix string `yaml:"key_prefix" env:"WH_CACHE_REDIS_KEY_PREFIX" env-default:"wh"` - Timeout time.Duration `yaml:"timeout" env:"WH_CACHE_REDIS_TIMEOUT" env-default:"100ms"` - DialTimeout time.Duration `yaml:"dial_timeout" env:"WH_CACHE_REDIS_DIAL_TIMEOUT" env-default:"1s"` + KeyPrefix string `yaml:"key_prefix" env:"WH_CACHE_REDIS_KEY_PREFIX"` + Timeout time.Duration `yaml:"timeout" env:"WH_CACHE_REDIS_TIMEOUT"` + DialTimeout time.Duration `yaml:"dial_timeout" env:"WH_CACHE_REDIS_DIAL_TIMEOUT"` // MaxValueBytes is the largest value stored, after compression. - MaxValueBytes int `yaml:"max_value_bytes" env:"WH_CACHE_REDIS_MAX_VALUE_BYTES" env-default:"1048576"` - // CompressMinBytes is the smallest value zstd-compressed; -1 never - // compresses. Not 0: the loader reads a 0 in the file as unset and - // applies the default, so 0 cannot mean off. - CompressMinBytes int `yaml:"compress_min_bytes" env:"WH_CACHE_REDIS_COMPRESS_MIN_BYTES" env-default:"1024"` + MaxValueBytes int `yaml:"max_value_bytes" env:"WH_CACHE_REDIS_MAX_VALUE_BYTES"` + // CompressMinBytes is the smallest value zstd-compressed; 0 never + // compresses. + CompressMinBytes int `yaml:"compress_min_bytes" env:"WH_CACHE_REDIS_COMPRESS_MIN_BYTES"` // VersionTTL is how long a version token outlives its last bump. - VersionTTL time.Duration `yaml:"version_ttl" env:"WH_CACHE_REDIS_VERSION_TTL" env-default:"168h"` + VersionTTL time.Duration `yaml:"version_ttl" env:"WH_CACHE_REDIS_VERSION_TTL"` } // CacheRedisTLS is cache.redis.tls. The files are paths, read at boot. @@ -106,8 +105,8 @@ func (r CacheRedisConfig) validate() error { if r.MaxValueBytes <= 0 { return fmt.Errorf("cache.redis.max_value_bytes (WH_CACHE_REDIS_MAX_VALUE_BYTES) %d must be positive", r.MaxValueBytes) } - if r.CompressMinBytes == 0 || r.CompressMinBytes < -1 { - return fmt.Errorf("cache.redis.compress_min_bytes (WH_CACHE_REDIS_COMPRESS_MIN_BYTES) %d: want a positive size, or -1 to never compress", r.CompressMinBytes) + if r.CompressMinBytes < 0 { + return fmt.Errorf("cache.redis.compress_min_bytes (WH_CACHE_REDIS_COMPRESS_MIN_BYTES) %d is negative: want a size, or 0 to never compress", r.CompressMinBytes) } if _, err := r.TLS.Config(); err != nil { return err diff --git a/internal/config/cache_redis_test.go b/internal/config/cache_redis_test.go index e1b5783a8..5389715a9 100644 --- a/internal/config/cache_redis_test.go +++ b/internal/config/cache_redis_test.go @@ -60,7 +60,7 @@ func TestLoad_CacheRedisFromEnv(t *testing.T) { "WH_CACHE_REDIS_TIMEOUT": "250ms", "WH_CACHE_REDIS_DIAL_TIMEOUT": "3s", "WH_CACHE_REDIS_MAX_VALUE_BYTES": "2048", - "WH_CACHE_REDIS_COMPRESS_MIN_BYTES": "-1", + "WH_CACHE_REDIS_COMPRESS_MIN_BYTES": "0", "WH_CACHE_REDIS_VERSION_TTL": "24h", "WH_CACHE_REDIS_TLS_INSECURE_SKIP_VERIFY": "false", } { @@ -75,7 +75,7 @@ func TestLoad_CacheRedisFromEnv(t *testing.T) { Enabled: true, CAFile: caFile, CertFile: certFile, KeyFile: keyFile, ServerName: "redis.internal", }, KeyPrefix: "staging", Timeout: 250 * time.Millisecond, DialTimeout: 3 * time.Second, - MaxValueBytes: 2048, CompressMinBytes: -1, VersionTTL: 24 * time.Hour, + MaxValueBytes: 2048, CompressMinBytes: 0, VersionTTL: 24 * time.Hour, }, cfg.Cache.Redis) tc, err := cfg.Cache.Redis.TLS.Config() require.NoError(t, err) @@ -112,10 +112,8 @@ cache: assert.Equal(t, 1024, r.CompressMinBytes) } -// A 0 in the file is read as unset, so it takes the default instead of -// switching compression off: why "never" is -1. Pinned so a loader that -// starts honoring the 0 is noticed. -func TestLoad_CacheRedisCompressZeroInYAMLIsTheDefault(t *testing.T) { +// A 0 in the file is kept, and is the backend's "never compress". +func TestLoad_CacheRedisCompressZeroInYAMLIsNever(t *testing.T) { t.Parallel() path := filepath.Join(t.TempDir(), "config.yaml") require.NoError(t, os.WriteFile(path, []byte(` @@ -129,7 +127,7 @@ cache: `), 0o600)) cfg, err := Load(path) require.NoError(t, err) - assert.Equal(t, 1024, cfg.Cache.Redis.CompressMinBytes) + assert.Zero(t, cfg.Cache.Redis.CompressMinBytes) } func TestLoad_CacheRedisRefusesUnknownKeys(t *testing.T) { @@ -171,7 +169,7 @@ func TestUnboundEnv_KnowsTheCacheRedisVariables(t *testing.T) { t.Parallel() assert.Empty(t, unboundEnv([]string{ "WH_CACHE_REDIS_ADDRS=r:6379", "WH_CACHE_REDIS_PASSWORD=x", "WH_CACHE_REDIS_TLS_CA_FILE=/ca.pem", - "WH_CACHE_REDIS_VERSION_TTL=1h", "WH_CACHE_REDIS_COMPRESS_MIN_BYTES=-1", + "WH_CACHE_REDIS_VERSION_TTL=1h", "WH_CACHE_REDIS_COMPRESS_MIN_BYTES=0", })) assert.Equal(t, []string{"WH_CACHE_REDIS_ADDR"}, unboundEnv([]string{"WH_CACHE_REDIS_ADDR=r:6379"})) } @@ -203,9 +201,8 @@ func TestValidate_CacheRedis(t *testing.T) { {"negative dial timeout", func(r *CacheRedisConfig) { r.DialTimeout = -time.Second }, "cache.redis.dial_timeout"}, {"short version ttl", func(r *CacheRedisConfig) { r.VersionTTL = time.Second }, "cache.redis.version_ttl (WH_CACHE_REDIS_VERSION_TTL) 1s is under 2s"}, {"zero max value", func(r *CacheRedisConfig) { r.MaxValueBytes = 0 }, "cache.redis.max_value_bytes"}, - {"compress 0", func(r *CacheRedisConfig) { r.CompressMinBytes = 0 }, "or -1 to never compress"}, - {"compress -2", func(r *CacheRedisConfig) { r.CompressMinBytes = -2 }, "or -1 to never compress"}, - {"compress never", func(r *CacheRedisConfig) { r.CompressMinBytes = -1 }, ""}, + {"compress never", func(r *CacheRedisConfig) { r.CompressMinBytes = 0 }, ""}, + {"compress negative", func(r *CacheRedisConfig) { r.CompressMinBytes = -1 }, "-1 is negative: want a size, or 0 to never compress"}, {"tls files while off", func(r *CacheRedisConfig) { r.TLS.CAFile = caFile }, "cache.redis.tls.enabled (WH_CACHE_REDIS_TLS_ENABLED) is off"}, {"tls system roots", func(r *CacheRedisConfig) { r.TLS.Enabled = true }, ""}, {"tls full", func(r *CacheRedisConfig) { diff --git a/internal/config/config.go b/internal/config/config.go index 95bfacb15..27f4188cc 100644 --- a/internal/config/config.go +++ b/internal/config/config.go @@ -7,6 +7,7 @@ import ( "os" "slices" "strings" + "time" "github.com/ilyakaznacheev/cleanenv" ) @@ -255,9 +256,16 @@ func defaults() Config { Roles: AllRoles(), Server: Server{Port: 8080, ShutdownTimeout: 10}, MQ: MQ{Backend: MQEmbedded}, - Cache: Cache{Backend: CacheLocal, L1MaxCost: 64 << 20}, - Dedupe: Dedupe{Backend: DedupePebble}, - Coord: Coord{Backend: CoordLocal}, + Cache: Cache{ + Backend: CacheLocal, L1MaxCost: 64 << 20, + Redis: CacheRedisConfig{ + Mode: RedisStandalone, KeyPrefix: "wh", + Timeout: 100 * time.Millisecond, DialTimeout: time.Second, + MaxValueBytes: 1 << 20, CompressMinBytes: 1 << 10, VersionTTL: 168 * time.Hour, + }, + }, + Dedupe: Dedupe{Backend: DedupePebble}, + Coord: Coord{Backend: CoordLocal}, OTel: OTel{ Traces: OTelTraces{Enabled: true, SampleRate: 1.0}, Metrics: OTelMetrics{Enabled: true}, diff --git a/internal/config/defaults_test.go b/internal/config/defaults_test.go index 6877be4bb..1a8fba337 100644 --- a/internal/config/defaults_test.go +++ b/internal/config/defaults_test.go @@ -9,6 +9,7 @@ import ( "strconv" "strings" "testing" + "time" "github.com/stretchr/testify/assert" "github.com/stretchr/testify/require" @@ -37,6 +38,15 @@ var zeroCases = []zeroCase{ {"cache.l1_max_cost", "WH_CACHE_L1_MAX_COST", int64(0), int64(64 << 20), "1024", int64(1024), func(c *Config) any { return c.Cache.L1MaxCost }}, {"prometheus.path", "WH_PROMETHEUS_PATH", "", "/metrics", "/prom", "/prom", func(c *Config) any { return c.Prometheus.Path }}, {"data_dir", "WH_DATA_DIR", "", "./data", "/var/lib/wh", "/var/lib/wh", func(c *Config) any { return c.DataDir }}, + // The cache.redis block is validated only under backend=redis, so under + // the default backend its zeros load as written. + {"cache.redis.mode", "WH_CACHE_REDIS_MODE", "", RedisStandalone, RedisCluster, RedisCluster, func(c *Config) any { return c.Cache.Redis.Mode }}, + {"cache.redis.key_prefix", "WH_CACHE_REDIS_KEY_PREFIX", "", "wh", "staging", "staging", func(c *Config) any { return c.Cache.Redis.KeyPrefix }}, + {"cache.redis.timeout", "WH_CACHE_REDIS_TIMEOUT", time.Duration(0), 100 * time.Millisecond, "250ms", 250 * time.Millisecond, func(c *Config) any { return c.Cache.Redis.Timeout }}, + {"cache.redis.dial_timeout", "WH_CACHE_REDIS_DIAL_TIMEOUT", time.Duration(0), time.Second, "500ms", 500 * time.Millisecond, func(c *Config) any { return c.Cache.Redis.DialTimeout }}, + {"cache.redis.max_value_bytes", "WH_CACHE_REDIS_MAX_VALUE_BYTES", 0, 1 << 20, "2048", 2048, func(c *Config) any { return c.Cache.Redis.MaxValueBytes }}, + {"cache.redis.compress_min_bytes", "WH_CACHE_REDIS_COMPRESS_MIN_BYTES", 0, 1 << 10, "2048", 2048, func(c *Config) any { return c.Cache.Redis.CompressMinBytes }}, + {"cache.redis.version_ttl", "WH_CACHE_REDIS_VERSION_TTL", time.Duration(0), 168 * time.Hour, "1h", time.Hour, func(c *Config) any { return c.Cache.Redis.VersionTTL }}, } // refusedZeros are the non-zero defaults whose zero Validate refuses: written @@ -296,11 +306,15 @@ func parseDocDefault(t *testing.T, key, cell string, like any) any { v, err = strconv.ParseInt(cell, 10, 64) case float64: v, err = strconv.ParseFloat(cell, 64) + case time.Duration: + v, err = time.ParseDuration(cell) default: rt := reflect.TypeOf(like) switch { case rt.Kind() == reflect.String: v = reflect.ValueOf(cell).Convert(rt).Interface() + case rt.Kind() == reflect.Slice && rt.Elem().Kind() == reflect.String && cell == "": + v = reflect.Zero(rt).Interface() // cache.redis.addrs: nil case rt.Kind() == reflect.Slice && rt.Elem().Kind() == reflect.String: // roles: a comma-separated cell parts := strings.Split(cell, ",") sv := reflect.MakeSlice(rt, len(parts), len(parts)) From 6fa1ae6eb949998c548c34bdc6e1cc623402029c Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:00:52 -0400 Subject: [PATCH 34/79] fix(cache): give the two-instance test its roles; redis is the shared cache app.New now wires the API, ingest and cache only for the roles a config names (#622), and refuses a Config with none, so the shared-cache integration test's instances run every role, as the suite's own app does. The refusal of api without ingest over a local cache now names cache.backend=redis, the shared cache it asks for, and the configuration and deployment pages say this build has a shared cache but still refuses every split while the queue and leases are in-process. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- CHANGELOG.md | 2 +- docs/src/content/docs/configuration.mdx | 2 +- docs/src/content/docs/deployment.md | 2 +- internal/config/config.go | 2 +- internal/config/roles_test.go | 4 +++- tests/integration/shared_cache_test.go | 1 + 6 files changed, 8 insertions(+), 5 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 7d0e6ea2a..f8f2efab6 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone`, `cluster`, `sentinel`), `sentinel_master`, `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. A malformed block — no address, an address without a port, an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. +- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone`, `cluster`, `sentinel`), `sentinel_master`, `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. - **A Redis-compatible shared cache backend** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row open a circuit breaker that skips the server until a probe succeeds, and deferred invalidations are retried until they land, collapsing to one tenant-wide bump per tenant past 100,000 keys. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. `cache.backend: redis` selects it (the entry above). Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression and server-stops-answering cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 0b5348549..6425d6f77 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -70,7 +70,7 @@ Every process, whatever its roles, reads the settings directory and reloads it ( Boot refuses a role set the selected backends cannot serve: - **Any split with `mq.backend=embedded`.** The embedded queue lives inside its process and listens on no port, so a process without every role could not reach it. Until a shared `mq.backend` exists, every process runs every role. -- **`api` without `ingest`, or `ingest` without `api`, with `cache.backend=local`.** The ingest worker invalidates the cache the API reads, and a local cache in another process never sees that invalidation. Run `api` and `ingest` together, or choose a shared `cache.backend`. A `sweeper`-only process holds no cache, so this rule does not apply to it. +- **`api` without `ingest`, or `ingest` without `api`, with `cache.backend=local`.** The ingest worker invalidates the cache the API reads, and a local cache in another process never sees that invalidation. Run `api` and `ingest` together, or set [`cache.backend: redis`](#backends), one cache every process shares. A `sweeper`-only process holds no cache, so this rule does not apply to it. ### Server diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 32e508a68..2acf348c9 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -347,7 +347,7 @@ By default one process runs all of WaveHouse. [`roles`](/configuration#process-r - **Ingest.** Every ingest pod consumes the same shared durable consumer and competes for its messages, so throughput scales with the pod count. The rows of one table are then split across pods: each pod writes smaller batches, and rows written by different pods do not reach ClickHouse in publish order. - **Sweeper.** The sweeper runs under a lease held in the shared `coord.backend`, so only one pod sweeps at a time. A second replica waits and takes over when the first stops. -A split needs backends that every process can reach: a shared `mq.backend`, so that every process reaches the same queue; a shared `cache.backend`, so that the ingest pods' invalidations reach the API pods' cache; and a shared `coord.backend`, so that the sweeper lease spans pods. **This build has only the in-process backends, so boot refuses any split** and names the backend to change. Until shared backends ship, run every role in one process, the default. +A split needs backends that every process can reach: a shared `mq.backend`, so that every process reaches the same queue; a shared `cache.backend` ([`redis`](#multiple-instances-and-the-shared-cache)), so that the ingest pods' invalidations reach the API pods' cache; and a shared `coord.backend`, so that the sweeper lease spans pods. **This build has a shared cache but only the in-process queue and leases, so boot refuses any split** and names the backend to change. Until a shared `mq.backend` and `coord.backend` ship, run every role in one process, the default. A pod without the `api` role serves an ops listener on `:8080`: `/livez`, `/readyz` and their aliases, `/version`, the metrics path when `prometheus.port` is `0`, and `POST /v1/ops/settings/reload`. Every other route answers 404 (under `/v1/ops`, 403 without the operator key, and 401 for a bearer token). Point the same probes at it as at an API pod. `/livez` does not wait for schema discovery there, because only the API runs it. `/readyz` checks ClickHouse in an ingest pod, and is ready once a sweeper pod has booted. Every pod reads the settings directory, so mount it in every Deployment. The reload route on the ops listener accepts only the operator key, so whatever reloads your API pods over HTTP must send the operator key to the worker pods too, or rely on `SIGHUP` (or, over a flat directory, the directory watcher) instead. diff --git a/internal/config/config.go b/internal/config/config.go index 27f4188cc..0bc70f2ab 100644 --- a/internal/config/config.go +++ b/internal/config/config.go @@ -218,7 +218,7 @@ func (c *Config) validateTopology() error { return fmt.Errorf("roles %s with mq.backend=embedded: the embedded MQ lives inside this process, and a process without it cannot reach its queue — run every role (%s), or set a shared mq.backend", joinRoles(c.Roles), joinRoles(allRoles)) } if c.splitsCache() && c.Cache.Backend == CacheLocal { - return fmt.Errorf("roles %s with cache.backend=local: api and ingest run in different processes, and the ingest worker's cache invalidation would never reach the API's cache — run api and ingest together, or set a shared cache.backend", joinRoles(c.Roles)) + return fmt.Errorf("roles %s with cache.backend=local: api and ingest run in different processes, and the ingest worker's cache invalidation would never reach the API's cache — run api and ingest together, or set cache.backend=redis, one cache every process shares", joinRoles(c.Roles)) } return nil } diff --git a/internal/config/roles_test.go b/internal/config/roles_test.go index 18b28267c..914815bf9 100644 --- a/internal/config/roles_test.go +++ b/internal/config/roles_test.go @@ -114,12 +114,14 @@ func TestValidate_RoleSplits(t *testing.T) { {"every role, shared queue", all, "shared", CacheLocal, ""}, {"api+ingest, shared queue", []Role{RoleAPI, RoleIngest}, "shared", CacheLocal, ""}, {"sweeper, shared queue", []Role{RoleSweeper}, "shared", CacheLocal, ""}, - {"api, local cache", []Role{RoleAPI}, "shared", CacheLocal, "roles api with cache.backend=local: api and ingest run in different processes"}, + {"api, local cache", []Role{RoleAPI}, "shared", CacheLocal, "roles api with cache.backend=local: api and ingest run in different processes, and the ingest worker's cache invalidation would never reach the API's cache — run api and ingest together, or set cache.backend=redis"}, {"ingest, local cache", []Role{RoleIngest}, "shared", CacheLocal, "roles ingest with cache.backend=local"}, {"api+sweeper, local cache", []Role{RoleAPI, RoleSweeper}, "shared", CacheLocal, "roles api,sweeper with cache.backend=local"}, {"ingest+sweeper, local cache", []Role{RoleIngest, RoleSweeper}, "shared", CacheLocal, "roles ingest,sweeper with cache.backend=local"}, {"api, shared cache", []Role{RoleAPI}, "shared", "shared", ""}, {"ingest, shared cache", []Role{RoleIngest}, "shared", "shared", ""}, + {"api, redis cache", []Role{RoleAPI}, "shared", CacheRedis, ""}, + {"ingest, redis cache", []Role{RoleIngest}, "shared", CacheRedis, ""}, } { t.Run(tc.name, func(t *testing.T) { t.Parallel() diff --git a/tests/integration/shared_cache_test.go b/tests/integration/shared_cache_test.go index b665ccb67..f4a884d0e 100644 --- a/tests/integration/shared_cache_test.go +++ b/tests/integration/shared_cache_test.go @@ -85,6 +85,7 @@ func bootRedisApp(t *testing.T, redisAddr, prefix string, timeout time.Duration) }}, Dedupe: config.Dedupe{Backend: config.DedupePebble}, Coord: config.Coord{Backend: config.CoordLocal}, + Roles: config.AllRoles(), Settings: config.Settings{Dir: settingsDir}, } a, err := app.New(ctx, app.Options{Config: cfg, Listener: ln}) From 712b1db250585dd0dfed112e76754763cdc50c8e Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:02:57 -0400 Subject: [PATCH 35/79] fix(config): refuse cache.redis.mode=sentinel until #656 Sentinel mode passed only the master set name: the backend does not authenticate to the sentinels, sets no topology refresh, and has no test (#656). Rather than make it selectable, validation refuses it and points at the issue; standalone and cluster are unchanged. cache.redis.sentinel_master is removed rather than kept and refused: a key that no config can use is noise in the reference, and the strict loader now reports it (and WH_CACHE_REDIS_SENTINEL_MASTER) as unknown, so nobody sets it believing it works. #656 can add it back with the sentinel credentials it also needs. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/configuration.mdx | 11 ++++----- docs/src/content/docs/deployment.md | 2 +- internal/app/wire.go | 1 - internal/config/cache_redis.go | 30 ++++++++++++------------- internal/config/cache_redis_test.go | 17 +++++++------- 7 files changed, 30 insertions(+), 35 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index f8f2efab6..ac09db1ae 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone`, `cluster`, `sentinel`), `sentinel_master`, `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. +- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. - **A Redis-compatible shared cache backend** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row open a circuit breaker that skips the server until a probe succeeds, and deferred invalidations are retried until they land, collapsing to one tenant-wide bump per tenant past 100,000 keys. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. `cache.backend: redis` selects it (the entry above). Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression and server-stops-answering cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 2aba0526f..4c627ee7c 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -126,7 +126,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — the implementation each layer runs on, the process's `roles`, resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `cache.redis.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). - **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load` when `Config.NeedsDataDir` says a selected backend keeps state there, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. - **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` in `Validate`, between `validateRoles` and `validateTopology`), which refuses a value not on the list and names the ones that are. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns what a valid configuration is still likely to get wrong — the combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), a `cache.redis` block that is not read, certificate verification turned off — which `app.New` logs at `WARN`. -- **cache_redis.go** — `CacheRedisConfig`, the `cache.redis` sub-block, and its checks: an address, each `host:port`, a known mode, a sentinel's master name, `db` 0 in cluster mode, positive timeouts and sizes, a `version_ttl` of at least 2 s, and a `compress_min_bytes` that is not negative (`0` never compresses); its defaults are in `defaults()` with the rest. `CacheRedisTLS.Config` builds the `tls.Config`, reading the files; `Validate` calls it so an unreadable file refuses boot, and `internal/app` calls it again to build the connection. A TLS key set while `tls.enabled` is off is an error rather than a plaintext connection. +- **cache_redis.go** — `CacheRedisConfig`, the `cache.redis` sub-block, and its checks: an address, each `host:port`, a known mode (`sentinel` is refused until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `db` 0 in cluster mode, positive timeouts and sizes, a `version_ttl` of at least 2 s, and a `compress_min_bytes` that is not negative (`0` never compresses); its defaults are in `defaults()` with the rest. `CacheRedisTLS.Config` builds the `tls.Config`, reading the files; `Validate` calls it so an unreadable file refuses boot, and `internal/app` calls it again to build the connection. A TLS key set while `tls.enabled` is off is an error rather than a plaintext connection. - **config.go**, roles — `roles` (`[]Role`: `api`, `ingest`, `sweeper`; `AllRoles` by default; `Has(Role)`) picks which components `internal/app` wires, and `instance_id` names the process (`-<8 hex>` when empty, resolved in `Load`; today only logged at boot, and a distributed coordinator will record it as a lease's holder). `validateRoles` refuses an empty list, an empty entry, an unknown or a repeated role; `validateTopology` refuses a role set the backends cannot serve: any split over the embedded MQ, and a process with exactly one of `api` and `ingest` over a local cache. `NeedsDataDir` counts Pebble only for a process running `api`, and `Warnings` is empty without `api`, since only that role opens a cache it reads or a dedupe store. - **strict.go** — `rejectUnknownKeys`, the YAML half: re-reads the file as a generic tree and walks it against the struct's `yaml` tags, listing every key the struct doesn't declare. cleanenv itself is lenient by design, which is exactly wrong for boot config once keys have moved to the settings directory. - **persistence.go** — `WarnIfFreshDataDir` logs the startup `WARN` when `data_dir` is missing or empty (on a redeploy, the sign that the volume didn't persist); `LogStorageInitError` attaches the UID-65532 `permissionHint` to a NATS or Pebble open failure that looks like a permission denial — the same hint string `CheckDataDir` uses. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 6425d6f77..700045fc7 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -150,12 +150,11 @@ The `redis` backend's settings, read only when `cache.backend` is `redis`. It is | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | -| `cache.redis.addrs` | `WH_CACHE_REDIS_ADDRS` | *(empty)* | **Required** with `backend: redis`. `host:port` of the server; several are a cluster's seed nodes or the sentinels. Comma-separated in the env var. | -| `cache.redis.mode` | `WH_CACHE_REDIS_MODE` | `standalone` | `standalone`, `cluster` or `sentinel`. | -| `cache.redis.sentinel_master` | `WH_CACHE_REDIS_SENTINEL_MASTER` | *(empty)* | The master set name. Required with `mode: sentinel`. | +| `cache.redis.addrs` | `WH_CACHE_REDIS_ADDRS` | *(empty)* | **Required** with `backend: redis`. `host:port` of the server; several are a cluster's seed nodes. Comma-separated in the env var. | +| `cache.redis.mode` | `WH_CACHE_REDIS_MODE` | `standalone` | `standalone` or `cluster`. `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656): the cache does not yet authenticate to the sentinels or refresh their topology, so it could not be trusted to follow a failover. | | `cache.redis.username` | `WH_CACHE_REDIS_USERNAME` | *(empty)* | ACL user. Empty uses the server's `default` user. | | `cache.redis.password` | `WH_CACHE_REDIS_PASSWORD` | *(empty)* | A secret: set it through the environment (or your secret store's env injection), not in a tracked `config.yaml`. | -| `cache.redis.db` | `WH_CACHE_REDIS_DB` | `0` | Database number (`SELECT`). Standalone and sentinel only: a cluster has only database `0`, and any other value refuses boot. | +| `cache.redis.db` | `WH_CACHE_REDIS_DB` | `0` | Database number (`SELECT`). Standalone only: a cluster has only database `0`, and any other value refuses boot. | | `cache.redis.tls.enabled` | `WH_CACHE_REDIS_TLS_ENABLED` | `false` | Connect over TLS, verifying the server against the system roots or `ca_file`. Any other `tls` key set while this is off refuses boot, rather than connecting in plaintext. | | `cache.redis.tls.ca_file` | `WH_CACHE_REDIS_TLS_CA_FILE` | *(empty)* | PEM file of the authorities to trust instead of the system roots. | | `cache.redis.tls.cert_file` | `WH_CACHE_REDIS_TLS_CERT_FILE` | *(empty)* | Client certificate (PEM) for mutual TLS. Set together with `key_file`. | @@ -268,8 +267,7 @@ cache: l1_max_cost: 67108864 # the local backend's size redis: # read only with backend: redis addrs: [] # required with backend: redis, e.g. ["redis:6379"] - mode: standalone # standalone | cluster | sentinel - sentinel_master: "" + mode: standalone # standalone | cluster username: "" password: "" # a secret: set WH_CACHE_REDIS_PASSWORD instead db: 0 @@ -348,7 +346,6 @@ WH_CACHE_L1_MAX_COST=67108864 # Read only with WH_CACHE_BACKEND=redis; WH_CACHE_REDIS_ADDRS is then required. WH_CACHE_REDIS_ADDRS= WH_CACHE_REDIS_MODE=standalone -WH_CACHE_REDIS_SENTINEL_MASTER= WH_CACHE_REDIS_USERNAME= WH_CACHE_REDIS_PASSWORD= WH_CACHE_REDIS_DB=0 diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 2acf348c9..92c0cb95b 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -427,7 +427,7 @@ The folder name is the tenant id, and each folder is a complete settings directo Several WaveHouse instances can serve one ClickHouse behind a load balancer, but most of what each one holds is its own. The message queue is embedded, so an event is inserted by the instance that took its `POST /v1/ingest`, and reaches only that instance's SSE subscribers. The dedupe store is per instance too, so an id one instance has seen is new to another. -The query-result cache is the layer that can be shared today. With the default `cache.backend: local`, each instance caches in its own memory, and an insert invalidates only the cache of the instance that made it. Every other instance keeps serving its cached results for the rows before the insert until each entry's TTL runs out, between 10 s and 1 h depending on how long the query took. With [`cache.backend: redis`](/configuration#cache), every instance reads and fills one Redis-compatible server, and an insert on any instance invalidates the cached results of every instance. +The query-result cache is the layer that can be shared today. With the default `cache.backend: local`, each instance caches in its own memory, and an insert invalidates only the cache of the instance that made it. Every other instance keeps serving its cached results for the rows before the insert until each entry's TTL runs out, between 10 s and 1 h depending on how long the query took. With [`cache.backend: redis`](/configuration#cache), every instance reads and fills one Redis-compatible server, and an insert on any instance invalidates the cached results of every instance. The server is a standalone one or a Redis Cluster; Sentinel (`mode: sentinel`) refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656), since the cache does not yet authenticate to the sentinels or refresh their topology. **What another instance can see.** Ingest is already asynchronous: `/v1/ingest` answers before the batch is inserted. Once the inserting instance's worker has written the batch to ClickHouse, it replaces the table's version token in Redis, and from then on a lookup on any instance misses and reads the new rows. The cache adds no delay of its own beyond that single write. The exceptions: diff --git a/internal/app/wire.go b/internal/app/wire.go index ea23606d0..2ba112a4c 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -667,7 +667,6 @@ func redisConfig(r config.CacheRedisConfig) (cache.RedisConfig, error) { return cache.RedisConfig{ Addrs: r.Addrs, Mode: r.Mode, - SentinelMaster: r.SentinelMaster, Username: r.Username, Password: r.Password, DB: r.DB, diff --git a/internal/config/cache_redis.go b/internal/config/cache_redis.go index 997d19a3a..83ff056ca 100644 --- a/internal/config/cache_redis.go +++ b/internal/config/cache_redis.go @@ -11,7 +11,8 @@ import ( "time" ) -// Redis deployment modes for cache.redis.mode. +// Redis deployment modes for cache.redis.mode. RedisSentinel is refused +// until the backend supports it (#656). const ( RedisStandalone = "standalone" RedisCluster = "cluster" @@ -22,15 +23,13 @@ const ( // (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB) shared by every process. // Read only when that backend is selected. type CacheRedisConfig struct { - // Addrs are host:port pairs: the server, or seeds for a cluster, or the - // sentinels. - Addrs []string `yaml:"addrs" env:"WH_CACHE_REDIS_ADDRS"` - Mode string `yaml:"mode" env:"WH_CACHE_REDIS_MODE"` - SentinelMaster string `yaml:"sentinel_master" env:"WH_CACHE_REDIS_SENTINEL_MASTER"` - Username string `yaml:"username" env:"WH_CACHE_REDIS_USERNAME"` - Password string `yaml:"password" env:"WH_CACHE_REDIS_PASSWORD"` - DB int `yaml:"db" env:"WH_CACHE_REDIS_DB"` - TLS CacheRedisTLS `yaml:"tls"` + // Addrs are host:port pairs: the server, or seeds for a cluster. + Addrs []string `yaml:"addrs" env:"WH_CACHE_REDIS_ADDRS"` + Mode string `yaml:"mode" env:"WH_CACHE_REDIS_MODE"` + Username string `yaml:"username" env:"WH_CACHE_REDIS_USERNAME"` + Password string `yaml:"password" env:"WH_CACHE_REDIS_PASSWORD"` + DB int `yaml:"db" env:"WH_CACHE_REDIS_DB"` + TLS CacheRedisTLS `yaml:"tls"` // KeyPrefix leads every key, so deployments can share one server. KeyPrefix string `yaml:"key_prefix" env:"WH_CACHE_REDIS_KEY_PREFIX"` Timeout time.Duration `yaml:"timeout" env:"WH_CACHE_REDIS_TIMEOUT"` @@ -61,7 +60,7 @@ func (r CacheRedisConfig) hasAddrs() bool { func (r CacheRedisConfig) validate() error { if !r.hasAddrs() { - return errors.New("cache.backend=redis needs cache.redis.addrs (WH_CACHE_REDIS_ADDRS): the server's host:port, or a cluster's seeds, or the sentinels") + return errors.New("cache.backend=redis needs cache.redis.addrs (WH_CACHE_REDIS_ADDRS): the server's host:port, or a cluster's seeds") } for _, a := range r.Addrs { if strings.TrimSpace(a) != a { @@ -72,12 +71,11 @@ func (r CacheRedisConfig) validate() error { } } switch r.Mode { - case RedisStandalone, RedisCluster, RedisSentinel: + case RedisStandalone, RedisCluster: + case RedisSentinel: + return fmt.Errorf("cache.redis.mode (WH_CACHE_REDIS_MODE) %q is not supported yet: the cache neither authenticates to the sentinels nor refreshes their topology (https://github.com/Wave-RF/WaveHouse/issues/656); valid: %s, %s", r.Mode, RedisStandalone, RedisCluster) default: - return fmt.Errorf("cache.redis.mode (WH_CACHE_REDIS_MODE) %q: valid: %s, %s, %s", r.Mode, RedisStandalone, RedisCluster, RedisSentinel) - } - if r.Mode == RedisSentinel && r.SentinelMaster == "" { - return errors.New("cache.redis.mode=sentinel needs cache.redis.sentinel_master (WH_CACHE_REDIS_SENTINEL_MASTER), the master set name") + return fmt.Errorf("cache.redis.mode (WH_CACHE_REDIS_MODE) %q: valid: %s, %s", r.Mode, RedisStandalone, RedisCluster) } if r.DB < 0 { return fmt.Errorf("cache.redis.db (WH_CACHE_REDIS_DB) %d is negative", r.DB) diff --git a/internal/config/cache_redis_test.go b/internal/config/cache_redis_test.go index 5389715a9..d9d3adab2 100644 --- a/internal/config/cache_redis_test.go +++ b/internal/config/cache_redis_test.go @@ -45,9 +45,8 @@ func TestLoad_CacheRedisFromEnv(t *testing.T) { caFile, certFile, keyFile := writeTestPKI(t, dir) for k, v := range map[string]string{ "WH_CACHE_BACKEND": "redis", - "WH_CACHE_REDIS_ADDRS": "s1:26379,s2:26379", - "WH_CACHE_REDIS_MODE": "sentinel", - "WH_CACHE_REDIS_SENTINEL_MASTER": "mymaster", + "WH_CACHE_REDIS_ADDRS": "r1:6379,r2:6379", + "WH_CACHE_REDIS_MODE": "standalone", "WH_CACHE_REDIS_USERNAME": "wavehouse", "WH_CACHE_REDIS_PASSWORD": "s3cret", "WH_CACHE_REDIS_DB": "2", @@ -69,7 +68,7 @@ func TestLoad_CacheRedisFromEnv(t *testing.T) { cfg, err := Load("nonexistent.yaml") require.NoError(t, err) assert.Equal(t, CacheRedisConfig{ - Addrs: []string{"s1:26379", "s2:26379"}, Mode: RedisSentinel, SentinelMaster: "mymaster", + Addrs: []string{"r1:6379", "r2:6379"}, Mode: RedisStandalone, Username: "wavehouse", Password: "s3cret", DB: 2, TLS: CacheRedisTLS{ Enabled: true, CAFile: caFile, CertFile: certFile, KeyFile: keyFile, ServerName: "redis.internal", @@ -142,6 +141,7 @@ cache: addr: r:6379 near_cache: max_cost: 1 + sentinel_master: mymaster tls: ca: /x memcached: @@ -149,7 +149,7 @@ cache: `), 0o600)) _, err := Load(path) require.Error(t, err) - assert.Contains(t, err.Error(), "cache.memcached, cache.redis.addr, cache.redis.near_cache, cache.redis.tls.ca") + assert.Contains(t, err.Error(), "cache.memcached, cache.redis.addr, cache.redis.near_cache, cache.redis.sentinel_master, cache.redis.tls.ca") } // The documented env file lists WH_CACHE_REDIS_ADDRS blank: that is no @@ -172,6 +172,7 @@ func TestUnboundEnv_KnowsTheCacheRedisVariables(t *testing.T) { "WH_CACHE_REDIS_VERSION_TTL=1h", "WH_CACHE_REDIS_COMPRESS_MIN_BYTES=0", })) assert.Equal(t, []string{"WH_CACHE_REDIS_ADDR"}, unboundEnv([]string{"WH_CACHE_REDIS_ADDR=r:6379"})) + assert.Equal(t, []string{"WH_CACHE_REDIS_SENTINEL_MASTER"}, unboundEnv([]string{"WH_CACHE_REDIS_SENTINEL_MASTER=m"}), "no sentinel mode until #656") } func TestValidate_CacheRedis(t *testing.T) { @@ -189,9 +190,9 @@ func TestValidate_CacheRedis(t *testing.T) { {"no addrs", func(r *CacheRedisConfig) { r.Addrs = nil }, "cache.backend=redis needs cache.redis.addrs (WH_CACHE_REDIS_ADDRS)"}, {"addr with space", func(r *CacheRedisConfig) { r.Addrs = []string{"a:6379", " b:6379"} }, "no spaces around an address"}, {"addr without port", func(r *CacheRedisConfig) { r.Addrs = []string{"redis"} }, `cache.redis.addrs (WH_CACHE_REDIS_ADDRS) "redis": want host:port`}, - {"mode", func(r *CacheRedisConfig) { r.Mode = "replica" }, `cache.redis.mode (WH_CACHE_REDIS_MODE) "replica": valid: standalone, cluster, sentinel`}, - {"sentinel without master", func(r *CacheRedisConfig) { r.Mode = RedisSentinel }, "needs cache.redis.sentinel_master"}, - {"sentinel", func(r *CacheRedisConfig) { r.Mode, r.SentinelMaster = RedisSentinel, "m" }, ""}, + {"mode", func(r *CacheRedisConfig) { r.Mode = "replica" }, `cache.redis.mode (WH_CACHE_REDIS_MODE) "replica": valid: standalone, cluster`}, + {"sentinel refused", func(r *CacheRedisConfig) { r.Mode = RedisSentinel }, `cache.redis.mode (WH_CACHE_REDIS_MODE) "sentinel" is not supported yet: the cache neither authenticates to the sentinels nor refreshes their topology (https://github.com/Wave-RF/WaveHouse/issues/656)`}, + {"cluster", func(r *CacheRedisConfig) { r.Mode = RedisCluster }, ""}, {"cluster db", func(r *CacheRedisConfig) { r.Mode, r.DB = RedisCluster, 1 }, "a Redis cluster has only database 0"}, {"standalone db", func(r *CacheRedisConfig) { r.DB = 3 }, ""}, {"negative db", func(r *CacheRedisConfig) { r.DB = -1 }, "is negative"}, From a7ef71fccd3d13f647d21e2a9c12ff7a42363692 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:03:55 -0400 Subject: [PATCH 36/79] fix(config): refuse a URL-style cache.redis address without echoing it WH_CACHE_REDIS_ADDRS=redis://default:s3cret@redis:6379 failed boot with an error that quoted the address, password included, twice: once in the key's message and once in net.SplitHostPort's. An address holding "://" or "@" is now refused by its position alone, pointing at the username, password and TLS settings that take those parts. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/configuration.mdx | 2 +- internal/config/cache_redis.go | 7 ++++++- internal/config/cache_redis_test.go | 21 +++++++++++++++++++++ 5 files changed, 30 insertions(+), 4 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index ac09db1ae..471ccc173 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. +- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. - **A Redis-compatible shared cache backend** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row open a circuit breaker that skips the server until a probe succeeds, and deferred invalidations are retried until they land, collapsing to one tenant-wide bump per tenant past 100,000 keys. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. `cache.backend: redis` selects it (the entry above). Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression and server-stops-answering cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 4c627ee7c..c0f83326f 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -126,7 +126,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — the implementation each layer runs on, the process's `roles`, resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `cache.redis.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). - **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load` when `Config.NeedsDataDir` says a selected backend keeps state there, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. - **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` in `Validate`, between `validateRoles` and `validateTopology`), which refuses a value not on the list and names the ones that are. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns what a valid configuration is still likely to get wrong — the combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), a `cache.redis` block that is not read, certificate verification turned off — which `app.New` logs at `WARN`. -- **cache_redis.go** — `CacheRedisConfig`, the `cache.redis` sub-block, and its checks: an address, each `host:port`, a known mode (`sentinel` is refused until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `db` 0 in cluster mode, positive timeouts and sizes, a `version_ttl` of at least 2 s, and a `compress_min_bytes` that is not negative (`0` never compresses); its defaults are in `defaults()` with the rest. `CacheRedisTLS.Config` builds the `tls.Config`, reading the files; `Validate` calls it so an unreadable file refuses boot, and `internal/app` calls it again to build the connection. A TLS key set while `tls.enabled` is off is an error rather than a plaintext connection. +- **cache_redis.go** — `CacheRedisConfig`, the `cache.redis` sub-block, and its checks: an address, each `host:port` (a URL or `user:password@` form refused without repeating it, since it may hold a password), a known mode (`sentinel` is refused until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `db` 0 in cluster mode, positive timeouts and sizes, a `version_ttl` of at least 2 s, and a `compress_min_bytes` that is not negative (`0` never compresses); its defaults are in `defaults()` with the rest. `CacheRedisTLS.Config` builds the `tls.Config`, reading the files; `Validate` calls it so an unreadable file refuses boot, and `internal/app` calls it again to build the connection. A TLS key set while `tls.enabled` is off is an error rather than a plaintext connection. - **config.go**, roles — `roles` (`[]Role`: `api`, `ingest`, `sweeper`; `AllRoles` by default; `Has(Role)`) picks which components `internal/app` wires, and `instance_id` names the process (`-<8 hex>` when empty, resolved in `Load`; today only logged at boot, and a distributed coordinator will record it as a lease's holder). `validateRoles` refuses an empty list, an empty entry, an unknown or a repeated role; `validateTopology` refuses a role set the backends cannot serve: any split over the embedded MQ, and a process with exactly one of `api` and `ingest` over a local cache. `NeedsDataDir` counts Pebble only for a process running `api`, and `Warnings` is empty without `api`, since only that role opens a cache it reads or a dedupe store. - **strict.go** — `rejectUnknownKeys`, the YAML half: re-reads the file as a generic tree and walks it against the struct's `yaml` tags, listing every key the struct doesn't declare. cleanenv itself is lenient by design, which is exactly wrong for boot config once keys have moved to the settings directory. - **persistence.go** — `WarnIfFreshDataDir` logs the startup `WARN` when `data_dir` is missing or empty (on a redeploy, the sign that the volume didn't persist); `LogStorageInitError` attaches the UID-65532 `permissionHint` to a NATS or Pebble open failure that looks like a permission denial — the same hint string `CheckDataDir` uses. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 700045fc7..f15788f37 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -150,7 +150,7 @@ The `redis` backend's settings, read only when `cache.backend` is `redis`. It is | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | -| `cache.redis.addrs` | `WH_CACHE_REDIS_ADDRS` | *(empty)* | **Required** with `backend: redis`. `host:port` of the server; several are a cluster's seed nodes. Comma-separated in the env var. | +| `cache.redis.addrs` | `WH_CACHE_REDIS_ADDRS` | *(empty)* | **Required** with `backend: redis`. `host:port` of the server; several are a cluster's seed nodes. Comma-separated in the env var. Not a URL: `redis://user:password@host:port` refuses boot, without repeating the value, so set the credentials through `username` and `password`, and `rediss://` through `tls.enabled`. | | `cache.redis.mode` | `WH_CACHE_REDIS_MODE` | `standalone` | `standalone` or `cluster`. `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656): the cache does not yet authenticate to the sentinels or refresh their topology, so it could not be trusted to follow a failover. | | `cache.redis.username` | `WH_CACHE_REDIS_USERNAME` | *(empty)* | ACL user. Empty uses the server's `default` user. | | `cache.redis.password` | `WH_CACHE_REDIS_PASSWORD` | *(empty)* | A secret: set it through the environment (or your secret store's env injection), not in a tracked `config.yaml`. | diff --git a/internal/config/cache_redis.go b/internal/config/cache_redis.go index 83ff056ca..0e41b0ff1 100644 --- a/internal/config/cache_redis.go +++ b/internal/config/cache_redis.go @@ -62,7 +62,12 @@ func (r CacheRedisConfig) validate() error { if !r.hasAddrs() { return errors.New("cache.backend=redis needs cache.redis.addrs (WH_CACHE_REDIS_ADDRS): the server's host:port, or a cluster's seeds") } - for _, a := range r.Addrs { + for i, a := range r.Addrs { + // A redis:// URL, or user:pass@host, may carry a password: refuse it + // without echoing it into the boot error and the logs. + if strings.Contains(a, "://") || strings.Contains(a, "@") { + return fmt.Errorf("cache.redis.addrs (WH_CACHE_REDIS_ADDRS) entry %d is a URL or holds credentials (not echoed): give host:port, and set the user and password with WH_CACHE_REDIS_USERNAME and WH_CACHE_REDIS_PASSWORD, and TLS (rediss://) with WH_CACHE_REDIS_TLS_ENABLED", i+1) + } if strings.TrimSpace(a) != a { return fmt.Errorf("cache.redis.addrs (WH_CACHE_REDIS_ADDRS) %q: no spaces around an address", a) } diff --git a/internal/config/cache_redis_test.go b/internal/config/cache_redis_test.go index d9d3adab2..21ec15fcf 100644 --- a/internal/config/cache_redis_test.go +++ b/internal/config/cache_redis_test.go @@ -165,6 +165,27 @@ func TestLoad_CacheRedisBlankAddrs(t *testing.T) { require.ErrorContains(t, err, "cache.backend=redis needs cache.redis.addrs") } +// A URL-style address carries its password, and a boot error reaches the +// logs: the refusal names the entry, never its value. +func TestLoad_CacheRedisRefusesURLAddrsWithoutEchoingThem(t *testing.T) { + for _, addr := range []string{ + "redis://default:s3cret@redis:6379", + "rediss://default:s3cret@redis:6380", + "default:s3cret@redis:6379", + "redis://redis:6379", + } { + t.Run(addr, func(t *testing.T) { + t.Setenv("WH_CACHE_BACKEND", "redis") + t.Setenv("WH_CACHE_REDIS_ADDRS", "ok:6379,"+addr) + _, err := Load("nonexistent.yaml") + require.ErrorContains(t, err, "cache.redis.addrs (WH_CACHE_REDIS_ADDRS) entry 2 is a URL or holds credentials") + assert.Contains(t, err.Error(), "WH_CACHE_REDIS_USERNAME and WH_CACHE_REDIS_PASSWORD") + assert.NotContains(t, err.Error(), "s3cret") + assert.NotContains(t, err.Error(), addr) + }) + } +} + func TestUnboundEnv_KnowsTheCacheRedisVariables(t *testing.T) { t.Parallel() assert.Empty(t, unboundEnv([]string{ From de27b600f670ec6ae53777836fd3a1f85951eb1c Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:04:55 -0400 Subject: [PATCH 37/79] fix(config): cap cache.redis.dial_timeout at 2s A dial is a TCP connect and a handshake, each bounded by dial_timeout. NewRedis dials once synchronously, so boot waits out one, and the cache's Close waits out a reconnect dial in flight: rueidis.NewClient takes no context. An unbounded dial_timeout could therefore push the cache's close past the 5s release budget and leave the queue, dedupe and ClickHouse pools unreleased. At 2s the wait is at most 4s; the default (1s) is unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/configuration.mdx | 2 +- internal/config/cache_redis.go | 9 +++++++++ internal/config/cache_redis_test.go | 6 ++++-- 5 files changed, 16 insertions(+), 5 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 471ccc173..1454d47ae 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. +- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`, at most `2s`: boot and shutdown each wait out a dial), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. - **A Redis-compatible shared cache backend** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row open a circuit breaker that skips the server until a probe succeeds, and deferred invalidations are retried until they land, collapsing to one tenant-wide bump per tenant past 100,000 keys. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. `cache.backend: redis` selects it (the entry above). Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression and server-stops-answering cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index c0f83326f..cd44c8d1b 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -126,7 +126,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — the implementation each layer runs on, the process's `roles`, resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `cache.redis.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). - **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load` when `Config.NeedsDataDir` says a selected backend keeps state there, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. - **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` in `Validate`, between `validateRoles` and `validateTopology`), which refuses a value not on the list and names the ones that are. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns what a valid configuration is still likely to get wrong — the combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), a `cache.redis` block that is not read, certificate verification turned off — which `app.New` logs at `WARN`. -- **cache_redis.go** — `CacheRedisConfig`, the `cache.redis` sub-block, and its checks: an address, each `host:port` (a URL or `user:password@` form refused without repeating it, since it may hold a password), a known mode (`sentinel` is refused until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `db` 0 in cluster mode, positive timeouts and sizes, a `version_ttl` of at least 2 s, and a `compress_min_bytes` that is not negative (`0` never compresses); its defaults are in `defaults()` with the rest. `CacheRedisTLS.Config` builds the `tls.Config`, reading the files; `Validate` calls it so an unreadable file refuses boot, and `internal/app` calls it again to build the connection. A TLS key set while `tls.enabled` is off is an error rather than a plaintext connection. +- **cache_redis.go** — `CacheRedisConfig`, the `cache.redis` sub-block, and its checks: an address, each `host:port` (a URL or `user:password@` form refused without repeating it, since it may hold a password), a known mode (`sentinel` is refused until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `db` 0 in cluster mode, positive timeouts and sizes, a `dial_timeout` of at most 2 s (boot and `Close` each wait out a dial, a connect and a handshake bounded by it), a `version_ttl` of at least 2 s, and a `compress_min_bytes` that is not negative (`0` never compresses); its defaults are in `defaults()` with the rest. `CacheRedisTLS.Config` builds the `tls.Config`, reading the files; `Validate` calls it so an unreadable file refuses boot, and `internal/app` calls it again to build the connection. A TLS key set while `tls.enabled` is off is an error rather than a plaintext connection. - **config.go**, roles — `roles` (`[]Role`: `api`, `ingest`, `sweeper`; `AllRoles` by default; `Has(Role)`) picks which components `internal/app` wires, and `instance_id` names the process (`-<8 hex>` when empty, resolved in `Load`; today only logged at boot, and a distributed coordinator will record it as a lease's holder). `validateRoles` refuses an empty list, an empty entry, an unknown or a repeated role; `validateTopology` refuses a role set the backends cannot serve: any split over the embedded MQ, and a process with exactly one of `api` and `ingest` over a local cache. `NeedsDataDir` counts Pebble only for a process running `api`, and `Warnings` is empty without `api`, since only that role opens a cache it reads or a dedupe store. - **strict.go** — `rejectUnknownKeys`, the YAML half: re-reads the file as a generic tree and walks it against the struct's `yaml` tags, listing every key the struct doesn't declare. cleanenv itself is lenient by design, which is exactly wrong for boot config once keys have moved to the settings directory. - **persistence.go** — `WarnIfFreshDataDir` logs the startup `WARN` when `data_dir` is missing or empty (on a redeploy, the sign that the volume didn't persist); `LogStorageInitError` attaches the UID-65532 `permissionHint` to a NATS or Pebble open failure that looks like a permission denial — the same hint string `CheckDataDir` uses. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index f15788f37..e00fdaec0 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -163,7 +163,7 @@ The `redis` backend's settings, read only when `cache.backend` is `redis`. It is | `cache.redis.tls.insecure_skip_verify` | `WH_CACHE_REDIS_TLS_INSECURE_SKIP_VERIFY` | `false` | Accept any server certificate. Logged at `WARN` at boot: whoever can intercept the connection can read and replace cached results. | | `cache.redis.key_prefix` | `WH_CACHE_REDIS_KEY_PREFIX` | `wh` | Leads every key, so several deployments can share one server, provided you trust each as much as the others: any of them can overwrite what the rest serve. No `{` or `}`. | | `cache.redis.timeout` | `WH_CACHE_REDIS_TIMEOUT` | `100ms` | Per operation. A lookup or fill that takes longer is a miss or a skipped fill, never a failed query. | -| `cache.redis.dial_timeout` | `WH_CACHE_REDIS_DIAL_TIMEOUT` | `1s` | Per connection attempt. | +| `cache.redis.dial_timeout` | `WH_CACHE_REDIS_DIAL_TIMEOUT` | `1s` | Per connection attempt, at most `2s`. An attempt is a connect and a handshake, each bounded by this, and both boot and shutdown wait out one in flight, so the cap keeps a stop inside the release's fixed 5s budget (see [Stopping](/deployment#stopping)). | | `cache.redis.max_value_bytes` | `WH_CACHE_REDIS_MAX_VALUE_BYTES` | `1048576` | Largest result stored, after compression (1 MiB). A larger one is returned to the caller but not cached. | | `cache.redis.compress_min_bytes` | `WH_CACHE_REDIS_COMPRESS_MIN_BYTES` | `1024` | Results at least this large are zstd-compressed when that makes them smaller. `0` never compresses. | | `cache.redis.version_ttl` | `WH_CACHE_REDIS_VERSION_TTL` | `168h` | How long a table's or tenant's version token outlives its last write, so the tokens of dropped tables and removed tenants eventually expire. At least `2s`. An expired token only causes misses. | diff --git a/internal/config/cache_redis.go b/internal/config/cache_redis.go index 0e41b0ff1..920630157 100644 --- a/internal/config/cache_redis.go +++ b/internal/config/cache_redis.go @@ -19,6 +19,12 @@ const ( RedisSentinel = "sentinel" ) +// maxRedisDialTimeout caps cache.redis.dial_timeout. A dial is a connect +// and a handshake, each bounded by it, and both boot and the cache's Close +// wait out one in flight: at 2s that is at most 4s, inside the 5s budget +// Close shares with the stores released after the cache. +const maxRedisDialTimeout = 2 * time.Second + // CacheRedisConfig configures cache.backend=redis: one Redis-compatible server // (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB) shared by every process. // Read only when that backend is selected. @@ -102,6 +108,9 @@ func (r CacheRedisConfig) validate() error { return fmt.Errorf("%s %s must be positive", d.key, d.v) } } + if r.DialTimeout > maxRedisDialTimeout { + return fmt.Errorf("cache.redis.dial_timeout (WH_CACHE_REDIS_DIAL_TIMEOUT) %s is over %s: boot and shutdown each wait out a dial, up to twice this", r.DialTimeout, maxRedisDialTimeout) + } if r.VersionTTL < 2*time.Second { return fmt.Errorf("cache.redis.version_ttl (WH_CACHE_REDIS_VERSION_TTL) %s is under 2s", r.VersionTTL) } diff --git a/internal/config/cache_redis_test.go b/internal/config/cache_redis_test.go index 21ec15fcf..7f91a0499 100644 --- a/internal/config/cache_redis_test.go +++ b/internal/config/cache_redis_test.go @@ -57,7 +57,7 @@ func TestLoad_CacheRedisFromEnv(t *testing.T) { "WH_CACHE_REDIS_TLS_SERVER_NAME": "redis.internal", "WH_CACHE_REDIS_KEY_PREFIX": "staging", "WH_CACHE_REDIS_TIMEOUT": "250ms", - "WH_CACHE_REDIS_DIAL_TIMEOUT": "3s", + "WH_CACHE_REDIS_DIAL_TIMEOUT": "2s", "WH_CACHE_REDIS_MAX_VALUE_BYTES": "2048", "WH_CACHE_REDIS_COMPRESS_MIN_BYTES": "0", "WH_CACHE_REDIS_VERSION_TTL": "24h", @@ -73,7 +73,7 @@ func TestLoad_CacheRedisFromEnv(t *testing.T) { TLS: CacheRedisTLS{ Enabled: true, CAFile: caFile, CertFile: certFile, KeyFile: keyFile, ServerName: "redis.internal", }, - KeyPrefix: "staging", Timeout: 250 * time.Millisecond, DialTimeout: 3 * time.Second, + KeyPrefix: "staging", Timeout: 250 * time.Millisecond, DialTimeout: 2 * time.Second, MaxValueBytes: 2048, CompressMinBytes: 0, VersionTTL: 24 * time.Hour, }, cfg.Cache.Redis) tc, err := cfg.Cache.Redis.TLS.Config() @@ -221,6 +221,8 @@ func TestValidate_CacheRedis(t *testing.T) { {"brace prefix", func(r *CacheRedisConfig) { r.KeyPrefix = "a{b}" }, "hash-tag brace"}, {"zero timeout", func(r *CacheRedisConfig) { r.Timeout = 0 }, "cache.redis.timeout (WH_CACHE_REDIS_TIMEOUT) 0s must be positive"}, {"negative dial timeout", func(r *CacheRedisConfig) { r.DialTimeout = -time.Second }, "cache.redis.dial_timeout"}, + {"dial timeout at the cap", func(r *CacheRedisConfig) { r.DialTimeout = 2 * time.Second }, ""}, + {"dial timeout over the cap", func(r *CacheRedisConfig) { r.DialTimeout = 2*time.Second + time.Millisecond }, "cache.redis.dial_timeout (WH_CACHE_REDIS_DIAL_TIMEOUT) 2.001s is over 2s"}, {"short version ttl", func(r *CacheRedisConfig) { r.VersionTTL = time.Second }, "cache.redis.version_ttl (WH_CACHE_REDIS_VERSION_TTL) 1s is under 2s"}, {"zero max value", func(r *CacheRedisConfig) { r.MaxValueBytes = 0 }, "cache.redis.max_value_bytes"}, {"compress never", func(r *CacheRedisConfig) { r.CompressMinBytes = 0 }, ""}, From 972ff30c2b1c201bf9cfe0bc30e6db7dcc452a89 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:05:55 -0400 Subject: [PATCH 38/79] docs(cache): only a move to another address or database bumps a tenant Pools.Reconcile reports a tenant stale only when it moves to another ClickHouse address or database, or returns to a pool after an absence; a username or tls change reads the same tables and bumps nothing. Say so where the snapshot ordering is described, and replace the last "ahead of the cache" with what the no-pool 503 now precedes: a cached result being served, or a query running. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- AGENTS.md | 2 +- CHANGELOG.md | 4 ++-- docs/src/content/docs/api.md | 4 ++-- docs/src/content/docs/architecture.md | 4 ++-- internal/api/cache_tenant_test.go | 8 ++++---- internal/api/pipes.go | 4 ++-- internal/api/structured_query.go | 4 ++-- internal/api/tenant_clickhouse_test.go | 4 ++-- 8 files changed, 17 insertions(+), 17 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 60474bc35..12805f8b8 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -31,7 +31,7 @@ Twenty internal packages under `internal/` (plus `internal/testutil/` for shared - **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers; `ch_errors.go` (`writeCHError`) is the one mapping from a failed ClickHouse query to status, `code` and `retryable` - **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` handing the sweeper each tenant's own gap window (a rejected tenant's as its folder last had it, unbounded for one rejected since boot) and the `mq.max_bytes_gb` reconcile each served tenant's byte budget, and `defaultPolicy` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, and the same hook's `Hub.Prune` ends the open streams of a tenant no longer served), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `New` wires only what the process's `roles` need (discovery, dedupe, auth verifiers, the hub bridge and keepalive per API process; the ingest worker per ingest process; the sweeper under its lease through `elected`); a process without `api` serves `api.NewOpsRouter` — probes, `/version`, metrics, and the settings reload behind the operator key alone. `cmd/wavehouse` and `tests/integration` both boot through it - **`auth/`** — JWT auth middleware: HMAC **or** JWKS verification with `alg` pinned to the active verifier, role extraction from a configurable claim path; always runs, never rejects (bad token → empty role + stashed reason). One verifier per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 9): `Authenticator` keys them by `tenant.ID` — the request store's `settings.Store.Tenant()`, through an injected `TenantSource`; `tenant.Default` on the tenant-exempt routes — built from each tenant's `auth` block by `Reconfigure`, dropped by `Prune` once the tenant stops being served (rejected or removed), released by `Close`; the secrets (`Config`) are boot-level and shared. A JWKS key set is fetched off the boot and reload paths: until one has been stored the verifier is pending and a token-bearing request gets `503` + `Retry-After` from `api.refuseUnverifiable` (`auth.ErrVerifierPending`), never a `default_role` evaluation; refresh is library-managed (Eric, 2026-09-22), response capped at 1 MiB; the operator key's admin role is the request tenant's -- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for the caller's query key and its singleflight (escaped whole as the lead field of the stored key, `|.||…`), `..
.
.` for a namespace, each field escaped by `keyenc` where the key is built (a `Namespace` carries the raw table and scope, so no caller escapes) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version folded into every namespace and query key of one tenant, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read, taken before the handler chooses any input a bump invalidates — the tenant's connection included — and `Set` files the fill under it, so a write landing mid-query, or a reload repointing the tenant after the request took its connection, orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) +- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for the caller's query key and its singleflight (escaped whole as the lead field of the stored key, `|.||…`), `..
.
.` for a namespace, each field escaped by `keyenc` where the key is built (a `Namespace` carries the raw table and scope, so no caller escapes) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version folded into every namespace and query key of one tenant, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read, taken before the handler chooses any input a bump invalidates — the tenant's connection included — and `Set` files the fill under it, so a write landing mid-query, or a reload moving the tenant to another address or database after the request took its connection, orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config. `Classify` (`errclass.go`) says what a failed ClickHouse request means for the request — `Unavailable`, `Denied`, `Rejected` (any unlisted exception code: the server read it and refused it), or `Unknown` (no code, no recognizable transport failure) — over the driver's error types and the HTTP interface's `HTTPError`; the ingest worker and the query handlers (`api/ch_errors.go` `writeCHError`, [#403](https://github.com/Wave-RF/WaveHouse/issues/403), [#271](https://github.com/Wave-RF/WaveHouse/issues/271)) both use it - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) - **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability when a selected backend keeps state there (`NeedsDataDir`); `backends.go` holds each layer's `.backend` (only the in-process value today); `config.go` holds `roles` (`Has(Role)`) and `instance_id`, and `Validate` refuses a role split the backends cannot serve (any split over the embedded MQ; `api` without `ingest`, or the reverse, over a local cache) — boot is the validator, there is no dry run diff --git a/CHANGELOG.md b/CHANGELOG.md index e8ccf372c..6baafa281 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -16,7 +16,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/config/defaults_test.go`, `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block; no backend has settings yet, so every such sub-block is an unknown key for now and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name the `mq`, `cache` and `dedupe` backends: the zero value is not the default, and `app.New` refuses it. The defaults live in `defaults()`, like every boot key's since [#631](https://github.com/Wave-RF/WaveHouse/issues/631), so an explicit `backend: ""` in `config.yaml` refuses boot rather than becoming the default. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. -- **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (since #614 this drops the tenant's cached pipe results too; no insert invalidates a pipe result, which names no table). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. +- **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, before a cached result is served or a query runs. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (since #614 this drops the tenant's cached pipe results too; no insert invalidates a pipe result, which names no table). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. - **ClickHouse TLS, HTTP-interface headers, pool sizes and a connection ceiling** (`internal/settings/{settings,validate,store}.go` + seed, `internal/chconn/chconn.go`, `internal/ingest/worker.go`, `internal/api/query.go`, `internal/config/config.go`, `internal/app/wire.go`, `deployments/compose/settings/config.json`, `deployments/compose/standalone.yaml`, `config.yaml`, `docs/src/content/docs/{settings-directory,configuration,reverse-proxy}.mdx`, `docs/src/content/docs/{architecture,deployment}.md`): the tenant-agnostic first slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). `config.json`'s `clickhouse` block gains a `tls` block (`enabled`, `ca_file`, `cert_file`, `key_file`, `insecure_skip_verify`, `server_name`), a `headers` map for the HTTP interface, and `max_open_conns` / `max_idle_conns`. **Every key is required, so an existing `config.json` must add them**; the seed values (`tls.enabled: false` with the other `tls` keys empty, `headers: {}`, `10` / `5`) change nothing. `tls.enabled` switches the native hop to TLS, `http_scheme` stays the HTTP hop's switch, and the material applies to whichever hop uses TLS: the driver gets the TLS config and the pool sizes, the ingest worker and the raw-SQL proxy get the TLS config and the headers, set ahead of their own so the credentials win (naming `X-ClickHouse-User`, `X-ClickHouse-Key` or `Authorization` is a validation error, and so are two spellings of one name). Validation checks shape only — the paths are not opened, so `wavehouse validate` runs anywhere — warns when `insecure_skip_verify` is on or when only one of the two hops is on TLS (each carries the credentials in the clear without it), and a certificate file that cannot be read or parsed refuses boot or leaves a reload's connection unchanged; the files are read when the connection is built, so a file replaced in place needs a restart. Boot config gains the optional `clickhouse.max_total_conns` (`WH_CH_MAX_TOTAL_CONNS`, `0` = no ceiling): a settings pool above it refuses boot, naming both numbers in the error, and a reload that raises the pool above it is refused and logged (the reload still reports adopted), leaving the connection as it was. @@ -88,7 +88,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **A failed ClickHouse query answers by what went wrong, not a flat `500`/`502`** (`internal/api/ch_errors.go` (new, + tests), `internal/api/{errors,query,structured_query,pipes,schema,ch_settings}.go`, `internal/chconn/errclass.go` (`HTTPStatus` exported), `clients/ts/src/errors.ts` (+ tests), `tests/integration/query_errors_test.go` (new), `tests/integration/query_limits_test.go`, `internal/app/app_test.go`, `tests/e2e/sdk/{admin,query}.test.ts`, `AGENTS.md`, `docs/src/content/docs/{api,architecture}.md`, `docs/src/content/docs/{access-control,configuration}.mdx`, `docs/src/content/docs/sdk/{reference.md,index.mdx}`): fixes [#403](https://github.com/Wave-RF/WaveHouse/issues/403) and [#271](https://github.com/Wave-RF/WaveHouse/issues/271), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). ClickHouse answers a syntax error, a missing grant and an overloaded server alike with HTTP `500`, so `/v1/ops/query` turned a bad statement into a `502` and `/v1/query` and pipes into a `500` the SDK retried. All three now class the failure with `chconn.Classify` through one helper, `writeCHError`: a statement ClickHouse refused is `400 clickhouse.rejected`; a query over a rows/bytes limit, the role's own memory cap, or its time cap where that is no longer than `query_timeout` is `400 clickhouse.limit_exceeded`; `ACCESS_DENIED` is `403 clickhouse.access_denied`; credentials, user or database refused, or a redirect or `4xx` with no exception code from whatever fronts ClickHouse, is `502 clickhouse.misconfigured`; ClickHouse down, unreachable or overloaded is `503 clickhouse.unavailable` with `Retry-After: 5`; a failure with no verdict stays `500` (`502` on the proxy) as `clickhouse.unknown`. The error envelope gains `code` and `retryable` next to `error` on these responses — additive. A role with `max_execution_time` now queries with no context deadline and a cancel two seconds past the cap instead: clickhouse-go overwrote the cap's `max_execution_time` with deadline+5s for any deadline over 1s, so an overrun came back as a bare deadline, indistinguishable from waiting for a pooled connection; ClickHouse now enforces the cap itself and reports `TIMEOUT_EXCEEDED`. `POST /v1/ops/schema/refresh` against an unreachable ClickHouse is a `503` with `Retry-After` instead of a `500`. **SDK:** `WaveHouseError.code` and `retryable` now take the server's `code`/`retryable` when the body has them (`HTTP_` and "5xx retries" otherwise), so a rejected query is `clickhouse.rejected` rather than `HTTP_500`, and is not retried. - **An unavailable ClickHouse is retried with backoff instead of dead-lettering every row** (`internal/chconn/errclass.go` (new, + tests), `internal/ingest/{worker,backoff}.go` (`backoff.go` new, + tests), `internal/mq/{mq,embedded}.go`, `internal/testutil/mocks.go`, `tests/integration/ingest_outage_test.go` (new), `AGENTS.md`, `README.md`, `docs/src/content/docs/{ingest-pipeline,architecture,api,deployment,why-wavehouse}.md`, `docs/src/content/docs/{settings-directory,index,access-control}.mdx`): workstream A of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). A failed batch insert used to go through row-by-row isolation whatever the failure, so a ClickHouse that was down, overloaded or read-only failed every row twice and parked the whole batch on the DLQ. `chconn.Classify` now classes the failure first — `Rejected` (any ClickHouse exception code outside the availability and credential lists: the server read the row and refused it), `Unavailable` (connection refused/reset, timeouts, `TOO_MANY_SIMULTANEOUS_QUERIES`, `SERVER_OVERLOADED`, `MEMORY_LIMIT_EXCEEDED`, `TOO_MANY_PARTS`, `READONLY`, `TABLE_IS_READ_ONLY`, `KEEPER_EXCEPTION`, …), `Denied` (`AUTHENTICATION_FAILED`, `ACCESS_DENIED`, …) or `Unknown` (no code, no recognizable transport failure). Only `Rejected` is isolated and dead-lettered as before, and a multi-row batch refused with `TOO_MANY_PARTS` or `MEMORY_LIMIT_EXCEEDED` is split row by row first (`chconn.Splittable`), because a batch spanning too many partitions or too much memory can fail where each of its rows inserts; every other class hands the batch back to the queue with a delayed nak (`mq.Message.NakWithDelay`, new) under a jittered 1 s → 30 s backoff shared by every table on the same ClickHouse pool (a failure of one table — read-only, too many parts or mutations, a grant missing on it, `chconn.TableScoped` — backs off that table alone), which turns rows away without a request while it runs and probes once per window, and ClickHouse going away mid-isolation stops isolation and retries the rows it had not settled. Counted by the new `wavehouse_ingest_retries_total{table, reason}`; logged at `WARN` when an outage starts and at most every 30 s during it. A long outage now shows as a growing ingest stream and, at `mq.max_bytes_gb`, ingest `503`s — not as a full DLQ; a lasting failure of one table holds back its tenant's other tables once its waiting rows reach `maxAckPending`. Retried rows come back out of arrival order, which matters only to a `ReplacingMergeTree` without a version column or a `CollapsingMergeTree`. - **Schema discovery's retry loop jitters its backoff** (`internal/discovery/discovery.go` (+ tests), `internal/app/wire.go`, `internal/api/errors.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`): `RetryRefresh` slept exactly `2s * 2^n` capped at 60s, so instances retrying against one recovering ClickHouse fired in lockstep, every 60s on the same second. Each sleep is now drawn uniformly from below the backoff (full jitter), spreading the retries over the whole window and halving the mean wait — so a failing tenant's retries, their log lines and `wavehouse_schema_refresh_failures_total` come about twice as often ([#141](https://github.com/Wave-RF/WaveHouse/issues/141)). -- **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/query/ident.go` (removed, + tests), `internal/app/wire.go`, `docs/src/content/docs/{api,architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL (a pipe result's key folded no version, so pipes were unaffected; with the tenant's version in every key they now take the same snapshot). The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. The snapshot is taken before any input a bump invalidates is chosen, the tenant's ClickHouse connection included: both handlers now look up before they resolve the tenant's pool, so a reload that repoints the tenant after a request took the old pool orphans that request's fill instead of filing the old database's rows as fresh under the new tenant version. A tenant on no pool is still a `503` before anything is served, a cached result included. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. The cache now escapes table and scope names itself: a `Namespace` carries them raw and the version index builds its keys with `internal/keyenc`, so neither the structured-query read nor the ingest worker's invalidation escapes them (`query.SafeEncodeToken` is gone) and no name reaches a key unescaped. The suite pins that a name holding a dot, a space or a `%` is read and bumped under one key, and that names which would run together unescaped (`a.0.b` against `a` with scope `b.0.`) stay two entries. Namespace keys are unchanged; an entry's key now carries the tenant's version and escapes the caller's query key whole. All of them live in the process, so nothing stored is orphaned. +- **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/query/ident.go` (removed, + tests), `internal/app/wire.go`, `docs/src/content/docs/{api,architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL (a pipe result's key folded no version, so pipes were unaffected; with the tenant's version in every key they now take the same snapshot). The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. The snapshot is taken before any input a bump invalidates is chosen, the tenant's ClickHouse connection included: both handlers now look up before they resolve the tenant's pool, so a reload that moves the tenant to another address or database after a request took the old pool orphans that request's fill instead of filing the old database's rows as fresh under the new tenant version. A tenant on no pool is still a `503` before a cached result is served or a query runs. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. The cache now escapes table and scope names itself: a `Namespace` carries them raw and the version index builds its keys with `internal/keyenc`, so neither the structured-query read nor the ingest worker's invalidation escapes them (`query.SafeEncodeToken` is gone) and no name reaches a key unescaped. The suite pins that a name holding a dot, a space or a `%` is read and bumped under one key, and that names which would run together unescaped (`a.0.b` against `a` with scope `b.0.`) stay two entries. Namespace keys are unchanged; an entry's key now carries the tenant's version and escapes the caller's query key whole. All of them live in the process, so nothing stored is orphaned. - **An explicit `false`, `0` or `""` in `config.yaml` is no longer replaced by the key's default** (`internal/config/config.go`, `internal/config/defaults_test.go` (new), `docs/src/content/docs/configuration.mdx`, `AGENTS.md`): [#631](https://github.com/Wave-RF/WaveHouse/issues/631). Defaults lived in cleanenv `env-default` tags, which cleanenv applies after the YAML decode to any field still at its zero value, so it could not tell a key the file set to its zero value from one the file left out. `otel.traces.enabled: false`, `otel.metrics.enabled: false` and `otel.logs.enabled: false` came back `true`; `otel.traces.sample_rate: 0` and `otel.logs.sample_rate: 0` came back `1.0`; `server.shutdown_timeout: 0` came back `10`; `cache.l1_max_cost: 0`, `prometheus.path: ""` and `data_dir: ""` came back as their defaults; `server.port: 0` came back `8080`. All of it was silent. Defaults now live in one Go function, `defaults()`, which `Load` starts from before decoding the file and then applying `WH_*` variables, so the order is env > YAML > default and a key the file sets always wins. **Behaviour change if your file relied on the bug:** a zero you wrote now takes effect. A file that says `sample_rate: 0` now exports no traces (or no DEBUG/INFO logs), where it silently exported everything; a signal set `enabled: false` is now off; `shutdown_timeout: 0` now skips the drain. `cache.l1_max_cost: 0`, `server.port: 0`, and `data_dir: ""` now refuse boot (`cache init: MaxCost can't be zero`, `server.port 0 out of range`, `data_dir (WH_DATA_DIR) is required`) instead of running on the default; an empty `prometheus.path` refuses boot when `prometheus.enabled` is true. Delete the key to get the default back. Env vars are unchanged: they already honoured an explicit zero. New tests load through `config.Load` for every affected key (a YAML zero is kept, an absent key gets the default, env wins in both directions), refuse an `env-default` tag on any field, and pin each documented default in `configuration.mdx` to `defaults()`. - **An embedded queue store that cannot be created fails boot at once, naming the cause** (`internal/mq/embedded.go` (+ tests)): part of [#617](https://github.com/Wave-RF/WaveHouse/issues/617). A regular file at `/nats`, or a `nats` directory that could not be created there, failed JetStream in the background, so boot waited out the server's 5s readiness check and reported only `nats server not ready`. `NewEmbedded` now creates the directory first (at `0700`, as the server does) and refuses boot with the mkdir error. An existing but unwritable `nats` directory still takes the old path. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its cached results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index 3f95cac89..b80e73b12 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -578,7 +578,7 @@ The inbound request body is capped at 1 MiB; a body over the cap is rejected wit | 400 / 403 / 502 / 503 | `{"error":"clickhouse query: …","code":"clickhouse.…","retryable":…}` | ClickHouse failed the query: a column dropped since the schema was discovered (`400 clickhouse.rejected`), the role's `max_rows_to_read`/`max_memory_usage` cap, or a `max_execution_time` no longer than `clickhouse.query_timeout` (`400 clickhouse.limit_exceeded`), ClickHouse down (`503 clickhouse.unavailable`, `Retry-After: 5`), … — see [ClickHouse errors on the query paths](#clickhouse-errors-on-the-query-paths) | | 500 | `{"error":"…","code":"clickhouse.unknown","retryable":true}` | A failure with no verdict | | 503 | `{"error":"schema not loaded yet"}` | The tenant's first schema discovery has not succeeded yet, so whether the table exists is not known; `Retry-After: 5` | -| 503 | `{"error":"no ClickHouse connection is open for this tenant"}` | The tenant is on no ClickHouse pool — [no pool could be opened for it](/settings-directory#clickhouse), such as one the connection ceiling refused — so the query cannot run; decided before anything is served, so nothing cached before is served either; `Retry-After: 30`, a settings reload retries the pool | +| 503 | `{"error":"no ClickHouse connection is open for this tenant"}` | The tenant is on no ClickHouse pool — [no pool could be opened for it](/settings-directory#clickhouse), such as one the connection ceiling refused — so the query cannot run; decided before a cached result is served, so nothing cached before is served either; `Retry-After: 30`, a settings reload retries the pool | | 503 | `{"error":"token verifier not ready: the tenant's JWKS has not been fetched yet"}` | A token was supplied, with no valid operator key, while the tenant's JWKS has not been fetched yet; refused before any policy runs, with a `Retry-After: 30` header — see [Authentication](#authentication) | --- @@ -609,7 +609,7 @@ The POST parameter body is capped at 1 MiB; a body over the cap is rejected with | Status | Body | Cause | | ------ | ---- | ----- | | 404 | `{"error":"pipe not found"}` | Pipe name not registered | -| 503 | `{"error":"no ClickHouse connection is open for this tenant"}` | The tenant is on no ClickHouse pool — [no pool could be opened for it](/settings-directory#clickhouse), such as one the connection ceiling refused; decided before anything is served, a cached result included; `Retry-After: 30` | +| 503 | `{"error":"no ClickHouse connection is open for this tenant"}` | The tenant is on no ClickHouse pool — [no pool could be opened for it](/settings-directory#clickhouse), such as one the connection ceiling refused; decided before a cached result is served or a query runs; `Retry-After: 30` | | 403 | `{"error":"forbidden"}` | Role not in pipe's `allowed_roles` (and not the admin role). Fails closed: a request with no role (no token, or a JWT missing `auth.role_claim`) is denied unless a `default_role` resolves it into the list; a pipe with no `allowed_roles` denies everyone but the admin role. | | 400 | `{"error":"missing required parameter: x"}` | Required parameter not supplied | | 400 | `{"error":"parameter \"x\": unsupported parameter type object"}` | A non-scalar value with no SQL literal form — a JSON object, whether supplied directly or nested as an array element. A JSON **array** is valid and renders as an `IN`-style `(…)` list. | diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 6a4f83ff8..202111750 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -86,7 +86,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request - **ingest.go** — Accepts `POST /v1/ingest?table={table}` in three body shapes: one flat JSON object, a JSON array of them, or NDJSON. The **required** `Content-Type` chooses the format *family* — `application/json` versus the four NDJSON spellings — and within the JSON family the body's first non-whitespace byte picks array versus single object; the bytes never choose the family. Anything that is not exactly one readable media type is a `415`, decided before the body is read: the header is parsed per RFC 9110 §8.3, and because `Content-Type` is a singleton field, repeated header lines must all resolve to the same format and a value carrying a comma is refused unless the value as a whole parses as one media type — a comma inside a *quoted* parameter value is data, so `application/json; a=", application/x-ndjson; b="` is accepted. It then reads the whole (`MaxBytesReader`-capped) body into a pooled buffer and runs the per-format record readers over those bytes, so the `413` lands before any record is processed and peak memory per request is O(body) rather than O(record). Then it validates each record against the discovered schema, optional dedup, and publishes each row through `mq.Publisher` on `mq.Topic{Tenant, Table, Scope}` (the request's tenant, read off its resolved store — `store.Tenant()` — and raw names; the subject it becomes is `internal/mq`'s; a full queue comes back as `mq.ErrQueueFull`, which is the `503` + `Retry-After: 30`, and a broker that cannot be reached or does not answer in time as `mq.ErrUnavailable`, the `503` + `Retry-After: 5`). When dedup is on, a row missing the configured `id_field` can't be deduped: it is logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total` (labeled by `table`), then published un-deduped — or rejected when `dedupe.require_id` is set ([#219](https://github.com/Wave-RF/WaveHouse/issues/219)). - **query.go** — Proxies raw SQL for `POST /v1/ops/query` straight to the `?tenant=`'s ClickHouse HTTP interface (`chconn.Pools.Target` by the resolved store's tenant; the zero target — no pool — is a `503` with `Retry-After`). **Not cached** — sets `Cache-Control: no-store` so every request hits ClickHouse; DateTime is rendered ISO-8601 via `date_time_output_format=iso` (the Go-side type conversion lives in the structured-query / pipes path, not here). - **stream.go** — Real-time streaming via SSE. Callers select a table with the `?table=` query parameter. Each connection registers one `Subscriber` (the `stream/` package) with both the event `Hub` (under its `(topic, role)`) and the shared keepalive wheel, then drains both from a single byte-pump — so idle streams keep emitting `:` keepalive comments (surviving reverse-proxy idle timeouts) while live events arrive already projected and serialized. Per-event projection/serialization happens **once per role** in the `Hub`, not once per subscriber ([#294](https://github.com/Wave-RF/WaveHouse/issues/294)); the handler also snapshots the connection's JWT claims onto the `Subscriber`, which the `Hub` evaluates per subscriber when the role carries a row-level `filter` ([#319](https://github.com/Wave-RF/WaveHouse/issues/319)). Gap-fill replay (`mq.Replayer.ReplaySince` on the connection's `mq.Topic` — a `DeliverByStartTime` consumer inside `internal/mq`) stays per-connection (low-volume, one-time on connect). A stream ends, a gap-fill in progress included, when the server begins shutting down (`Closing`) or its `Subscriber` is evicted because its tenant is no longer served (`Hub.Prune`); one admitted just before the reload that stopped serving its tenant, and registered just after the prune, is ended right after it registers (`Served`). -- **schema.go** — Schema discovery API of one tenant, the `?tenant=` (`opsStore`): list all schemas, get one table, trigger refresh. `lookupSchema`, shared with the ingest and structured-query handlers, is the one reading of a `SchemaRegistry.Lookup` miss: `503` with `Retry-After` before the tenant's first discovery (`ErrNotLoaded`, or no registry built yet), `404` for a table the discovered schema lacks; the list answers the same `503` rather than `[]`. A refresh of a tenant on no pool (`discovery.ErrNoConnection`) is a `503` with `Retry-After` too. The handlers hold `RegistrySource`, `func(*settings.Store) *discovery.SchemaRegistry`, and the query paths a `func(*settings.Store) driver.Conn` beside it — each resolves the request's tenant per call, and a nil connection (a tenant no pool could be opened for, such as by the connection ceiling) is a `503` before anything is served, a cached result included. The cached paths resolve it after their cache `Lookup`, so the snapshot predates the connection (see `cache.go` below). +- **schema.go** — Schema discovery API of one tenant, the `?tenant=` (`opsStore`): list all schemas, get one table, trigger refresh. `lookupSchema`, shared with the ingest and structured-query handlers, is the one reading of a `SchemaRegistry.Lookup` miss: `503` with `Retry-After` before the tenant's first discovery (`ErrNotLoaded`, or no registry built yet), `404` for a table the discovered schema lacks; the list answers the same `503` rather than `[]`. A refresh of a tenant on no pool (`discovery.ErrNoConnection`) is a `503` with `Retry-After` too. The handlers hold `RegistrySource`, `func(*settings.Store) *discovery.SchemaRegistry`, and the query paths a `func(*settings.Store) driver.Conn` beside it — each resolves the request's tenant per call, and a nil connection (a tenant no pool could be opened for, such as by the connection ceiling) is a `503` before a cached result is served or a query runs. The cached paths resolve it after their cache `Lookup`, so the snapshot predates the connection (see `cache.go` below). - **dlq.go** — DLQ stats endpoint (`GET /v1/ops/dlq/stats`): asks `mq.DeadLetterStats.DeadLetterCounts` for one tenant's per-table parked counts (optionally one table) and its total — the tenant `?tenant=` names, read strictly by `opsTenant`, tenant `0` without it. The tenant is looked up in the MQ, not the settings registry, so a rejected or removed tenant's parked rows are read like a served one's; a tenant with no dead-letter queue (`mq.ErrNoDeadLetterQueue`) is a 404, and any other failure to read it a 500. The queue itself is `internal/mq`'s. - **health.go** — Liveness (`/livez`), readiness (`/readyz`), and a content-free `Online` ping (`/v1/health`, the SDK's public liveness check); `/healthz` is a permanent alias of `/livez`, and `/health`/`/ready` are deprecated aliases. All three consult an optional `BootState` so they can return 503 while boot-time schema discovery is still failing in the retry loop (see `internal/app`; over a nested directory, while no tenant's has succeeded); once `BootState.Set(nil)` fires, `/livez` returns 200 and stays there. `/readyz` additionally runs a `Ping` each call — `chconn.Pools.Ping` in production: every open pool at once, ready at the first answer, every pool's error joined when none answers; `/v1/health` deliberately does not. @@ -112,7 +112,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ ### `cache/` — Query Cache -- **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on — one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)) — each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope by their raw names, which the cache escapes where it builds a key. `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The snapshot is taken before any input a bump invalidates is chosen, the tenant's connection included: a reload that repoints the tenant runs `Pools.Reconcile` and then `InvalidateTenant`, so a request that took the old pool files the old database's rows under a version the bump has already orphaned. The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. +- **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on — one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)) — each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope by their raw names, which the cache escapes where it builds a key. `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The snapshot is taken before any input a bump invalidates is chosen, the tenant's connection included: a reload that moves the tenant to another address or database runs `Pools.Reconcile` and then `InvalidateTenant` (a repoint that keeps both, such as a username or `tls` change, reads the same tables and bumps nothing), so a request that took the old pool files the old database's rows under a version that bump orphans, whether its `Set` lands before the bump or after. The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. - **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: a namespace key is `..
.
.`, its fields joined with `keyenc.Join` so a dot in a table or scope name is escaped rather than read as a separator, and an entry's key, `|.||…` with the caller's query key escaped whole (its `:` become `%3A`), folds in the tenant's version and each dependency's namespace key and namespace version, so bumping a table (a scopeless write) or one scope — scope is reserved and empty today, so every write is the whole-table bump — orphans every dependent entry without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results, and the flat directory's single tenant simply carries the `0` prefix. `BumpTenant` (behind `InvalidateTenant`) advances the tenant version that leads every namespace key of one tenant, orphaning its every namespace and every cached query in one step (the tenant version is folded into every query key, so a pipe result with no dependencies is orphaned too) — a table no bump ever keyed included, which an enumeration of the index would miss — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. diff --git a/internal/api/cache_tenant_test.go b/internal/api/cache_tenant_test.go index 38ea67267..3b335e1a3 100644 --- a/internal/api/cache_tenant_test.go +++ b/internal/api/cache_tenant_test.go @@ -263,10 +263,10 @@ func TestCachedRoutes_BumpDuringQueryOrphansTheFill(t *testing.T) { } // The snapshot is taken before the tenant's pool is chosen. A reload that -// repoints the tenant — Pools.Reconcile, then InvalidateTenant — landing -// between the two leaves the request on the old pool: its fill, read from -// the old database, is orphaned by the bump rather than filed as fresh under -// the new tenant version. And a tenant on no pool is a 503 even when its +// moves the tenant to another address or database — Pools.Reconcile, then +// InvalidateTenant — landing between the two leaves the request on the old +// pool: its fill, read from the old database, is orphaned by the bump rather +// than filed as fresh under the new tenant version. And a tenant on no pool is a 503 even when its // Lookup hit (#583 story 6). func TestCachedRoutes_ReloadAsThePoolIsTakenOrphansTheFill(t *testing.T) { for _, route := range cachedRoutes { diff --git a/internal/api/pipes.go b/internal/api/pipes.go index 81dd9d710..428d3b8d0 100644 --- a/internal/api/pipes.go +++ b/internal/api/pipes.go @@ -157,8 +157,8 @@ func (h *PipesHandler) Execute(w http.ResponseWriter, r *http.Request) { // the tenant's version alone, so InvalidateTenant orphans it but no insert // does (TTL-bound until #343). The snapshot is of the versions before // anything the query reads is chosen, so a bump landing after — mid-query - // (#382), or a reload repointing the tenant once its pool below is taken — - // orphans the fill. + // (#382), or a reload moving the tenant to another address or database + // once its pool below is taken — orphans the fill. // TODO: once pipes expose their tables/scopes, pass them as deps here so writes // invalidate cached pipe results. cacheKey := queryCacheKey(store.Tenant(), sql, params) diff --git a/internal/api/structured_query.go b/internal/api/structured_query.go index 9b192a0b3..6fb62eeb5 100644 --- a/internal/api/structured_query.go +++ b/internal/api/structured_query.go @@ -172,8 +172,8 @@ func (h *StructuredQueryHandler) Handle(w http.ResponseWriter, r *http.Request) // The snapshot is of the versions before anything the query reads is // chosen, so a bump landing after — an insert mid-query (#382), or a - // reload repointing the tenant once its pool below is taken — orphans the - // fill. + // reload moving the tenant to another address or database once its pool + // below is taken — orphans the fill. var entry cache.Entry var snap cache.Snapshot if h.Cache != nil { diff --git a/internal/api/tenant_clickhouse_test.go b/internal/api/tenant_clickhouse_test.go index faab375ca..8181baa2d 100644 --- a/internal/api/tenant_clickhouse_test.go +++ b/internal/api/tenant_clickhouse_test.go @@ -109,8 +109,8 @@ func selectAllQuery() query.StructuredQuery { return query.StructuredQuery{Selec // A tenant on no pool — its tuple could not be opened, such as by the // connection ceiling — fails closed on every route that reaches its -// ClickHouse: a 503 with Retry-After before anything is served, so nothing -// it cached before is served either (TestCachedRoutes_ReloadAsThePoolIsTakenOrphansTheFill +// ClickHouse: a 503 with Retry-After before a cached result is served or a +// query runs, so nothing it cached before is served either (TestCachedRoutes_ReloadAsThePoolIsTakenOrphansTheFill // pins that with a hit), and on the refresh, which cannot run. func TestClickHouseRoutes_NoPoolIs503(t *testing.T) { t.Parallel() From 8189783eb2ab4972753c27357c3bedfd8351690a Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:08:18 -0400 Subject: [PATCH 39/79] test(cache): the local backend's insert-invalidates lifecycle, end to end With e2e on Redis, no test above unit level showed an ingest invalidating LocalCache, the default backend. The integration suite's own app runs cache.backend=local, so it gets the shared-cache test's twin: a fill is a hit, an ingested row is first served as a miss well inside the stale entry's TTL, and the refill is a hit again. With LocalCache.Invalidate made a no-op it fails: the row landed only ~10.0s after the fill, every round. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- CHANGELOG.md | 2 +- tests/integration/shared_cache_test.go | 42 ++++++++++++++++++++++++++ 2 files changed, 43 insertions(+), 1 deletion(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 1454d47ae..48eb10147 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`, at most `2s`: boot and shutdown each wait out a dial), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. +- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`, at most `2s`: boot and shutdown each wait out a dial), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits; a third runs the first one's hit, insert and fresh-miss lifecycle on the suite's own `cache.backend: local` app, since e2e no longer exercises that backend. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. - **A Redis-compatible shared cache backend** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row open a circuit breaker that skips the server until a probe succeeds, and deferred invalidations are retried until they land, collapsing to one tenant-wide bump per tenant past 100,000 keys. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. `cache.backend: redis` selects it (the entry above). Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression and server-stops-answering cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/tests/integration/shared_cache_test.go b/tests/integration/shared_cache_test.go index f4a884d0e..e9857b725 100644 --- a/tests/integration/shared_cache_test.go +++ b/tests/integration/shared_cache_test.go @@ -216,6 +216,48 @@ func TestSharedCache_IngestOnOneInstanceInvalidatesAnother(t *testing.T) { } } +// The same lifecycle on the default backend, against the suite's own app +// (cache.backend=local), which e2e no longer runs: a fill is a hit, and an +// ingest invalidates it, so the first answer carrying the new row is a miss +// well inside the stale entry's TTL, and the refill is a hit again. +func TestLocalCache_IngestInvalidates(t *testing.T) { + base := env(t).baseURL + // As above: a round whose row lands past the TTL floor proves nothing, + // and runs again on a fresh table, so a fresh fill. + for round := 1; ; round++ { + table := createTable(t, "user_id String, value Float64", "ORDER BY user_id") + status, xc, body := structuredQuery(t, base, table) + require.Equal(t, http.StatusOK, status, body) + require.Equal(t, "MISS", xc) + filled := time.Now() + status, xc, _ = structuredQuery(t, base, table) + require.Equal(t, http.StatusOK, status) + require.Equal(t, "HIT", xc) + + ingestRow(t, base, table, "u1") + var seenAt time.Time + require.Eventually(t, func() bool { + var err error + status, xc, body, err = tryStructuredQuery(base, table) + if err == nil && status == http.StatusOK && strings.Contains(body, "u1") { + seenAt = time.Now() + return true + } + return false + }, 30*time.Second, 100*time.Millisecond, "the ingested row is never served") + if seenAt.Sub(filled) < minCacheTTL-time.Second { + assert.Equal(t, "MISS", xc, "the first answer carrying the new row is a refill") + status, xc, body = structuredQuery(t, base, table) + require.Equal(t, http.StatusOK, status) + assert.Equal(t, "HIT", xc, "the refill is cached again") + assert.Contains(t, body, "u1") + return + } + require.Less(t, round, 3, "the new row was served only once the stale entry could have expired, in every round: the ingest never invalidated it, or ingest is too slow here to tell") + t.Logf("round %d: row landed %s after the fill, past the TTL floor; retrying", round, seenAt.Sub(filled)) + } +} + // A Redis that stops answering costs queries nothing but the cache: they // keep succeeding, straight from ClickHouse, each a miss; an ingest made // meanwhile is visible at once. Once it answers again, the cache serves hits. From 5e8476271d03db7fc5ca51aba68e3a48db5353c4 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:10:05 -0400 Subject: [PATCH 40/79] fix(ingest): log an invalidation that did not land at WARN With cache.backend=redis a bump the server does not take is deferred and retried, so a Redis outage logged an ERROR for every batch the worker inserted. It is a WARN now; wavehouse_cache_invalidations_pending is the signal to alert on. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- CHANGELOG.md | 2 +- internal/ingest/worker.go | 4 +++- 2 files changed, 4 insertions(+), 2 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 48eb10147..27ee2eaf7 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`, at most `2s`: boot and shutdown each wait out a dial), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits; a third runs the first one's hit, insert and fresh-miss lifecycle on the suite's own `cache.backend: local` app, since e2e no longer exercises that backend. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. +- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `internal/ingest/worker.go`, `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`, at most `2s`: boot and shutdown each wait out a dial), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The ingest worker's log of an invalidation that did not land drops from `ERROR` to `WARN`, since the shared backend defers and retries it: an outage would otherwise log an `ERROR` for every batch. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits; a third runs the first one's hit, insert and fresh-miss lifecycle on the suite's own `cache.backend: local` app, since e2e no longer exercises that backend. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. - **A Redis-compatible shared cache backend** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row open a circuit breaker that skips the server until a probe succeeds, and deferred invalidations are retried until they land, collapsing to one tenant-wide bump per tenant past 100,000 keys. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. `cache.backend: redis` selects it (the entry above). Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression and server-stops-answering cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/internal/ingest/worker.go b/internal/ingest/worker.go index 70517cd2e..529b8485e 100644 --- a/internal/ingest/worker.go +++ b/internal/ingest/worker.go @@ -910,7 +910,9 @@ func (w *IngestWorker) invalidate(ctx context.Context, id tenant.ID, tableName s } invCtx := trace.ContextWithSpanContext(context.WithoutCancel(ctx), trace.SpanContextFromContext(ctx)) if _, err := w.cache.Invalidate(invCtx, namespaces); err != nil { - slog.ErrorContext(invCtx, "failed to invalidate cache after insert - your cache is holding stale data now!", "tenant", id, "table", tableName, "error", err) + // WARN, not ERROR: a shared backend defers and retries the bump, and + // an outage would otherwise log an ERROR for every batch. + slog.WarnContext(invCtx, "cache invalidation after insert did not land; the table's cached results may be stale until it does", "tenant", id, "table", tableName, "error", err) } } From 0696702c730393a9a18795cadc8b5d0532744243 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:18:13 -0400 Subject: [PATCH 41/79] fix(cache): owed bumps hold lookups; refused writes open the breaker A process that owes a version bump no longer serves what the bump would orphan: a lookup reading a token key still pending is a bypass with a zero snapshot, so it neither hits nor fills. An atomic count keeps the check free while nothing is owed, and the tenant token a pending set collapses to past PendingMax is one of every lookup's keys, so the collapse holds the whole tenant. An error reply saying the server takes no writes (READONLY from a demoted primary, OOM under noeviction, MASTERDOWN, NOREPLICAS, MISCONF, LOADING, BUSY, CLUSTERDOWN) opens the breaker at once instead of counting as the server being up. The probe is now a write, so a server that answers but refuses writes stays bypassed, and only the probe's success closes an open breaker. This is the READONLY/OOM half of #656. The first deferral wakes the drain at once, and while the breaker is open the drain wakes when the probe is due rather than at its own backoff (up to 10 s), so a process that makes no lookups recovers as soon as one that does. Invalidate sends its bumps in batches of 1,000, as the drain does. The zstd encoder window drops to 1 MiB: its pool kept about 121 MiB at 14 procs with 240 KB values, now about 36 MiB, at the same speed and ratio. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 8 +- internal/cache/breaker.go | 37 ++++- internal/cache/breaker_test.go | 38 ++++++ internal/cache/export_test.go | 3 + internal/cache/metrics.go | 4 +- internal/cache/pending.go | 32 +++-- internal/cache/pending_test.go | 26 ++++ internal/cache/redis.go | 132 +++++++++++++----- internal/cache/redis_codec.go | 10 +- internal/cache/redis_codec_test.go | 35 +++++ internal/cache/redis_integration_test.go | 164 +++++++++++++++++++++++ internal/cache/redis_test.go | 45 +++++++ 13 files changed, 479 insertions(+), 57 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index d02ae4ad3..b21658065 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A Redis-compatible shared cache backend, built but not yet selectable** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row open a circuit breaker that skips the server until a probe succeeds, and deferred invalidations are retried until they land, collapsing to one tenant-wide bump per tenant past 100,000 keys. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. Nothing selects it yet: the `cache.backend` switch and the `cache.redis.*` settings arrive with the wiring (E4), so every deployment still runs `LocalCache`. Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression and server-stops-answering cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. +- **A Redis-compatible shared cache backend, built but not yet selectable** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), open a circuit breaker that skips the server until a probe write succeeds, and deferred invalidations are retried until they land — the first at once, and at the probe's cadence while the breaker is open — collapsing to one tenant-wide bump per tenant past 100,000 keys; until one lands, the process that owes it bypasses the lookups it would orphan. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. Nothing selects it yet: the `cache.backend` switch and the `cache.redis.*` settings arrive with the wiring (E4), so every deployment still runs `LocalCache`. Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression, server-stops-answering, refused-writes (a demoted primary, a full `noeviction` server) and owed-invalidation cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (since #614 this drops the tenant's cached pipe results too; no insert invalidates a pipe result, which names no table). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index e18ace9c1..f124f8d7c 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -113,11 +113,11 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on — one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)) — each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope by their raw names, which the cache escapes where it builds a key. `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. - **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: a namespace key is `..
.
.`, its fields joined with `keyenc.Join` so a dot in a table or scope name is escaped rather than read as a separator, and an entry's key, `|.||…` with the caller's query key escaped whole (its `:` become `%3A`), folds in the tenant's version and each dependency's namespace key and namespace version, so bumping a table (a scopeless write) or one scope — scope is reserved and empty today, so every write is the whole-table bump — orphans every dependent entry without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results, and the flat directory's single tenant simply carries the `0` prefix. `BumpTenant` (behind `InvalidateTenant`) advances the tenant version that leads every namespace key of one tenant, orphaning its every namespace and every cached query in one step (the tenant version is folded into every query key, so a pipe result with no dependencies is orphaned too) — a table no bump ever keyed included, which an enumeration of the index would miss — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. -- **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET`, `MGET` and `PING`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot; the table and scope are escaped with `keyenc` after the fixed prefix, so a `:` in a name is never read as the separator. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, the hash over their escaped forms, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token (a string of another length, a list, a hash) is replaced, so a token lost to eviction, expiry or a restart is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing. Nothing selects this backend yet — the boot config and wiring arrive with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s E4. +- **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET` and `MGET`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot; the table and scope are escaped with `keyenc` after the fixed prefix, so a `:` in a name is never read as the separator. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, the hash over their escaped forms, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token (a string of another length, a list, a hash) is replaced, so a token lost to eviction, expiry or a restart is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. While this process owes a bump on any of a lookup's tokens, that lookup is a bypass that files nothing: the bump would orphan whatever it found (other processes, which cannot know, serve those entries until it lands). `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing. Nothing selects this backend yet — the boot config and wiring arrive with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s E4. - **redis_codec.go** — the key schema (`tokenKeys`, `bumpKeys`, `valueKey`, which take a `Namespace`'s raw names and escape them) and the value frame: format, flags, expiry, the token list, the payload, zstd-compressed from `CompressMinBytes` when that is smaller. A stored value is capped at `MaxValueBytes` and a decoded one at eight times that, which refuses a zip bomb planted in a shared server; an unknown format, as a newer process writes during a rolling upgrade, is a miss. -- **breaker.go** — the circuit breaker: `BreakerThreshold` (5) consecutive failures open it, and while open every operation skips the server at once; after `BreakerOpenFor` (5 s) one background `PING` decides whether it closes. Only a transport failure or a timeout counts against the server: any reply, an error reply or one the backend cannot use included, counts as a success, for operations and the probe alike, and a caller that gave up first counts as nothing. -- **pending.go** — invalidations the server did not take, kept per token key (repeats coalesce) and retried by a background loop with backoff from 100 ms to 10 s until they land. Landing late is still correct: a fresh token orphans the pre-write entries and any fill made meanwhile. Past `PendingMax` (100,000) keys the set collapses to one tenant bump per affected tenant. `Close` makes one last attempt at the bumps still owed, past the breaker (an open one is why they are owed); what that attempt cannot deliver is lost, and the entries they would orphan are served until their TTL — the same failure as a worker stopping between an insert and its invalidation. -- **metrics.go** — the shared backend's instruments (meter `wavehouse-cache`, `backend="redis"`, no tenant attribute): `wavehouse_cache_lookups_total` by `result` (`hit`, `miss`, `stale` — filed under since-bumped tokens, `bypass` — server skipped, `error`), `wavehouse_cache_op_duration_seconds` by `op` (`lookup`, `set`, `invalidate`), `wavehouse_cache_breaker_open` (1 while bypassed, including before the first connection), `wavehouse_cache_invalidations_total` by `result` (`ok`, `deferred`), `wavehouse_cache_invalidations_pending` (bumps owed: entries they would orphan may be served stale meanwhile), `wavehouse_cache_value_bytes` (stored size), `wavehouse_cache_oversize_total` and `wavehouse_cache_set_failures_total` by `reason` (`oom`, `timeout`, `other`). +- **breaker.go** — the circuit breaker: `BreakerThreshold` (5) consecutive failures open it, and while open every operation skips the server at once; after `BreakerOpenFor` (5 s) one background write (`SET :probe`) decides whether it closes, and only that probe's success closes it. A transport failure or a timeout counts against the server; an error reply saying it takes no writes right now — `READONLY` (a demoted primary), `MASTERDOWN`, `OOM` (full under `maxmemory-policy noeviction`, Redis's default, so give the cache an evicting policy), `NOREPLICAS`, `MISCONF`, `LOADING`, `BUSY`, `CLUSTERDOWN` — opens it at once, whatever the threshold, since no bump can land. Any other reply, an error reply about one key (`WRONGTYPE`, `NOPERM`, `TRYAGAIN`) or one the backend cannot use included, counts as a success, and a caller that gave up first counts as nothing. +- **pending.go** — invalidations the server did not take, kept per token key (repeats coalesce) and retried by a background loop until they land: the first one owed at once, then with backoff from 100 ms to 10 s, and while the breaker is open at each probe, which the loop starts when due, so a process that makes no lookups (ingest only) recovers as soon as one that does. `Invalidate` sends its bumps 1,000 to a round trip, as the retry does. Landing late is still correct: a fresh token orphans the pre-write entries and any fill made meanwhile. Past `PendingMax` (100,000) keys the set collapses to one tenant bump per affected tenant. `Close` makes one last attempt at the bumps still owed, past the breaker (an open one is why they are owed); what that attempt cannot deliver is lost, and the entries they would orphan are served until their TTL — the same failure as a worker stopping between an insert and its invalidation. +- **metrics.go** — the shared backend's instruments (meter `wavehouse-cache`, `backend="redis"`, no tenant attribute): `wavehouse_cache_lookups_total` by `result` (`hit`, `miss`, `stale` — filed under since-bumped tokens, `bypass` — server skipped, or held by a bump this process owes, `error`), `wavehouse_cache_op_duration_seconds` by `op` (`lookup`, `set`, `invalidate`), `wavehouse_cache_breaker_open` (1 while bypassed, including before the first connection), `wavehouse_cache_invalidations_total` by `result` (`ok`, `deferred`), `wavehouse_cache_invalidations_pending` (bumps owed: entries they would orphan may be served stale meanwhile), `wavehouse_cache_value_bytes` (stored size), `wavehouse_cache_oversize_total` and `wavehouse_cache_set_failures_total` by `reason` (`oom`, `timeout`, `other`). ### `config/` — Configuration diff --git a/internal/cache/breaker.go b/internal/cache/breaker.go index c90ab3a91..a33c94dae 100644 --- a/internal/cache/breaker.go +++ b/internal/cache/breaker.go @@ -6,9 +6,10 @@ import ( ) // breaker stops a failing cache server from costing every request its full -// timeout: after threshold consecutive failures it opens, and while open -// callers skip the server entirely. Once openFor has passed, one caller is -// told to probe; the probe's outcome closes the breaker or reopens it. +// timeout: after threshold consecutive failures it opens — at once, for a +// reply refusing the work — and while open callers skip the server +// entirely. Once openFor has passed, one caller is told to probe; the +// probe's outcome closes the breaker or reopens it. type breaker struct { threshold int openFor time.Duration @@ -40,10 +41,15 @@ func (b *breaker) allow() (ok, probe bool) { return false, true } -// success records a call the server answered, and closes the breaker. +// success records a call the server answered. Only a probe's closes an +// open breaker: any other set out before it opened, and a server that +// answers reads may still be refusing writes. func (b *breaker) success() { b.mu.Lock() defer b.mu.Unlock() + if b.open && !b.probing { + return + } b.failures, b.open, b.probing = 0, false, false } @@ -58,8 +64,31 @@ func (b *breaker) failure() { } } +// trip opens the breaker at once, for a reply that says the server cannot +// do the work: one is as conclusive as any number. +func (b *breaker) trip() { + b.mu.Lock() + defer b.mu.Unlock() + b.open, b.openedAt, b.probing = true, b.now(), false +} + func (b *breaker) isOpen() bool { b.mu.Lock() defer b.mu.Unlock() return b.open } + +// untilProbe reports how long until an open breaker is due its probe, and +// whether it is open. While a probe runs it reports openFor: the probe ends +// by closing the breaker or by opening it afresh. +func (b *breaker) untilProbe() (time.Duration, bool) { + b.mu.Lock() + defer b.mu.Unlock() + switch { + case !b.open: + return 0, false + case b.probing: + return b.openFor, true + } + return b.openedAt.Add(b.openFor).Sub(b.now()), true +} diff --git a/internal/cache/breaker_test.go b/internal/cache/breaker_test.go index e6dc2296c..3b2109d96 100644 --- a/internal/cache/breaker_test.go +++ b/internal/cache/breaker_test.go @@ -5,6 +5,7 @@ import ( "time" "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" ) type fakeClock struct{ t time.Time } @@ -57,3 +58,40 @@ func TestBreaker(t *testing.T) { assert.True(t, ok) assert.False(t, probe) } + +// A reply refusing the work opens the breaker at once, and only the probe's +// success closes it: a call that set out before it opened proves nothing. +func TestBreaker_TripAndProbeSchedule(t *testing.T) { + t.Parallel() + clock := &fakeClock{t: time.Unix(0, 0)} + b := newBreaker(100, 5*time.Second, clock.now) + _, open := b.untilProbe() + assert.False(t, open) + + b.trip() + assert.True(t, b.isOpen(), "no threshold for a refusal") + b.success() + assert.True(t, b.isOpen(), "a success that is not the probe's leaves it open") + d, open := b.untilProbe() + assert.True(t, open) + assert.Equal(t, 5*time.Second, d) + + clock.t = clock.t.Add(2 * time.Second) + d, _ = b.untilProbe() + assert.Equal(t, 3*time.Second, d) + + clock.t = clock.t.Add(3 * time.Second) + _, probe := b.allow() + require.True(t, probe) + d, _ = b.untilProbe() + assert.Equal(t, 5*time.Second, d, "while the probe runs, wait out a whole period") + b.trip() // the probe was refused too + d, _ = b.untilProbe() + assert.Equal(t, 5*time.Second, d) + + clock.t = clock.t.Add(5 * time.Second) + _, probe = b.allow() + require.True(t, probe) + b.success() + assert.False(t, b.isOpen()) +} diff --git a/internal/cache/export_test.go b/internal/cache/export_test.go index 7c392c13e..aaf712f7f 100644 --- a/internal/cache/export_test.go +++ b/internal/cache/export_test.go @@ -8,5 +8,8 @@ func Pending(r *RedisCache) int { return r.pending.len() } // Bypassed reports whether r is skipping the server. func Bypassed(r *RedisCache) bool { return r.bypassed() } +// ZeroSnapshot reports whether s files nothing. +func ZeroSnapshot(s Snapshot) bool { return s.key == "" && s.tokens == nil } + // DecodedFactor is how many times MaxValueBytes a value may decompress to. const DecodedFactor = decodedFactor diff --git a/internal/cache/metrics.go b/internal/cache/metrics.go index cf76eb7b3..34539f13b 100644 --- a/internal/cache/metrics.go +++ b/internal/cache/metrics.go @@ -15,7 +15,7 @@ const ( resultHit = "hit" resultMiss = "miss" // nothing stored resultStale = "stale" // stored under versions since bumped - resultBypass = "bypass" // server skipped: breaker open or not yet connected + resultBypass = "bypass" // server skipped: breaker open, not yet connected, or a bump this process owes would orphan the entry resultError = "error" // server failed or timed out ) @@ -40,7 +40,7 @@ func newMetrics(backend string, breakerOpen func() bool, pending func() int) (*m m := &metrics{backend: attribute.String("backend", backend)} var errs [9]error m.lookups, errs[0] = meter.Int64Counter("wavehouse_cache_lookups_total", - metric.WithDescription("Shared-cache lookups by result: hit, miss, stale (stored under since-bumped versions), bypass (server skipped), error")) + metric.WithDescription("Shared-cache lookups by result: hit, miss, stale (stored under since-bumped versions), bypass (server skipped, or held by an invalidation this process has yet to deliver), error")) m.duration, errs[1] = meter.Float64Histogram("wavehouse_cache_op_duration_seconds", metric.WithDescription("Shared-cache round-trip time by op: lookup, set, invalidate"), metric.WithUnit("s"), metric.WithExplicitBucketBoundaries(.0001, .00025, .0005, .001, .0025, .005, .01, .025, .05, .1, .25)) diff --git a/internal/cache/pending.go b/internal/cache/pending.go index c19b21786..481ab9ac7 100644 --- a/internal/cache/pending.go +++ b/internal/cache/pending.go @@ -2,6 +2,7 @@ package cache import ( "sync" + "sync/atomic" "github.com/Wave-RF/WaveHouse/internal/tenant" ) @@ -18,6 +19,7 @@ type pendingBumps struct { mu sync.Mutex keys map[string]pendingKey gen uint64 + n atomic.Int64 // len(keys), read without mu on every lookup } type pendingKey struct { @@ -37,14 +39,14 @@ func (p *pendingBumps) add(id tenant.ID, keys ...string) { for _, k := range keys { p.keys[k] = pendingKey{tenant: id, gen: p.gen} } - if len(p.keys) <= p.max { - return - } - collapsed := make(map[string]pendingKey, len(p.keys)) - for _, pk := range p.keys { - collapsed[tenantTokenKey(p.prefix, pk.tenant)] = pendingKey{tenant: pk.tenant, gen: p.gen} + if len(p.keys) > p.max { + collapsed := make(map[string]pendingKey, len(p.keys)) + for _, pk := range p.keys { + collapsed[tenantTokenKey(p.prefix, pk.tenant)] = pendingKey{tenant: pk.tenant, gen: p.gen} + } + p.keys = collapsed } - p.keys = collapsed + p.n.Store(int64(len(p.keys))) } // snapshot returns the keys owed a bump with the generation each was added @@ -69,10 +71,22 @@ func (p *pendingBumps) done(landed map[string]uint64) { delete(p.keys, k) } } + p.n.Store(int64(len(p.keys))) } -func (p *pendingBumps) len() int { +// owesAny reports whether any of keys is owed a bump. +func (p *pendingBumps) owesAny(keys []string) bool { + if p.n.Load() == 0 { + return false + } p.mu.Lock() defer p.mu.Unlock() - return len(p.keys) + for _, k := range keys { + if _, ok := p.keys[k]; ok { + return true + } + } + return false } + +func (p *pendingBumps) len() int { return int(p.n.Load()) } diff --git a/internal/cache/pending_test.go b/internal/cache/pending_test.go index cc8f331da..1c8390896 100644 --- a/internal/cache/pending_test.go +++ b/internal/cache/pending_test.go @@ -43,3 +43,29 @@ func TestPendingBumps_OverflowCollapsesToTenants(t *testing.T) { assert.Contains(t, got, "wh:{acme}:T") assert.Contains(t, got, "wh:{globex}:T") } + +// A lookup is held by a bump owed on any of its token keys, including the +// tenant token a set past its maximum collapses to. +func TestPendingBumps_OwesAny(t *testing.T) { + t.Parallel() + events := tokenKeys("wh", "acme", []Namespace{{Tenant: "acme", Table: "events", Scope: "org_1"}}) + orders := tokenKeys("wh", "acme", []Namespace{{Tenant: "acme", Table: "orders"}}) + globex := tokenKeys("wh", "globex", []Namespace{{Tenant: "globex", Table: "events"}}) + p := newPendingBumps("wh", 3) + assert.False(t, p.owesAny(events)) + + p.add("acme", bumpKeys("wh", Namespace{Tenant: "acme", Table: "events", Scope: "org_1"})...) + assert.True(t, p.owesAny(events)) + assert.False(t, p.owesAny(orders), "another table's lookups are not held") + assert.False(t, p.owesAny(globex)) + + p.add("acme", "wh:{acme}:B:a", "wh:{acme}:B:b") // past the maximum + assert.Equal(t, map[string]uint64{"wh:{acme}:T": 2}, p.snapshot()) + assert.True(t, p.owesAny(orders), "collapsed to the tenant token, which every lookup of the tenant reads") + assert.True(t, p.owesAny(tokenKeys("wh", "acme", nil))) + assert.False(t, p.owesAny(globex)) + + p.done(p.snapshot()) + assert.False(t, p.owesAny(events)) + assert.Zero(t, p.len()) +} diff --git a/internal/cache/redis.go b/internal/cache/redis.go index 01c6a6c41..ec24893d1 100644 --- a/internal/cache/redis.go +++ b/internal/cache/redis.go @@ -7,7 +7,9 @@ import ( "errors" "fmt" "log/slog" + "maps" "net" + "slices" "strings" "sync" "sync/atomic" @@ -163,8 +165,8 @@ func (c RedisConfig) clientOption() rueidis.ClientOption { } // RedisCache is a Cache shared by every process pointed at one Redis — -// or Valkey, Dragonfly, ElastiCache, MemoryDB: it uses only GET, SET, MGET -// and PING, no scripts and no client tracking. +// or Valkey, Dragonfly, ElastiCache, MemoryDB: it uses only GET, SET and +// MGET, no scripts and no client tracking. // // Versions are random tokens, one per tenant, per table and per scope, // under the tenant's hash tag; a bump sets a fresh one. A value carries the @@ -291,10 +293,12 @@ func (r *RedisCache) conn() rueidis.Client { return *cp } +// probe decides whether an open breaker closes. It writes: a server that +// answers but refuses writes (refusesWork) would take no bump either. func (r *RedisCache) probe(c rueidis.Client) { ctx, cancel := context.WithTimeout(r.ctx, r.cfg.Timeout) defer cancel() - r.record(r.ctx, c.Do(ctx, c.B().Ping().Build()).Error()) + r.record(r.ctx, c.Do(ctx, c.B().Set().Key(r.cfg.KeyPrefix+":probe").Value("1").Ex(time.Minute).Build()).Error()) if r.breaker.isOpen() { return } @@ -308,14 +312,23 @@ func (r *RedisCache) bypassed() bool { } // record feeds an operation's outcome to the breaker. A reply from the -// server, even an error reply or one this code cannot use, shows it is up; a caller that gave up first -// shows nothing about it. +// server, even an error reply or one this code cannot use, shows it is up — +// unless it refuses the work outright, which opens the breaker at once. A +// caller that gave up first shows nothing about it. func (r *RedisCache) record(parent context.Context, err error) { if err == nil || rueidis.IsRedisNil(err) { r.breaker.success() return } - if _, ok := rueidis.IsRedisErr(err); ok || errors.Is(err, errMalformedReply) { + if re, ok := rueidis.IsRedisErr(err); ok { + if refusesWork(re.Error()) { + r.breaker.trip() + } else { + r.breaker.success() + } + return + } + if errors.Is(err, errMalformedReply) { r.breaker.success() return } @@ -325,6 +338,23 @@ func (r *RedisCache) record(parent context.Context, err error) { r.breaker.failure() } +// refusesWork reports whether an error reply says the server takes no +// writes from anyone right now, so no bump can land: a replica (READONLY, +// or MASTERDOWN, which refuses reads too), memory full under noeviction +// (OOM), writes stopped by min-replicas-to-write (NOREPLICAS) or a failed +// snapshot (MISCONF), a dataset still loading (LOADING), a script holding +// the server (BUSY), a cluster not serving the slot (CLUSTERDOWN). Replies +// about one key or one moment — WRONGTYPE, NOPERM, TRYAGAIN during a slot +// migration — are not. +func refusesWork(msg string) bool { + code, _, _ := strings.Cut(msg, " ") + switch code { + case "READONLY", "MASTERDOWN", "OOM", "NOREPLICAS", "MISCONF", "LOADING", "BUSY", "CLUSTERDOWN": + return true + } + return false +} + // Lookup reads the tokens deps fold and the entry for sha in one pipelined // round trip. A token that does not exist yet is created, never read as a // value, so its first use is a miss. @@ -338,6 +368,16 @@ func (r *RedisCache) Lookup(ctx context.Context, id tenant.ID, sha string, deps if len(keys) > maxTokenKeys { return Entry{}, Snapshot{}, fmt.Errorf("cache: %d dependencies is more than a value can record", len(deps)) } + // A bump this process owes would orphan what a lookup reading its key + // finds, so that lookup is a bypass: no hit, and a zero snapshot, so no + // fill either. A lookup's token keys are exactly those whose bumps orphan + // its entry, and include the tenant token a set past PendingMax collapses + // to. Other processes cannot know what this one owes, and serve those + // entries until the bump lands. + if r.pending.owesAny(keys) { + r.metrics.lookup(resultBypass) + return Entry{}, Snapshot{}, nil + } c := r.conn() if c == nil { r.metrics.lookup(resultBypass) @@ -540,20 +580,19 @@ func (r *RedisCache) bump(ctx context.Context, owner map[string]tenant.ID) error return fmt.Errorf("%w: %d invalidations deferred", errBypassed, len(owner)) } defer r.metrics.op("invalidate", time.Now()) - opCtx, cancel := context.WithTimeout(ctx, r.cfg.Timeout) - defer cancel() - keys := make([]string, 0, len(owner)) - cmds := make(rueidis.Commands, 0, len(owner)) - for k := range owner { - keys = append(keys, k) - cmds = append(cmds, r.bumpCmd(c, k)) - } failed := map[string]tenant.ID{} var firstErr error - for i, rr := range c.DoMulti(opCtx, cmds...) { - if err := rr.Error(); err != nil { - failed[keys[i]] = owner[keys[i]] - firstErr = cmpOr(firstErr, err) + // In batches, so a wide fan-out is several round trips each within the + // timeout; past a failure the rest are deferred unsent. + for batch := range slices.Chunk(slices.Collect(maps.Keys(owner)), drainBatch) { + var landed []bool + if firstErr == nil { + landed, firstErr = r.sendBumps(ctx, c, batch) + } + for i, k := range batch { + if landed == nil || !landed[i] { + failed[k] = owner[k] + } } } r.record(ctx, firstErr) @@ -570,11 +609,36 @@ func (r *RedisCache) bumpCmd(c rueidis.Client, key string) rueidis.Completed { return c.B().Set().Key(key).Value(rueidis.BinaryString(tok)).Ex(jitter(r.cfg.VersionTTL, tok)).Build() } +// sendBumps sets a fresh token under each key in one pipeline, within the +// op timeout, reporting which landed and the first error. +func (r *RedisCache) sendBumps(ctx context.Context, c rueidis.Client, keys []string) (landed []bool, firstErr error) { + cmds := make(rueidis.Commands, 0, len(keys)) + for _, k := range keys { + cmds = append(cmds, r.bumpCmd(c, k)) + } + opCtx, cancel := context.WithTimeout(ctx, r.cfg.Timeout) + defer cancel() + landed = make([]bool, len(keys)) + for i, rr := range c.DoMulti(opCtx, cmds...) { + err := rr.Error() + landed[i] = err == nil + firstErr = cmpOr(firstErr, err) + } + return landed, firstErr +} + func (r *RedisCache) deferBumps(owner map[string]tenant.ID) { + first := r.pending.len() == 0 for k, id := range owner { r.pending.add(id, k) } r.metrics.invalidated("deferred", len(owner)) + // The first bump owed wakes the drain now, not at its next tick: it is + // retried at once, or, past an open breaker, when the probe is due. + // Later ones join the retry already backing off. + if first { + r.nudge() + } } // nudge wakes the drain loop now, rather than at its next tick. @@ -600,7 +664,7 @@ func (r *RedisCache) drainLoop() { } wait := drainIdle if r.breaker.isOpen() { - r.conn() // starts the probe when due, so an idle process recovers too + r.conn() // starts the probe when due } if r.pending.len() > 0 { if r.drain(r.ctx, r.conn) { @@ -609,6 +673,12 @@ func (r *RedisCache) drainLoop() { wait, backoff = backoff, min(backoff*2, drainMaxBackoff) } } + // Lookups start the probe that closes an open breaker; a process + // with none (ingest only) has this loop, which wakes when the probe + // is due rather than at the drain's backoff, so it recovers as soon. + if d, open := r.breaker.untilProbe(); open { + wait = min(wait, d) + } timer.Reset(wait) } } @@ -621,32 +691,22 @@ func (r *RedisCache) drain(ctx context.Context, conn func() rueidis.Client) bool for k := range owed { keys = append(keys, k) } - for len(keys) > 0 { - batch := keys[:min(drainBatch, len(keys))] - keys = keys[len(batch):] + for batch := range slices.Chunk(keys, drainBatch) { c := conn() if c == nil { return false } - cmds := make(rueidis.Commands, 0, len(batch)) - for _, k := range batch { - cmds = append(cmds, r.bumpCmd(c, k)) - } - opCtx, cancel := context.WithTimeout(ctx, r.cfg.Timeout) + sent, err := r.sendBumps(ctx, c, batch) landed := map[string]uint64{} - var firstErr error - for i, rr := range c.DoMulti(opCtx, cmds...) { - if err := rr.Error(); err != nil { - firstErr = cmpOr(firstErr, err) - continue + for i, k := range batch { + if sent[i] { + landed[k] = owed[k] } - landed[batch[i]] = owed[batch[i]] } - cancel() - r.record(ctx, firstErr) + r.record(ctx, err) r.pending.done(landed) r.metrics.invalidated("ok", len(landed)) - if firstErr != nil { + if err != nil { return false } } diff --git a/internal/cache/redis_codec.go b/internal/cache/redis_codec.go index 96cddb452..bb9a853ef 100644 --- a/internal/cache/redis_codec.go +++ b/internal/cache/redis_codec.go @@ -26,6 +26,8 @@ const tokenLen = 8 // stored-size limit, refusing a zip bomb planted in a shared server. const decodedFactor = 8 +const encoderWindow = 1 << 20 + // Value layout: format, flags, expires-at (unix ms), token count, tokens, // payload. Big-endian. const ( @@ -122,6 +124,12 @@ func valueKey(prefix string, id tenant.ID, sha string, deps []Namespace) string // codec compresses and frames values. Its zstd encoder and decoder are safe // for concurrent EncodeAll/DecodeAll. +// +// The encoder keeps one window-sized history per concurrent caller for the +// life of the process; 1 MiB, not SpeedFastest's 4, cuts that about +// threefold at no measurable cost in speed or ratio on row payloads. The +// decoder takes any window up to maxDecoded, so a value written with a +// larger one still reads. type codec struct { enc *zstd.Encoder dec *zstd.Decoder @@ -130,7 +138,7 @@ type codec struct { } func newCodec(compressMin, maxDecoded int) (*codec, error) { - enc, err := zstd.NewWriter(nil, zstd.WithEncoderLevel(zstd.SpeedFastest)) + enc, err := zstd.NewWriter(nil, zstd.WithEncoderLevel(zstd.SpeedFastest), zstd.WithWindowSize(encoderWindow)) if err != nil { return nil, err } diff --git a/internal/cache/redis_codec_test.go b/internal/cache/redis_codec_test.go index c956a6097..a61d84dd4 100644 --- a/internal/cache/redis_codec_test.go +++ b/internal/cache/redis_codec_test.go @@ -6,6 +6,7 @@ import ( "testing" "time" + "github.com/klauspost/compress/zstd" "github.com/stretchr/testify/assert" "github.com/stretchr/testify/require" ) @@ -192,6 +193,40 @@ func TestCodec_RefusesBadValues(t *testing.T) { } } +// The encoder's window bounds the history it keeps per concurrent caller, +// and a value an encoder with a wider one wrote (an earlier build's) still +// decodes. +func TestCodec_EncoderWindow(t *testing.T) { + t.Parallel() + window := func(t *testing.T, b []byte) uint64 { + t.Helper() + var h zstd.Header + require.NoError(t, h.Decode(b[headerLen:])) + if h.SingleSegment { + return h.FrameContentSize + } + return h.WindowSize + } + c := newTestCodec(t, 1, 8<<20) + wide, err := zstd.NewWriter(nil, zstd.WithEncoderLevel(zstd.SpeedFastest)) + require.NoError(t, err) + t.Cleanup(func() { _ = wide.Close() }) + earlier := &codec{enc: wide, compressMin: 1} + + for _, n := range []int{2 << 20, 6 << 20} { + payload := bytes.Repeat([]byte(`{"user_id":"u-1","event":"click","value":42.5},`), n/48) + b := c.encode(nil, time.Now(), payload) + require.Equal(t, byte(flagZstd), b[1]) + assert.LessOrEqual(t, window(t, b), uint64(encoderWindow), n) + + b = earlier.encode(nil, time.Now(), payload) + require.Greater(t, window(t, b), uint64(encoderWindow), n) + _, _, got, err := c.decode(b) + require.NoError(t, err, n) + assert.Equal(t, payload, got, n) + } +} + func TestJitter(t *testing.T) { t.Parallel() d := 100 * time.Second diff --git a/internal/cache/redis_integration_test.go b/internal/cache/redis_integration_test.go index 422e372e7..882ed6a52 100644 --- a/internal/cache/redis_integration_test.go +++ b/internal/cache/redis_integration_test.go @@ -7,6 +7,7 @@ import ( "context" "fmt" "net" + "slices" "strconv" "strings" "sync/atomic" @@ -495,3 +496,166 @@ func TestRedis_CloseDeliversPastAnOpenBreaker(t *testing.T) { require.NoError(t, err) assert.Nil(t, e.Value, "the bump Close delivered orphans the fill") } + +// command runs cmd on s through c, for the test to reconfigure the server. +func command(t *testing.T, c rueidis.Client, cmd ...string) { + t.Helper() + require.NoError(t, c.Do(context.Background(), c.B().Arbitrary(cmd[0]).Args(cmd[1:]...).Build()).Error(), "%q", cmd) +} + +// A bump this process owes holds the lookups it would orphan — a bypass +// that files nothing — until it lands, with the breaker closed and every +// other lookup served. The ACL lets the process read the tokens but not +// replace them, so the bump stays owed until the test grants the write. +func TestRedis_OwedBumpHoldsItsLookups(t *testing.T) { + t.Parallel() + ctx := context.Background() + s := startRedis(t) + r := raw(t, s) + prefix := uniquePrefix() + command(t, r, "ACL", "SETUSER", "limited", "on", ">pw", "+@all", "%R~*", "%W~"+prefix+":q:*") + + events := []cache.Namespace{{Tenant: "acme", Table: "events"}} + orders := []cache.Namespace{{Tenant: "acme", Table: "orders"}} + seed := open(t, s, prefix) + for _, deps := range [][]cache.Namespace{events, orders, nil} { + _, snap, err := seed.Lookup(ctx, "acme", "q", deps) + require.NoError(t, err) + require.NoError(t, seed.Set(ctx, snap, []byte("pre-write rows"), time.Minute)) + } + a := open(t, s, prefix, func(c *cache.RedisConfig) { c.Username, c.Password, c.BreakerThreshold = "limited", "pw", 1000 }) + lookup := func(deps []cache.Namespace) (string, cache.Snapshot) { + t.Helper() + e, snap, err := a.Lookup(ctx, "acme", "q", deps) + require.NoError(t, err) + return string(e.Value), snap + } + got, _ := lookup(events) + require.Equal(t, "pre-write rows", got) + + _, err := a.Invalidate(ctx, events) + require.ErrorContains(t, err, "NOPERM") + require.Equal(t, 1, cache.Pending(a)) + require.False(t, cache.Bypassed(a), "a refused key is not a refusing server") + + got, snap := lookup(events) + assert.Empty(t, got, "the owed bump would orphan it") + assert.True(t, cache.ZeroSnapshot(snap), "and a fill under the token it replaces would be orphaned too") + for _, deps := range [][]cache.Namespace{orders, nil} { + got, _ = lookup(deps) + assert.Equal(t, "pre-write rows", got, "%v: lookups the bump does not orphan are served", deps) + } + + command(t, r, "ACL", "SETUSER", "limited", "~*") + deadline := time.Now().Add(10 * time.Second) + for cache.Pending(a) > 0 { + require.True(t, time.Now().Before(deadline), "the bump lands once the server takes it") + got, _ = lookup(events) + require.NotEqual(t, "pre-write rows", got, "served before the owed bump landed") + time.Sleep(time.Millisecond) + } + got, _ = lookup(events) + assert.Empty(t, got, "the landed bump orphaned the pre-write rows") +} + +// A server that answers but refuses writes — a primary demoted to a replica +// (READONLY), memory full under noeviction (OOM) — takes no bump, so its +// first refusal opens the breaker whatever the threshold, and the bump stays +// owed. The probe writes, so it keeps the breaker open until the server +// takes writes again; then the owed bump lands before anything it would +// orphan is served, and a process that only invalidates, whose drain loop +// is all that probes for it, recovers at the probe's cadence too. +func TestRedis_RefusedWrites(t *testing.T) { + t.Parallel() + for _, tt := range []struct { + name string + refuse, restore [][]string + reply string + }{ + { + "demoted to a replica", + [][]string{{"REPLICAOF", "127.0.0.1", "1"}}, // nothing listens: it stays a replica, serving reads + [][]string{{"REPLICAOF", "NO", "ONE"}}, + "READONLY", + }, + { + "full under noeviction", + [][]string{{"CONFIG", "SET", "maxmemory-policy", "noeviction"}, {"CONFIG", "SET", "maxmemory", "1"}}, + [][]string{{"CONFIG", "SET", "maxmemory", "0"}}, + "OOM", + }, + } { + t.Run(tt.name, func(t *testing.T) { + t.Parallel() + ctx := context.Background() + s := startRedis(t) + r := raw(t, s) + prefix := uniquePrefix() + const openFor = 200 * time.Millisecond + tune := func(c *cache.RedisConfig) { c.BreakerThreshold, c.BreakerOpenFor = 1000, openFor } + a := open(t, s, prefix, tune) + ingest := open(t, s, prefix, tune) + reader := open(t, s, prefix, tune) + events := []cache.Namespace{{Tenant: "acme", Table: "events"}} + orders := []cache.Namespace{{Tenant: "acme", Table: "orders"}} + for _, deps := range [][]cache.Namespace{events, orders} { + _, snap, err := a.Lookup(ctx, "acme", "q", deps) + require.NoError(t, err) + require.NoError(t, a.Set(ctx, snap, []byte("pre-write rows"), time.Minute)) + } + lookup := func(c *cache.RedisCache, deps []cache.Namespace) string { + t.Helper() + e, _, err := c.Lookup(ctx, "acme", "q", deps) + require.NoError(t, err) + return string(e.Value) + } + + for _, cmd := range tt.refuse { + command(t, r, cmd...) + } + _, err := a.Invalidate(ctx, events) + require.ErrorContains(t, err, tt.reply) + assert.Equal(t, 1, cache.Pending(a)) + assert.True(t, cache.Bypassed(a), "the first refusal opens the breaker") + wide := slices.Clone(orders) // past one batch: the rest go unsent after the refusal, and drain in two + for i := range 1500 { + wide = append(wide, cache.Namespace{Tenant: "acme", Table: fmt.Sprintf("t%d", i)}) + } + _, err = ingest.Invalidate(ctx, wide) + require.ErrorContains(t, err, tt.reply) + assert.True(t, cache.Bypassed(ingest)) + assert.Equal(t, len(wide), cache.Pending(ingest)) + _, snap, err := reader.Lookup(ctx, "acme", "another query", orders) + require.NoError(t, err) + require.ErrorContains(t, reader.Set(ctx, snap, []byte("rows"), time.Minute), tt.reply) + assert.True(t, cache.Bypassed(reader), "a refused fill opens the breaker too") + + // Long enough for many probes to be refused, and for a drain + // backing off unchecked to be seconds from its next attempt. + time.Sleep(3500 * time.Millisecond) + assert.True(t, cache.Bypassed(reader), "owing nothing, it is held open by probes that write and are refused") + assert.True(t, cache.Bypassed(a)) + assert.Empty(t, lookup(a, events)) + assert.Equal(t, 1, cache.Pending(a)) + assert.Equal(t, len(wide), cache.Pending(ingest)) + + for _, cmd := range tt.restore { + command(t, r, cmd...) + } + restored := time.Now() + for cache.Pending(a) > 0 { + require.Less(t, time.Since(restored), 10*time.Second, "the bump lands once the server takes writes") + require.NotEqual(t, "pre-write rows", lookup(a, events), "served before the owed bump landed") + time.Sleep(time.Millisecond) + } + require.Eventually(t, func() bool { return cache.Pending(ingest) == 0 }, 10*time.Second, 5*time.Millisecond) + assert.Less(t, time.Since(restored), openFor+time.Second, "an ingest-only process probes when due, not at the drain's backoff") + require.Eventually(t, func() bool { return !cache.Bypassed(reader) }, 10*time.Second, 5*time.Millisecond) + + fresh := open(t, s, prefix) + for _, deps := range [][]cache.Namespace{events, orders} { + assert.Empty(t, lookup(fresh, deps), "%v: the landed bumps orphaned the pre-write rows", deps) + } + }) + } +} diff --git a/internal/cache/redis_test.go b/internal/cache/redis_test.go index 9aafb4f12..3fe67a2f5 100644 --- a/internal/cache/redis_test.go +++ b/internal/cache/redis_test.go @@ -15,6 +15,8 @@ import ( "go.opentelemetry.io/otel/attribute" sdkmetric "go.opentelemetry.io/otel/sdk/metric" "go.opentelemetry.io/otel/sdk/metric/metricdata" + + "github.com/Wave-RF/WaveHouse/internal/tenant" ) func TestRedisConfig_Validation(t *testing.T) { @@ -149,6 +151,49 @@ func TestRedis_Record(t *testing.T) { assert.True(t, r.breaker.isOpen()) } +func TestRefusesWork(t *testing.T) { + t.Parallel() + for _, msg := range []string{ + "READONLY You can't write against a read only replica.", + "OOM command not allowed when used memory > 'maxmemory'.", + "MASTERDOWN Link with MASTER is down and replica-serve-stale-data is set to 'no'.", + "NOREPLICAS Not enough good replicas to write.", + "MISCONF Errors writing to the AOF file: No space left on device", + "LOADING Redis is loading the dataset in memory", + "BUSY Redis is busy running a script. You can only call SCRIPT KILL or SHUTDOWN NOSAVE.", + "CLUSTERDOWN The cluster is down", + } { + assert.True(t, refusesWork(msg), msg) + } + for _, msg := range []string{ + "WRONGTYPE Operation against a key holding the wrong kind of value", + "NOPERM this user has no permissions to access one of the keys used as arguments", + "TRYAGAIN Multiple keys request during rehashing of slot", + "BUSYKEY Target key name already exists.", + "ERR unknown command", + "", + } { + assert.False(t, refusesWork(msg), msg) + } +} + +// The first bump owed wakes the drain at once; later ones ride the retry +// already under way rather than resetting its backoff. +func TestRedis_FirstDeferralWakesTheDrain(t *testing.T) { + t.Parallel() + r := &RedisCache{breaker: newBreaker(1, time.Hour, time.Now), pending: newPendingBumps("wh", 10), wake: make(chan struct{}, 1)} + var err error + r.metrics, err = newMetrics("redis", r.bypassed, r.pending.len) + require.NoError(t, err) + t.Cleanup(r.metrics.close) + + r.deferBumps(map[string]tenant.ID{"wh:{acme}:B:events": "acme"}) + assert.Len(t, r.wake, 1) + <-r.wake + r.deferBumps(map[string]tenant.ID{"wh:{acme}:B:orders": "acme"}) + assert.Empty(t, r.wake) +} + func cancelledCtx() context.Context { ctx, cancel := context.WithCancel(context.Background()) cancel() From c6df13ad35f5311c6cb8f0fd3010c339ac5f111a Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:21:09 -0400 Subject: [PATCH 42/79] fix(cache): bound a cluster client's topology read like a dial rueidis reads the cluster topology after the handshake under ConnWriteTimeout, 10 s when unset. A node that answered the handshake and then went quiet held NewClient for 10 s (measured), and with it NewRedis's first dial and a Close waiting on the dial loop; a refused or silent node already failed within DialTimeout. ConnWriteTimeout is now the larger of DialTimeout and the op Timeout: bounded like a dial, and never cutting a connection an operation may still wait on. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- docs/src/content/docs/architecture.md | 2 +- internal/cache/export_test.go | 8 ++++++ internal/cache/redis.go | 5 ++++ internal/cache/redis_integration_test.go | 34 ++++++++++++++++++++++++ internal/cache/redis_test.go | 4 +++ 5 files changed, 52 insertions(+), 1 deletion(-) diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index f124f8d7c..d9c08f81e 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -113,7 +113,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on — one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)) — each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope by their raw names, which the cache escapes where it builds a key. `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. - **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: a namespace key is `..
.
.`, its fields joined with `keyenc.Join` so a dot in a table or scope name is escaped rather than read as a separator, and an entry's key, `|.||…` with the caller's query key escaped whole (its `:` become `%3A`), folds in the tenant's version and each dependency's namespace key and namespace version, so bumping a table (a scopeless write) or one scope — scope is reserved and empty today, so every write is the whole-table bump — orphans every dependent entry without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results, and the flat directory's single tenant simply carries the `0` prefix. `BumpTenant` (behind `InvalidateTenant`) advances the tenant version that leads every namespace key of one tenant, orphaning its every namespace and every cached query in one step (the tenant version is folded into every query key, so a pipe result with no dependencies is orphaned too) — a table no bump ever keyed included, which an enumeration of the index would miss — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. -- **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET` and `MGET`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot; the table and scope are escaped with `keyenc` after the fixed prefix, so a `:` in a name is never read as the separator. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, the hash over their escaped forms, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token (a string of another length, a list, a hash) is replaced, so a token lost to eviction, expiry or a restart is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. While this process owes a bump on any of a lookup's tokens, that lookup is a bypass that files nothing: the bump would orphan whatever it found (other processes, which cannot know, serve those entries until it lands). `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing. Nothing selects this backend yet — the boot config and wiring arrive with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s E4. +- **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET` and `MGET`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot; the table and scope are escaped with `keyenc` after the fixed prefix, so a `:` in a name is never read as the separator. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, the hash over their escaped forms, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token (a string of another length, a list, a hash) is replaced, so a token lost to eviction, expiry or a restart is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. While this process owes a bump on any of a lookup's tokens, that lookup is a bypass that files nothing: the bump would orphan whatever it found (other processes, which cannot know, serve those entries until it lands). `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing, each attempt bounded by `DialTimeout` (1 s), and a cluster client's topology read after the handshake by the larger of it and `Timeout`, which is also how long a connection waits on a silent server before it is redialed. Nothing selects this backend yet — the boot config and wiring arrive with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s E4. - **redis_codec.go** — the key schema (`tokenKeys`, `bumpKeys`, `valueKey`, which take a `Namespace`'s raw names and escape them) and the value frame: format, flags, expiry, the token list, the payload, zstd-compressed from `CompressMinBytes` when that is smaller. A stored value is capped at `MaxValueBytes` and a decoded one at eight times that, which refuses a zip bomb planted in a shared server; an unknown format, as a newer process writes during a rolling upgrade, is a miss. - **breaker.go** — the circuit breaker: `BreakerThreshold` (5) consecutive failures open it, and while open every operation skips the server at once; after `BreakerOpenFor` (5 s) one background write (`SET :probe`) decides whether it closes, and only that probe's success closes it. A transport failure or a timeout counts against the server; an error reply saying it takes no writes right now — `READONLY` (a demoted primary), `MASTERDOWN`, `OOM` (full under `maxmemory-policy noeviction`, Redis's default, so give the cache an evicting policy), `NOREPLICAS`, `MISCONF`, `LOADING`, `BUSY`, `CLUSTERDOWN` — opens it at once, whatever the threshold, since no bump can land. Any other reply, an error reply about one key (`WRONGTYPE`, `NOPERM`, `TRYAGAIN`) or one the backend cannot use included, counts as a success, and a caller that gave up first counts as nothing. - **pending.go** — invalidations the server did not take, kept per token key (repeats coalesce) and retried by a background loop until they land: the first one owed at once, then with backoff from 100 ms to 10 s, and while the breaker is open at each probe, which the loop starts when due, so a process that makes no lookups (ingest only) recovers as soon as one that does. `Invalidate` sends its bumps 1,000 to a round trip, as the retry does. Landing late is still correct: a fresh token orphans the pre-write entries and any fill made meanwhile. Past `PendingMax` (100,000) keys the set collapses to one tenant bump per affected tenant. `Close` makes one last attempt at the bumps still owed, past the breaker (an open one is why they are owed); what that attempt cannot deliver is lost, and the entries they would orphan are served until their TTL — the same failure as a worker stopping between an insert and its invalidation. diff --git a/internal/cache/export_test.go b/internal/cache/export_test.go index aaf712f7f..9501f6be5 100644 --- a/internal/cache/export_test.go +++ b/internal/cache/export_test.go @@ -1,7 +1,15 @@ package cache +import "github.com/redis/rueidis" + // Hooks for the integration tests in package cache_test. +// ClientOption is the rueidis option a RedisCache built from c dials with. +func ClientOption(c RedisConfig) (rueidis.ClientOption, error) { + c, err := c.withDefaults() + return c.clientOption(), err +} + // Pending reports how many token bumps r still owes the server. func Pending(r *RedisCache) int { return r.pending.len() } diff --git a/internal/cache/redis.go b/internal/cache/redis.go index ec24893d1..b8bd5e286 100644 --- a/internal/cache/redis.go +++ b/internal/cache/redis.go @@ -158,6 +158,11 @@ func (c RedisConfig) clientOption() rueidis.ClientOption { DisableCache: true, // no client-side caching until the near-cache (E5) ForceSingleClient: c.Mode == RedisStandalone, } + // How long a connection waits on a silent server, 10 s unset. A cluster + // client reads the topology under it after the handshake, which boot and + // Close would wait out; never under Timeout, so no connection is cut + // while an operation may still wait on it. + opt.ConnWriteTimeout = max(c.DialTimeout, c.Timeout) if c.Mode == RedisSentinel { opt.Sentinel = rueidis.SentinelOption{MasterSet: c.SentinelMaster, TLSConfig: c.TLS, Dialer: opt.Dialer} } diff --git a/internal/cache/redis_integration_test.go b/internal/cache/redis_integration_test.go index 882ed6a52..a6f0aae23 100644 --- a/internal/cache/redis_integration_test.go +++ b/internal/cache/redis_integration_test.go @@ -5,6 +5,7 @@ package cache_test import ( "bytes" "context" + "crypto/tls" "fmt" "net" "slices" @@ -497,6 +498,39 @@ func TestRedis_CloseDeliversPastAnOpenBreaker(t *testing.T) { assert.Nil(t, e.Value, "the bump Close delivered orphans the fill") } +// A cluster client reads the topology after the handshake; a node that +// answers the handshake and then goes quiet held that read, and so boot and +// Close, for rueidis's 10 s default. It is bounded like a dial now. Any +// server serves: what matters is that the read goes unanswered. +func TestRedis_ClusterTopologyReadIsBounded(t *testing.T) { + t.Parallel() + s := startRedis(t) + opt, err := cache.ClientOption(cache.RedisConfig{Addrs: []string{s.addr}, Mode: cache.RedisCluster, DialTimeout: 300 * time.Millisecond}) + require.NoError(t, err) + opt.DialCtxFn = func(ctx context.Context, addr string, d *net.Dialer, _ *tls.Config) (net.Conn, error) { + c, err := d.DialContext(ctx, "tcp", addr) + return unanswered{c}, err + } + start := time.Now() + c, err := rueidis.NewClient(opt) + if err == nil { + c.Close() + } + require.Error(t, err, "nothing answered the topology read") + assert.Less(t, time.Since(start), 3*time.Second) +} + +// unanswered drops CLUSTER commands unsent, as a node gone quiet would +// leave them unanswered. +type unanswered struct{ net.Conn } + +func (c unanswered) Write(b []byte) (int, error) { + if bytes.Contains(b, []byte("CLUSTER")) { + return len(b), nil + } + return c.Conn.Write(b) +} + // command runs cmd on s through c, for the test to reconfigure the server. func command(t *testing.T, c rueidis.Client, cmd ...string) { t.Helper() diff --git a/internal/cache/redis_test.go b/internal/cache/redis_test.go index 3fe67a2f5..17da785a6 100644 --- a/internal/cache/redis_test.go +++ b/internal/cache/redis_test.go @@ -60,6 +60,10 @@ func TestRedisConfig_Defaults(t *testing.T) { assert.Equal(t, DefaultRedisVersionTTL, c.VersionTTL) assert.Zero(t, c.CompressMinBytes, "0 means never compress, not the default") assert.True(t, c.clientOption().ForceSingleClient) + assert.Equal(t, DefaultRedisDialTimeout, c.clientOption().ConnWriteTimeout) + slow := c + slow.Timeout = 5 * time.Second + assert.Equal(t, slow.Timeout, slow.clientOption().ConnWriteTimeout, "never under the op timeout") c.Mode, c.SentinelMaster = RedisSentinel, "mymaster" assert.Equal(t, "mymaster", c.clientOption().Sentinel.MasterSet) From 3c19a857298460889966245076d8eae4eb5197fb Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:24:43 -0400 Subject: [PATCH 43/79] test(cache): count Redis values so the zero-snapshot case runs The conformance suite's zero-snapshot case needs Options.Entries and skipped for the Redis backend. It now counts the values under the case's own key prefix, scanning every node, plus the empty key a fill under the zero snapshot would land on, so the case runs on Redis, Valkey, Dragonfly and the Redis Cluster node. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- internal/cache/export_test.go | 3 +++ internal/cache/redis_integration_test.go | 28 ++++++++++++++++++++++++ 2 files changed, 31 insertions(+) diff --git a/internal/cache/export_test.go b/internal/cache/export_test.go index 8b4a503c8..0ef1f0b61 100644 --- a/internal/cache/export_test.go +++ b/internal/cache/export_test.go @@ -16,6 +16,9 @@ func Pending(r *RedisCache) int { return r.pending.len() } // Bypassed reports whether r is skipping the server. func Bypassed(r *RedisCache) bool { return r.bypassed() } +// KeyPrefix is the prefix every key r writes leads with. +func KeyPrefix(r *RedisCache) string { return r.cfg.KeyPrefix } + // ZeroSnapshot reports whether s files nothing. func ZeroSnapshot(s Snapshot) bool { return s.key == "" && s.tokens == nil } diff --git a/internal/cache/redis_integration_test.go b/internal/cache/redis_integration_test.go index a6f0aae23..68c3e71a2 100644 --- a/internal/cache/redis_integration_test.go +++ b/internal/cache/redis_integration_test.go @@ -222,6 +222,7 @@ func TestRedis_Conformance(t *testing.T) { p := uniquePrefix() return open(t, s, p), open(t, s, p) }, + Entries: valueCount(raw(t, s)), }) t.Run("cross-tenant invalidation spans slots", func(t *testing.T) { t.Parallel() @@ -235,6 +236,33 @@ func TestRedis_Conformance(t *testing.T) { } } +// valueCount counts the values under a cache's own key prefix, scanning +// every node r knows (the test cluster's one node has no replica), and the +// empty key, where a fill under the zero snapshot's empty key would land; +// -1 is a failed read. +func valueCount(r rueidis.Client) func(cache.Cache) int { + return func(c cache.Cache) int { + match := cache.KeyPrefix(c.(*cache.RedisCache)) + ":q:*" + n, err := r.Do(context.Background(), r.B().Exists().Key("").Build()).AsInt64() + if err != nil { + return -1 + } + for _, node := range r.Nodes() { + for cursor := uint64(0); ; { + e, err := node.Do(context.Background(), node.B().Scan().Cursor(cursor).Match(match).Count(1000).Build()).AsScanEntry() + if err != nil { + return -1 + } + if n += int64(len(e.Elements)); e.Cursor == 0 { + break + } + cursor = e.Cursor + } + } + return int(n) + } +} + // The ingest worker's shared-tables fan-out bumps one table under several // tenants in one call: tokens in as many slots, one pipeline. func testCrossTenantInvalidate(t *testing.T, c *cache.RedisCache) { From 05a536dd1bf9c9d1aa81ed8085125d26298758d0 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:28:11 -0400 Subject: [PATCH 44/79] fix(api): step over heredocs when classifying writes ClickHouse reads $$...$$ and $tag$...$tag$ as a string, but isMutation scanned into it: a paren or quote inside hid the INSERT after a WITH list, so the write went through Query and could run again on a retry, and a verb inside made a read look like a write. The scanner now skips a heredoc as ClickHouse's lexer does (tag of letters, digits and _, matched exactly; unclosed is not a heredoc), and reads $ as part of a bareword, as ClickHouse does (a$b, set$). Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- CHANGELOG.md | 2 +- internal/api/clickhouse_exec.go | 64 +++++++++++++++++++--------- internal/api/clickhouse_exec_test.go | 9 ++++ 3 files changed, 55 insertions(+), 20 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index f60b7e80f..819700a69 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -86,7 +86,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed - **A pipe that writes runs on every call instead of being answered from the cache** (`internal/api/pipes.go` (+ tests), `docs/src/content/docs/{pipes.mdx,api.md,architecture.md}`, `AGENTS.md`): fixes [#386](https://github.com/Wave-RF/WaveHouse/issues/386). `/v1/pipes/{name}` sent a write's SQL to ClickHouse through `Exec`, but still cached the `[]` it returned and coalesced identical calls in flight, so a repeat within the TTL answered `200` without executing and concurrent identical calls became one write — silently dropped writes, and with a shared cache ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)) on every instance. A pipe whose bound SQL `isMutation` classifies as a write — the same classifier that picks `Exec` — now skips the cache lookup, the fill and singleflight, and answers `X-Cache: BYPASS` with `Cache-Control: no-store`, so an HTTP cache in front of a `GET` cannot drop the write either. Classification stays automatic rather than a declared pipe property, so an operator cannot forget to mark one, and costs no ClickHouse round trip. Read pipes are unchanged. Not in this fix: a write pipe still does not invalidate cached reads of the table it writes ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). -- **The write classifier skips whitespace, comments and quoted text the way ClickHouse's lexer does** (`internal/api/clickhouse_exec.go` (+ tests)): `isMutation` picks `Exec` for a write, and since [#386](https://github.com/Wave-RF/WaveHouse/issues/386) keeps a write pipe out of the cache. It missed a write behind a backslash-escaped quote (`'it\'s'`, and the same inside `"…"` and `` `…` ``), behind a nested block comment (`/* a /* b */ SELECT */ INSERT …`), or behind leading whitespace other than space, tab, CR and LF: `\v`, `\f`, a no-break space, a byte-order mark, and the other Unicode spaces ClickHouse skips. A missed write went through `Query`, which ran it and then failed the call with a `5xx` the TypeScript SDK retries, so one call could write three times. The same gaps could make a read look like a write, which runs through `Exec` and answers `[]`. +- **The write classifier skips whitespace, comments and quoted text the way ClickHouse's lexer does** (`internal/api/clickhouse_exec.go` (+ tests)): `isMutation` picks `Exec` for a write, and since [#386](https://github.com/Wave-RF/WaveHouse/issues/386) keeps a write pipe out of the cache. It missed a write behind a backslash-escaped quote (`'it\'s'`, and the same inside `"…"` and `` `…` ``), a heredoc (`$$ ( $$`, `$tag$ … $tag$`), a nested block comment (`/* a /* b */ SELECT */ INSERT …`), or behind leading whitespace other than space, tab, CR and LF: `\v`, `\f`, a no-break space, a byte-order mark, and the other Unicode spaces ClickHouse skips. A missed write went through `Query`, which ran it and then failed the call with a `5xx` the TypeScript SDK retries, so one call could write three times. The same gaps could make a read look like a write, which runs through `Exec` and answers `[]`. - **The pipes page no longer says a parameter can never break out of its literal** (`docs/src/content/docs/pipes.mdx`, `internal/pipes/pipes.go`): that holds only for a placeholder written bare. A string value brings its own quotes, so inside a quoted placeholder they close the template's: the body `{"id": " OR 1=1 OR id = "}` turns `WHERE id = '{{id}}'` into `WHERE id = '' OR 1=1 OR id = ''`, which matches every row. The page now says to write each placeholder bare, never inside quotes. - **A failed ClickHouse query answers by what went wrong, not a flat `500`/`502`** (`internal/api/ch_errors.go` (new, + tests), `internal/api/{errors,query,structured_query,pipes,schema,ch_settings}.go`, `internal/chconn/errclass.go` (`HTTPStatus` exported), `clients/ts/src/errors.ts` (+ tests), `tests/integration/query_errors_test.go` (new), `tests/integration/query_limits_test.go`, `internal/app/app_test.go`, `tests/e2e/sdk/{admin,query}.test.ts`, `AGENTS.md`, `docs/src/content/docs/{api,architecture}.md`, `docs/src/content/docs/{access-control,configuration}.mdx`, `docs/src/content/docs/sdk/{reference.md,index.mdx}`): fixes [#403](https://github.com/Wave-RF/WaveHouse/issues/403) and [#271](https://github.com/Wave-RF/WaveHouse/issues/271), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). ClickHouse answers a syntax error, a missing grant and an overloaded server alike with HTTP `500`, so `/v1/ops/query` turned a bad statement into a `502` and `/v1/query` and pipes into a `500` the SDK retried. All three now class the failure with `chconn.Classify` through one helper, `writeCHError`: a statement ClickHouse refused is `400 clickhouse.rejected`; a query over a rows/bytes limit, the role's own memory cap, or its time cap where that is no longer than `query_timeout` is `400 clickhouse.limit_exceeded`; `ACCESS_DENIED` is `403 clickhouse.access_denied`; credentials, user or database refused, or a redirect or `4xx` with no exception code from whatever fronts ClickHouse, is `502 clickhouse.misconfigured`; ClickHouse down, unreachable or overloaded is `503 clickhouse.unavailable` with `Retry-After: 5`; a failure with no verdict stays `500` (`502` on the proxy) as `clickhouse.unknown`. The error envelope gains `code` and `retryable` next to `error` on these responses — additive. A role with `max_execution_time` now queries with no context deadline and a cancel two seconds past the cap instead: clickhouse-go overwrote the cap's `max_execution_time` with deadline+5s for any deadline over 1s, so an overrun came back as a bare deadline, indistinguishable from waiting for a pooled connection; ClickHouse now enforces the cap itself and reports `TIMEOUT_EXCEEDED`. `POST /v1/ops/schema/refresh` against an unreachable ClickHouse is a `503` with `Retry-After` instead of a `500`. **SDK:** `WaveHouseError.code` and `retryable` now take the server's `code`/`retryable` when the body has them (`HTTP_` and "5xx retries" otherwise), so a rejected query is `clickhouse.rejected` rather than `HTTP_500`, and is not retried. - **An unavailable ClickHouse is retried with backoff instead of dead-lettering every row** (`internal/chconn/errclass.go` (new, + tests), `internal/ingest/{worker,backoff}.go` (`backoff.go` new, + tests), `internal/mq/{mq,embedded}.go`, `internal/testutil/mocks.go`, `tests/integration/ingest_outage_test.go` (new), `AGENTS.md`, `README.md`, `docs/src/content/docs/{ingest-pipeline,architecture,api,deployment,why-wavehouse}.md`, `docs/src/content/docs/{settings-directory,index,access-control}.mdx`): workstream A of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). A failed batch insert used to go through row-by-row isolation whatever the failure, so a ClickHouse that was down, overloaded or read-only failed every row twice and parked the whole batch on the DLQ. `chconn.Classify` now classes the failure first — `Rejected` (any ClickHouse exception code outside the availability and credential lists: the server read the row and refused it), `Unavailable` (connection refused/reset, timeouts, `TOO_MANY_SIMULTANEOUS_QUERIES`, `SERVER_OVERLOADED`, `MEMORY_LIMIT_EXCEEDED`, `TOO_MANY_PARTS`, `READONLY`, `TABLE_IS_READ_ONLY`, `KEEPER_EXCEPTION`, …), `Denied` (`AUTHENTICATION_FAILED`, `ACCESS_DENIED`, …) or `Unknown` (no code, no recognizable transport failure). Only `Rejected` is isolated and dead-lettered as before, and a multi-row batch refused with `TOO_MANY_PARTS` or `MEMORY_LIMIT_EXCEEDED` is split row by row first (`chconn.Splittable`), because a batch spanning too many partitions or too much memory can fail where each of its rows inserts; every other class hands the batch back to the queue with a delayed nak (`mq.Message.NakWithDelay`, new) under a jittered 1 s → 30 s backoff shared by every table on the same ClickHouse pool (a failure of one table — read-only, too many parts or mutations, a grant missing on it, `chconn.TableScoped` — backs off that table alone), which turns rows away without a request while it runs and probes once per window, and ClickHouse going away mid-isolation stops isolation and retries the rows it had not settled. Counted by the new `wavehouse_ingest_retries_total{table, reason}`; logged at `WARN` when an outage starts and at most every 30 s during it. A long outage now shows as a growing ingest stream and, at `mq.max_bytes_gb`, ingest `503`s — not as a full DLQ; a lasting failure of one table holds back its tenant's other tables once its waiting rows reach `maxAckPending`. Retried rows come back out of arrival order, which matters only to a `ReplacingMergeTree` without a version column or a `CollapsingMergeTree`. diff --git a/internal/api/clickhouse_exec.go b/internal/api/clickhouse_exec.go index 52a6228d6..c86d88a40 100644 --- a/internal/api/clickhouse_exec.go +++ b/internal/api/clickhouse_exec.go @@ -164,10 +164,10 @@ var nonMutationVerbs = map[string]struct{}{ // containsMutationVerbAtTopLevel scans s for the statement-introducing // keyword at paren-depth 0, stepping over string literals and quoted -// identifiers (skipQuoted), parenthesized CTE subqueries, and comments -// (skipComment). The CTE list contains ordinary identifiers -// (CTE names, table/database names) that must not be matched as mutation -// verbs — `system` would otherwise pattern-match `SYSTEM` and route a +// identifiers (skipQuoted), heredocs (skipHeredoc), parenthesized CTE +// subqueries, and comments (skipComment). The CTE list contains ordinary +// identifiers (CTE names, table/database names) that must not be matched as +// mutation verbs — `system` would otherwise pattern-match `SYSTEM` and route a // `WITH … SELECT * FROM system.tables` read through `Exec` (silent empty- // array result instead of the actual rows). Two-part fix: // @@ -203,15 +203,17 @@ func containsMutationVerbAtTopLevel(s string) bool { i++ case c == '\'' || c == '"' || c == '`': i = skipQuoted(s, i) + case c == '$': + // A heredoc, else a bareword led by `$` (never a keyword) or a + // lone `$`. + if j := skipHeredoc(s, i); j > i { + i = j + } else { + i = skipWord(s, i+1) + } case (c >= 'A' && c <= 'Z') || (c >= 'a' && c <= 'z'): start := i - for i < len(s) { - c2 := s[i] - if (c2 < 'A' || c2 > 'Z') && (c2 < 'a' || c2 > 'z') && (c2 < '0' || c2 > '9') && c2 != '_' { - break - } - i++ - } + i = skipWord(s, i) if depth == 0 { kw := strings.ToUpper(s[start:i]) // Check non-mutation statement keywords (SELECT, SHOW, @@ -258,15 +260,39 @@ func isCTENameLookahead(s string, pos int) bool { if s[i] == '(' { return true } - end := i - for end < len(s) { - c := s[end] - if (c < 'A' || c > 'Z') && (c < 'a' || c > 'z') && (c < '0' || c > '9') && c != '_' { - break - } - end++ + return strings.EqualFold(s[i:skipWord(s, i)], "AS") +} + +// skipWord returns the index just past the bareword at s[i]: ClickHouse's +// barewords run over ASCII letters, digits, `_` and `$`. +func skipWord(s string, i int) int { + for i < len(s) && (isWordByte(s[i]) || s[i] == '$') { + i++ + } + return i +} + +func isWordByte(c byte) bool { + return (c >= 'A' && c <= 'Z') || (c >= 'a' && c <= 'z') || (c >= '0' && c <= '9') || c == '_' +} + +// skipHeredoc returns the index just past the heredoc opening at s[i] — +// `$tag$ … $tag$`, the tag a possibly empty run of letters, digits and `_`, +// matched exactly — or i if none does, as an unclosed one is not a heredoc to +// ClickHouse either. +func skipHeredoc(s string, i int) int { + j := i + 1 + for j < len(s) && isWordByte(s[j]) { + j++ + } + if j >= len(s) || s[j] != '$' { + return i } - return strings.EqualFold(s[i:end], "AS") + tag := s[i : j+1] + if k := strings.Index(s[j+1:], tag); k >= 0 { + return j + 1 + k + len(tag) + } + return i } // stripLeadingSQLComments trims whitespace and comments from the front of diff --git a/internal/api/clickhouse_exec_test.go b/internal/api/clickhouse_exec_test.go index 76441b6cf..86ad6c6e7 100644 --- a/internal/api/clickhouse_exec_test.go +++ b/internal/api/clickhouse_exec_test.go @@ -119,6 +119,15 @@ func TestIsMutation(t *testing.T) { {"with backslash-escaped backtick then insert", "WITH m AS (SELECT 'x' AS `a\\`(b`) INSERT INTO t SELECT * FROM m", true}, {"with backslash-escaped backtick then select", "WITH m AS (SELECT 1 AS `a\\`b`) SELECT 2 AS `x) INSERT` FROM m", false}, + // A heredoc ($$…$$, $tag$…$tag$) is a literal: its parens, quotes + // and words are not the statement's. + {"with heredoc holding a paren then insert", "WITH $$ ( $$ AS s INSERT INTO t SELECT s", true}, + {"with tagged heredoc holding a quote then insert", "WITH $x$ it's $x$ AS s INSERT INTO t SELECT s", true}, + {"with tagged heredoc holding a paren and another tag then insert", "WITH $x$ ( $y$ $x$ AS s INSERT INTO t SELECT s", true}, + {"with heredoc holding a verb then select", "WITH $$INSERT$$ AS s SELECT s", false}, + {"with tagged heredoc holding a paren and a verb then select", "WITH $x$ ) INSERT $x$ AS s SELECT s", false}, + {"with CTE alias set$ (read)", "WITH set$ AS (SELECT 1 AS v) SELECT * FROM set$", false}, + // ClickHouse's lexer skips \v, \f and Unicode spaces as whitespace // (TestIsMutation_ClickHouseWhitespace covers the whole set). {"leading form feed then insert", "\fINSERT INTO t VALUES (1)", true}, From 0880685e89c256ceac6a8bba2390d9e52f8912ca Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:29:44 -0400 Subject: [PATCH 45/79] fix(api): a failed write pipe is never retryable A write pipe's failure answered a bare 500, while a read's is classed by the ClickHouse error. It is now classed the same way (status and code), but always as retryable:false with no Retry-After, 503 included: once the statement is sent it may have run, and the SDK retries both a retryable 5xx and any 503 carrying Retry-After, which would run the write again. A write refused before it is sent (the tenant on no pool) keeps its 503 with Retry-After: 30. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- AGENTS.md | 2 +- CHANGELOG.md | 2 +- docs/src/content/docs/api.md | 4 +++- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/pipes.mdx | 2 ++ docs/src/content/docs/sdk/pipes.md | 2 +- internal/api/ch_errors.go | 18 ++++++++++++++++-- internal/api/ch_errors_test.go | 26 ++++++++++++++++++++++++++ internal/api/pipes.go | 2 +- internal/api/pipes_test.go | 9 +++++---- internal/api/tenant_clickhouse_test.go | 9 +++++++++ 11 files changed, 66 insertions(+), 12 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 942684c4b..ea7a87cae 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -28,7 +28,7 @@ One binary: Twenty internal packages under `internal/` (plus `internal/testutil/` for shared test helpers): -- **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers; `ch_errors.go` (`writeCHError`) is the one mapping from a failed ClickHouse query to status, `code` and `retryable` +- **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers; `ch_errors.go` (`writeCHError`) is the one mapping from a failed ClickHouse query to status, `code` and `retryable` (`writeCHWriteError` for a write pipe: never retryable, no `Retry-After`, since the write may have run) - **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` handing the sweeper each tenant's own gap window (a rejected tenant's as its folder last had it, unbounded for one rejected since boot) and the `mq.max_bytes_gb` reconcile each served tenant's byte budget, and `defaultPolicy` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, and the same hook's `Hub.Prune` ends the open streams of a tenant no longer served), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `New` wires only what the process's `roles` need (discovery, dedupe, auth verifiers, the hub bridge and keepalive per API process; the ingest worker per ingest process; the sweeper under its lease through `elected`); a process without `api` serves `api.NewOpsRouter` — probes, `/version`, metrics, and the settings reload behind the operator key alone. `cmd/wavehouse` and `tests/integration` both boot through it - **`auth/`** — JWT auth middleware: HMAC **or** JWKS verification with `alg` pinned to the active verifier, role extraction from a configurable claim path; always runs, never rejects (bad token → empty role + stashed reason). One verifier per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 9): `Authenticator` keys them by `tenant.ID` — the request store's `settings.Store.Tenant()`, through an injected `TenantSource`; `tenant.Default` on the tenant-exempt routes — built from each tenant's `auth` block by `Reconfigure`, dropped by `Prune` once the tenant stops being served (rejected or removed), released by `Close`; the secrets (`Config`) are boot-level and shared. A JWKS key set is fetched off the boot and reload paths: until one has been stored the verifier is pending and a token-bearing request gets `503` + `Retry-After` from `api.refuseUnverifiable` (`auth.ErrVerifierPending`), never a `default_role` evaluation; refresh is library-managed (Eric, 2026-09-22), response capped at 1 MiB; the operator key's admin role is the request tenant's - **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for the caller's query key and its singleflight (escaped whole as the lead field of the stored key, `|.||…`), `..
.
.` for a namespace, each field escaped by `keyenc` where the key is built (a `Namespace` carries the raw table and scope, so no caller escapes) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version folded into every namespace and query key of one tenant, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read, taken before the handler chooses any input a bump invalidates — the tenant's connection included — and `Set` files the fill under it, so a write landing mid-query, or a reload moving the tenant to another address or database after the request took its connection, orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) diff --git a/CHANGELOG.md b/CHANGELOG.md index 819700a69..9ed3da2cb 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -85,7 +85,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed -- **A pipe that writes runs on every call instead of being answered from the cache** (`internal/api/pipes.go` (+ tests), `docs/src/content/docs/{pipes.mdx,api.md,architecture.md}`, `AGENTS.md`): fixes [#386](https://github.com/Wave-RF/WaveHouse/issues/386). `/v1/pipes/{name}` sent a write's SQL to ClickHouse through `Exec`, but still cached the `[]` it returned and coalesced identical calls in flight, so a repeat within the TTL answered `200` without executing and concurrent identical calls became one write — silently dropped writes, and with a shared cache ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)) on every instance. A pipe whose bound SQL `isMutation` classifies as a write — the same classifier that picks `Exec` — now skips the cache lookup, the fill and singleflight, and answers `X-Cache: BYPASS` with `Cache-Control: no-store`, so an HTTP cache in front of a `GET` cannot drop the write either. Classification stays automatic rather than a declared pipe property, so an operator cannot forget to mark one, and costs no ClickHouse round trip. Read pipes are unchanged. Not in this fix: a write pipe still does not invalidate cached reads of the table it writes ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). +- **A pipe that writes runs on every call instead of being answered from the cache** (`internal/api/{pipes,ch_errors}.go` (+ tests), `docs/src/content/docs/{pipes.mdx,api.md,architecture.md,sdk/pipes.md}`, `AGENTS.md`): fixes [#386](https://github.com/Wave-RF/WaveHouse/issues/386). `/v1/pipes/{name}` sent a write's SQL to ClickHouse through `Exec`, but still cached the `[]` it returned and coalesced identical calls in flight, so a repeat within the TTL answered `200` without executing and concurrent identical calls became one write — silently dropped writes, and with a shared cache ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)) on every instance. A pipe whose bound SQL `isMutation` classifies as a write — the same classifier that picks `Exec` — now skips the cache lookup, the fill and singleflight, and answers `X-Cache: BYPASS` with `Cache-Control: no-store`, so an HTTP cache in front of a `GET` cannot drop the write either. Classification stays automatic rather than a declared pipe property, so an operator cannot forget to mark one, and costs no ClickHouse round trip. A failed write answers with the status and `code` a failed read gets (see the ClickHouse-errors entry below), but always `retryable: false` and with no `Retry-After`, `503 clickhouse.unavailable` included: the statement may have run, so the SDK does not retry it. A write refused before it is sent, the tenant on no pool, keeps its `503` with `Retry-After: 30`. Read pipes are unchanged. Not in this fix: a write pipe still does not invalidate cached reads of the table it writes ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). - **The write classifier skips whitespace, comments and quoted text the way ClickHouse's lexer does** (`internal/api/clickhouse_exec.go` (+ tests)): `isMutation` picks `Exec` for a write, and since [#386](https://github.com/Wave-RF/WaveHouse/issues/386) keeps a write pipe out of the cache. It missed a write behind a backslash-escaped quote (`'it\'s'`, and the same inside `"…"` and `` `…` ``), a heredoc (`$$ ( $$`, `$tag$ … $tag$`), a nested block comment (`/* a /* b */ SELECT */ INSERT …`), or behind leading whitespace other than space, tab, CR and LF: `\v`, `\f`, a no-break space, a byte-order mark, and the other Unicode spaces ClickHouse skips. A missed write went through `Query`, which ran it and then failed the call with a `5xx` the TypeScript SDK retries, so one call could write three times. The same gaps could make a read look like a write, which runs through `Exec` and answers `[]`. - **The pipes page no longer says a parameter can never break out of its literal** (`docs/src/content/docs/pipes.mdx`, `internal/pipes/pipes.go`): that holds only for a placeholder written bare. A string value brings its own quotes, so inside a quoted placeholder they close the template's: the body `{"id": " OR 1=1 OR id = "}` turns `WHERE id = '{{id}}'` into `WHERE id = '' OR 1=1 OR id = ''`, which matches every row. The page now says to write each placeholder bare, never inside quotes. - **A failed ClickHouse query answers by what went wrong, not a flat `500`/`502`** (`internal/api/ch_errors.go` (new, + tests), `internal/api/{errors,query,structured_query,pipes,schema,ch_settings}.go`, `internal/chconn/errclass.go` (`HTTPStatus` exported), `clients/ts/src/errors.ts` (+ tests), `tests/integration/query_errors_test.go` (new), `tests/integration/query_limits_test.go`, `internal/app/app_test.go`, `tests/e2e/sdk/{admin,query}.test.ts`, `AGENTS.md`, `docs/src/content/docs/{api,architecture}.md`, `docs/src/content/docs/{access-control,configuration}.mdx`, `docs/src/content/docs/sdk/{reference.md,index.mdx}`): fixes [#403](https://github.com/Wave-RF/WaveHouse/issues/403) and [#271](https://github.com/Wave-RF/WaveHouse/issues/271), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). ClickHouse answers a syntax error, a missing grant and an overloaded server alike with HTTP `500`, so `/v1/ops/query` turned a bad statement into a `502` and `/v1/query` and pipes into a `500` the SDK retried. All three now class the failure with `chconn.Classify` through one helper, `writeCHError`: a statement ClickHouse refused is `400 clickhouse.rejected`; a query over a rows/bytes limit, the role's own memory cap, or its time cap where that is no longer than `query_timeout` is `400 clickhouse.limit_exceeded`; `ACCESS_DENIED` is `403 clickhouse.access_denied`; credentials, user or database refused, or a redirect or `4xx` with no exception code from whatever fronts ClickHouse, is `502 clickhouse.misconfigured`; ClickHouse down, unreachable or overloaded is `503 clickhouse.unavailable` with `Retry-After: 5`; a failure with no verdict stays `500` (`502` on the proxy) as `clickhouse.unknown`. The error envelope gains `code` and `retryable` next to `error` on these responses — additive. A role with `max_execution_time` now queries with no context deadline and a cancel two seconds past the cap instead: clickhouse-go overwrote the cap's `max_execution_time` with deadline+5s for any deadline over 1s, so an overrun came back as a bare deadline, indistinguishable from waiting for a pooled connection; ClickHouse now enforces the cap itself and reports `TIMEOUT_EXCEEDED`. `POST /v1/ops/schema/refresh` against an unreachable ClickHouse is a `503` with `Retry-After` instead of a `500`. **SDK:** `WaveHouseError.code` and `retryable` now take the server's `code`/`retryable` when the body has them (`HTTP_` and "5xx retries" otherwise), so a rejected query is `clickhouse.rejected` rather than `HTTP_500`, and is not retried. diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index a0d21408f..2076e7340 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -101,6 +101,8 @@ When ClickHouse fails a query on [`POST /v1/query`](#post-v1querytabletable--str | 503 | `clickhouse.unavailable` | `true` | ClickHouse, or the way to it, could not take the query now: connection refused or dropped, a timeout, too many queries, memory pressure, lost replicas or Keeper, or a `502`/`503`/`504`/`429`/`408` from a proxy. `Retry-After: 5` | | 500 (`/v1/query`, pipes) / 502 (`/v1/ops/query`) | `clickhouse.unknown` | `true` | A failure with no verdict: no exception code and no recognizable transport error | +A [pipe that writes](/pipes#pipes-that-write) answers with the same status and `code`, but always `retryable: false` and with no `Retry-After`: the statement may have run, so a retry could run it twice. + On `/v1/query`, when the role sets `max_execution_time`, ClickHouse enforces it and reports an overrun as `TIMEOUT_EXCEEDED`, answered `400 clickhouse.limit_exceeded`; WaveHouse then waits two seconds past the cap before giving up itself, and that give-up — like a wait for a pooled connection or a dial timeout — is `503 clickhouse.unavailable`. If the tenant's `clickhouse.query_timeout` is shorter than the cap, that timeout is what ClickHouse enforces, and its overrun is answered as for a role with no cap. When the role sets `max_memory_usage`, every `MEMORY_LIMIT_EXCEEDED` is taken as that cap and answered `400`, even one caused by the server's total memory. Without a role cap of that kind, and always on pipes and `/v1/ops/query`, a timeout or memory limit is `503 clickhouse.unavailable`: it can be the server's state as much as the query's ([#620](https://github.com/Wave-RF/WaveHouse/issues/620)). The classes are the ones the ingest worker uses to decide between retrying a batch and dead-lettering it ([ingest pipeline](/ingest-pipeline#when-clickhouse-cannot-take-an-insert)); the lists of exception codes live in `internal/chconn/errclass.go`. **Why a missing grant is a `403`.** A query path runs as the ClickHouse user in the tenant's settings, not as the caller, so `ACCESS_DENIED` is in one sense WaveHouse's configuration. It is still a verdict on *this statement*: ClickHouse understood it and refused it, the same statement is refused every time, and other statements from the same caller succeed. That is a `403`, and it matters most on `/v1/ops/query`, where the admin wrote the statement — a `CREATE USER` through a user without the grant is the admin asking for something this deployment does not allow. A `5xx` would tell clients and monitors that ClickHouse is down and invite retries of a request that can never pass. Denials that refuse every query, not one statement — the credentials, the user, the database — are the operator's to fix, so they are `502 clickhouse.misconfigured`, still not retryable. @@ -615,7 +617,7 @@ The POST parameter body is capped at 1 MiB; a body over the cap is rejected with | 400 | `{"error":"parameter \"x\": unsupported parameter type object"}` | A non-scalar value with no SQL literal form — a JSON object, whether supplied directly or nested as an array element. A JSON **array** is valid and renders as an `IN`-style `(…)` list. | | 400 | `{"error":"parameter \"x\": array parameter must not be empty"}` | An empty array — it would render as the invalid `IN ()`. | | 413 | `{"error":"request body exceeded 1048576 bytes"}` | POST body over the 1 MiB cap | -| 400 / 403 / 500 / 502 / 503 | `{"error":"clickhouse query: …","code":"clickhouse.…","retryable":…}` | ClickHouse failed the pipe's query — for instance a parameter value it cannot use (`400 clickhouse.rejected`), or ClickHouse down (`503 clickhouse.unavailable`, `Retry-After: 5`); see [ClickHouse errors on the query paths](#clickhouse-errors-on-the-query-paths) | +| 400 / 403 / 500 / 502 / 503 | `{"error":"clickhouse query: …","code":"clickhouse.…","retryable":…}` | ClickHouse failed the pipe's query — for instance a parameter value it cannot use (`400 clickhouse.rejected`), or ClickHouse down (`503 clickhouse.unavailable`, `Retry-After: 5`); see [ClickHouse errors on the query paths](#clickhouse-errors-on-the-query-paths). A [pipe that writes](/pipes#pipes-that-write) answers `retryable: false` with no `Retry-After`, its message led by `clickhouse exec:` | | 503 | `{"error":"token verifier not ready: the tenant's JWKS has not been fetched yet"}` | A token was supplied, with no valid operator key, while the tenant's JWKS has not been fetched yet; refused before any policy runs, with a `Retry-After: 30` header — see [Authentication](#authentication) | --- diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 303c45dd9..e243dfc90 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -82,7 +82,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request - **tenant.go** — `TenantMW` resolves the request's tenant ahead of the auth middleware on every `/v1` route outside `/v1/ops/*`: the [`X-Tenant-ID`](/deployment#multi-tenant-deployments) header (absent means `tenant.Default`), validated by `tenant.Parse` (`400`), looked up in the `settings.Registry` (`resolveStore`: `404` for an id it does not hold, a bare `503` for a tenant whose folder was rejected — the findings stay out of a body answered before authentication), and the resolved `*settings.Store` stored in the request context (`WithStore` / `StoreFromContext` — here rather than in `tenant/`, because `settings` names `tenant.ID`). A handler reads the store once and passes it down as an argument — the per-tenant getters it holds take it as a parameter (`(*settings.Store).Policy`, `.DedupeFor`, `.DefaultMaxRows`, … in production) — and nothing below a handler reads the context; a tenant route reached without a resolved store answers `500` rather than fall back to a tenant. The probes, `/version`, the metrics path, and `/v1/ops/*` are tenant-exempt; an ops route that addresses one tenant — the admin pipe reads, the schema routes, the raw-SQL proxy, the settings reload and the DLQ stats — names it in `?tenant=` (`opsTenant`, and `opsStore` over it for the routes that need the tenant's store; the DLQ stats need none, since the MQ holds the queue), parsed strictly so that a query `url.ParseQuery` would half-read is a `400` rather than a read of the default tenant, which is what it means when absent on the reads (on the reload, absent is the whole directory). - **pipes.go** — Named query pipe handlers: admin listing (`GET /v1/ops/pipes[/{name}]`, read per request from its `pipes.Source`) and execution with parameter binding. A read is cached and coalesced; a write — bound SQL that `isMutation` (`clickhouse_exec.go`) classifies as one — bypasses both and runs every call. `pipes.json` is the only way to define or change a pipe. - **structured_query.go** — Handler for `POST /v1/query?table={table}`: validates query AST, enforces permissions, builds and executes SQL. -- **ch_errors.go** — `writeCHError`, the one mapping from a failed ClickHouse query to a response, shared by `/v1/query`, pipes and `/v1/ops/query` so they cannot drift apart: `chconn.Classify` decides the class, and the class the status, `code` and `retryable` ([ClickHouse errors on the query paths](/api#clickhouse-errors-on-the-query-paths)). +- **ch_errors.go** — `writeCHError`, the one mapping from a failed ClickHouse query to a response, shared by `/v1/query`, pipes and `/v1/ops/query` so they cannot drift apart: `chconn.Classify` decides the class, and the class the status, `code` and `retryable` ([ClickHouse errors on the query paths](/api#clickhouse-errors-on-the-query-paths)). A write pipe answers through `writeCHWriteError`, the same mapping with `retryable` always `false` and no `Retry-After`, since the write may have run. - **ingest.go** — Accepts `POST /v1/ingest?table={table}` in three body shapes: one flat JSON object, a JSON array of them, or NDJSON. The **required** `Content-Type` chooses the format *family* — `application/json` versus the four NDJSON spellings — and within the JSON family the body's first non-whitespace byte picks array versus single object; the bytes never choose the family. Anything that is not exactly one readable media type is a `415`, decided before the body is read: the header is parsed per RFC 9110 §8.3, and because `Content-Type` is a singleton field, repeated header lines must all resolve to the same format and a value carrying a comma is refused unless the value as a whole parses as one media type — a comma inside a *quoted* parameter value is data, so `application/json; a=", application/x-ndjson; b="` is accepted. It then reads the whole (`MaxBytesReader`-capped) body into a pooled buffer and runs the per-format record readers over those bytes, so the `413` lands before any record is processed and peak memory per request is O(body) rather than O(record). Then it validates each record against the discovered schema, optional dedup, and publishes each row through `mq.Publisher` on `mq.Topic{Tenant, Table, Scope}` (the request's tenant, read off its resolved store — `store.Tenant()` — and raw names; the subject it becomes is `internal/mq`'s; a full queue comes back as `mq.ErrQueueFull`, which is the `503` + `Retry-After: 30`, and a broker that cannot be reached or does not answer in time as `mq.ErrUnavailable`, the `503` + `Retry-After: 5`). When dedup is on, a row missing the configured `id_field` can't be deduped: it is logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total` (labeled by `table`), then published un-deduped — or rejected when `dedupe.require_id` is set ([#219](https://github.com/Wave-RF/WaveHouse/issues/219)). - **query.go** — Proxies raw SQL for `POST /v1/ops/query` straight to the `?tenant=`'s ClickHouse HTTP interface (`chconn.Pools.Target` by the resolved store's tenant; the zero target — no pool — is a `503` with `Retry-After`). **Not cached** — sets `Cache-Control: no-store` so every request hits ClickHouse; DateTime is rendered ISO-8601 via `date_time_output_format=iso` (the Go-side type conversion lives in the structured-query / pipes path, not here). - **stream.go** — Real-time streaming via SSE. Callers select a table with the `?table=` query parameter. Each connection registers one `Subscriber` (the `stream/` package) with both the event `Hub` (under its `(topic, role)`) and the shared keepalive wheel, then drains both from a single byte-pump — so idle streams keep emitting `:` keepalive comments (surviving reverse-proxy idle timeouts) while live events arrive already projected and serialized. Per-event projection/serialization happens **once per role** in the `Hub`, not once per subscriber ([#294](https://github.com/Wave-RF/WaveHouse/issues/294)); the handler also snapshots the connection's JWT claims onto the `Subscriber`, which the `Hub` evaluates per subscriber when the role carries a row-level `filter` ([#319](https://github.com/Wave-RF/WaveHouse/issues/319)). Gap-fill replay (`mq.Replayer.ReplaySince` on the connection's `mq.Topic` — a `DeliverByStartTime` consumer inside `internal/mq`) stays per-connection (low-volume, one-time on connect). A stream ends, a gap-fill in progress included, when the server begins shutting down (`Closing`) or its `Subscriber` is evicted because its tenant is no longer served (`Hub.Prune`); one admitted just before the reload that stopped serving its tenant, and registered just after the prune, is ended right after it registers (`Served`). diff --git a/docs/src/content/docs/pipes.mdx b/docs/src/content/docs/pipes.mdx index e8a6ea768..f06d1ebe0 100644 --- a/docs/src/content/docs/pipes.mdx +++ b/docs/src/content/docs/pipes.mdx @@ -182,6 +182,8 @@ The response is a JSON array of rows. Results flow through the shared in-process A pipe's SQL may be a write: a statement led by a write verb WaveHouse recognizes — `INSERT`, `UPDATE`, `DELETE`, `ALTER` (so `ALTER … DELETE`), `CREATE`, `DROP`, `TRUNCATE`, `RENAME`, `EXCHANGE`, `REPLACE`, `OPTIMIZE`, `ATTACH`, `DETACH`, `GRANT`, `REVOKE`, `KILL`, `SET`, `USE` or `SYSTEM` — directly or after a `WITH` list, as in `WITH … INSERT`. Such a pipe runs on **every** call: it never reads or fills the cache and is never coalesced with an identical call in flight, so ten identical calls are ten writes. The response is `[]` with `X-Cache: BYPASS` and `Cache-Control: no-store`, so an HTTP cache in front of a `GET` does not answer a repeat either. WaveHouse classifies the statement from its leading keyword (after any `WITH` list), with the same classifier that sends it to ClickHouse as a write, so no pipe property marks it; a statement led by any other keyword runs as a read. +A failed write is not retried automatically, because it may have run. It answers with the status and `code` a failed read would ([ClickHouse errors on the query paths](/api#clickhouse-errors-on-the-query-paths)), but always with `retryable: false` and no `Retry-After`, `503 clickhouse.unavailable` included: once the statement is on its way to ClickHouse, WaveHouse cannot tell whether it ran. The [SDK](/sdk/pipes) does not retry such an answer, so check whether the write landed before you send it again. A call refused before anything is sent — the tenant on no ClickHouse pool, `503` with `Retry-After: 30` — cannot have run, and the SDK retries it. The SDK also retries a request whose answer never arrives, such as one on a dropped connection, so a write can still run twice that way; give a client that runs write pipes [`options.maxRetries`](/sdk#clientconfigdb) `0` if that matters. + `allowed_roles` is a write pipe's only gate: the [policy engine](/access-control)'s insert rules do not apply to it, so any role you list — including a [`default_role`](/access-control#default_role--public-unauthenticated-access) that anonymous callers resolve to — can run the write. The operator fixes the statement and its predicate when authoring the pipe; callers supply only literal values, provided every placeholder is written bare ([how a value becomes SQL](#how-a-value-becomes-sql)). Two things a write pipe does not do yet: it does not invalidate cached reads of the table it writes — a structured query or read pipe over that table can serve pre-write rows until its TTL ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)) — and its rows do not reach [`/v1/stream`](/api#get-v1stream--server-sent-events-stream) subscribers, which only the [ingest pipeline](/ingest-pipeline) feeds ([#362](https://github.com/Wave-RF/WaveHouse/issues/362)). For writes that should be seen at once, use [`POST /v1/ingest`](/api#post-v1ingesttabletable--ingest-data). diff --git a/docs/src/content/docs/sdk/pipes.md b/docs/src/content/docs/sdk/pipes.md index 1bae86bf0..81bb77d22 100644 --- a/docs/src/content/docs/sdk/pipes.md +++ b/docs/src/content/docs/sdk/pipes.md @@ -17,7 +17,7 @@ const { data } = await wh.pipe('top_pages', { start_date: '2026-01-01', limit: 5 ### `.fetch(opts?)` -Execute and return results. Takes `PipeRequestOptions` — `{ signal }` only, narrower than the `.fetch(opts?)` on a [query builder](/sdk/queries), which also accepts `limit`. Passing a `limit` is a compile error rather than a silent no-op. +Execute and return results. A [pipe that writes](/pipes#pipes-that-write) returns `[]`, and a failed one comes back `retryable: false`, which the SDK does not retry. Takes `PipeRequestOptions` — `{ signal }` only, narrower than the `.fetch(opts?)` on a [query builder](/sdk/queries), which also accepts `limit`. Passing a `limit` is a compile error rather than a silent no-op. `limit` is typed `never` rather than left out, so the rejection also catches a value passed in a variable — leaving it out would only reject an inline object. That cuts both ways: a value *declared* as `RequestOptions` is rejected whether or not it actually carries a limit, since the type permits one. If you share one options object across calls, type it as `PipeRequestOptions` — the table and query-builder `.fetch()` accept that too — or inline `{ signal }` at the pipe call. diff --git a/internal/api/ch_errors.go b/internal/api/ch_errors.go index e2d6c8e75..bf5e3eb41 100644 --- a/internal/api/ch_errors.go +++ b/internal/api/ch_errors.go @@ -119,10 +119,24 @@ func chFailureOf(err error, unknownStatus int, caps queryCaps) chFailure { // it is WaveHouse's configuration being refused, which an operator should // hear about even when the caller only sees a 403. func writeCHError(w http.ResponseWriter, r *http.Request, err error, message string, unknownStatus int, caps queryCaps) { - f := chFailureOf(err, unknownStatus, caps) + writeCHFailure(w, r, err, message, chFailureOf(err, unknownStatus, caps)) +} + +// writeCHWriteError answers a failed write pipe as writeCHError does, but +// never as retryable and with no Retry-After: the statement may have reached +// ClickHouse and run, so a client retrying would run it again. +func writeCHWriteError(w http.ResponseWriter, r *http.Request, err error, message string) { + f := chFailureOf(err, http.StatusInternalServerError, queryCaps{}) + f.retryable = false + writeCHFailure(w, r, err, message, f) +} + +func writeCHFailure(w http.ResponseWriter, r *http.Request, err error, message string, f chFailure) { switch f.code { case codeCHUnavailable: - w.Header().Set("Retry-After", retryAfterClickHouse) + if f.retryable { + w.Header().Set("Retry-After", retryAfterClickHouse) + } case codeCHAccessDenied, codeCHMisconfigured: exCode, _ := chconn.ExceptionCode(err) slog.WarnContext(r.Context(), "clickhouse refused WaveHouse's configuration", diff --git a/internal/api/ch_errors_test.go b/internal/api/ch_errors_test.go index 6daf11ea5..ee2df2db6 100644 --- a/internal/api/ch_errors_test.go +++ b/internal/api/ch_errors_test.go @@ -149,6 +149,32 @@ func TestPipes_ClickHouseErrors(t *testing.T) { } } +// TestPipes_WriteClickHouseErrors: a failed write pipe answers with its +// class's status and code, but never as retryable and with no Retry-After: +// the statement may have run, so a client that retried would run it again. +func TestPipes_WriteClickHouseErrors(t *testing.T) { + t.Parallel() + for _, tc := range chErrorCases(t) { + if tc.caps.MaxExecutionTime > 0 || tc.caps.MaxMemoryUsage > 0 { + continue + } + t.Run(tc.name, func(t *testing.T) { + t.Parallel() + conn := &writeConn{err: tc.err} + h := writerPipesHandler(t, conn, nil, &pipes.NamedQuery{Name: "log", SQL: "INSERT INTO audit_log VALUES ({{msg}}, now())"}) + w := pipeCallAs(t, h, "log") + require.Equal(t, tc.wantStatus, w.Code, w.Body.String()) + var got errorBody + require.NoError(t, json.Unmarshal(w.Body.Bytes(), &got)) + assert.Equal(t, tc.wantCode, got.Code) + require.NotNil(t, got.Retryable) + assert.False(t, *got.Retryable) + assert.Empty(t, w.Header().Get("Retry-After")) + assert.Equal(t, int32(1), conn.execs.Load()) + }) + } +} + type errRow struct{ err error } func (r errRow) Err() error { return r.err } diff --git a/internal/api/pipes.go b/internal/api/pipes.go index 54cc0236e..9e28ea468 100644 --- a/internal/api/pipes.go +++ b/internal/api/pipes.go @@ -223,7 +223,7 @@ func (h *PipesHandler) executeWrite(w http.ResponseWriter, r *http.Request, stor } data, _, err := h.run(r.Context(), store, conn, sql, params) if err != nil { - writeJSONError(w, http.StatusInternalServerError, err.Error()) + writeCHWriteError(w, r, err, err.Error()) return } w.Header().Set("Content-Type", "application/json") diff --git a/internal/api/pipes_test.go b/internal/api/pipes_test.go index 3211c57c9..f3a163d44 100644 --- a/internal/api/pipes_test.go +++ b/internal/api/pipes_test.go @@ -517,13 +517,14 @@ func TestPipesHandler_Execute_NoAllowedRoles_AdminAllowed(t *testing.T) { assert.NotEqual(t, http.StatusNotFound, w.Code) } -// writeConn counts Exec and Query calls. With gate set, every Exec reports -// itself on entered and holds until gate is closed, so a test can hold -// requests in flight together. +// writeConn counts Exec and Query calls, and every Exec returns err. With +// gate set, every Exec reports itself on entered and holds until gate is +// closed, so a test can hold requests in flight together. type writeConn struct { driver.Conn execs, queries atomic.Int32 entered, gate chan struct{} + err error } func (c *writeConn) Exec(context.Context, string, ...any) error { @@ -532,7 +533,7 @@ func (c *writeConn) Exec(context.Context, string, ...any) error { c.entered <- struct{}{} <-c.gate } - return nil + return c.err } func (c *writeConn) Query(context.Context, string, ...any) (driver.Rows, error) { diff --git a/internal/api/tenant_clickhouse_test.go b/internal/api/tenant_clickhouse_test.go index 8181baa2d..781b3e4a3 100644 --- a/internal/api/tenant_clickhouse_test.go +++ b/internal/api/tenant_clickhouse_test.go @@ -135,6 +135,15 @@ func TestClickHouseRoutes_NoPoolIs503(t *testing.T) { h.Execute(w, withTenant(pipesRequest(t, http.MethodGet, "/v1/pipes/top_pages", "top_pages", nil))) assertUnavailable(t, w, noConnectionMessage, retryAfterPool) }) + // A write that never reached ClickHouse cannot have run, so this 503 + // keeps its Retry-After where a failed write's answer drops it. + t.Run("write pipe execute", func(t *testing.T) { + t.Parallel() + h := NewPipesHandler(staticPipes(&pipes.NamedQuery{Name: "log", SQL: "INSERT INTO audit_log VALUES (1)", AllowedRoles: []string{"viewer"}}), allowAll, noConn, nil, noTimeout) + w := httptest.NewRecorder() + h.Execute(w, withTenant(pipesRequest(t, http.MethodGet, "/v1/pipes/log", "log", nil))) + assertUnavailable(t, w, noConnectionMessage, retryAfterPool) + }) t.Run("raw-SQL proxy", func(t *testing.T) { t.Parallel() h := newTestQueryHandler(func(*settings.Store) chconn.Target { return chconn.Target{} }, noTimeout) From 0c31dc5a522d2ea730a6537158b6b10a7c0dbd05 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:44:04 -0400 Subject: [PATCH 46/79] fix(api): lex //, curly quotes and _-led words as ClickHouse does Review findings, each checked on ClickHouse 26.6.3.62: - `//` starts a line comment, so a `(` after it hid the INSERT behind a WITH list, and a leading one hid the verb. - ClickHouse reads a string literal in curly single quotes and a quoted identifier in curly double quotes, with no escapes inside; a paren in one hid a WITH's INSERT, and a verb in one looked like a write. - A bareword led by `_` (`_delete`, `_set`) was read from its second byte, so its tail matched a verb and a read ran through Exec. Words now start at any word byte. A heredoc tag stays letters, digits and `_`: ClickHouse rejects `$a b$ ... $a b$`. Docs: the SDK pipes page and error reference say a failed write pipe is returned on the first attempt, and that a dropped connection is still retried; configuration.mdx and ingest-pipeline.md name write pipes; "only write path" becomes "only way to define or change a pipe" on the SDK pipes page and in the PipesNamespace doc comment. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- CHANGELOG.md | 4 +- clients/ts/src/pipes.ts | 4 +- docs/src/content/docs/configuration.mdx | 2 +- docs/src/content/docs/ingest-pipeline.md | 2 +- docs/src/content/docs/sdk/pipes.md | 4 +- docs/src/content/docs/sdk/reference.md | 2 +- internal/api/clickhouse_exec.go | 55 ++++++++++++++++++------ internal/api/clickhouse_exec_test.go | 15 +++++++ 8 files changed, 67 insertions(+), 21 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 9ed3da2cb..06279b65a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -85,8 +85,8 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed -- **A pipe that writes runs on every call instead of being answered from the cache** (`internal/api/{pipes,ch_errors}.go` (+ tests), `docs/src/content/docs/{pipes.mdx,api.md,architecture.md,sdk/pipes.md}`, `AGENTS.md`): fixes [#386](https://github.com/Wave-RF/WaveHouse/issues/386). `/v1/pipes/{name}` sent a write's SQL to ClickHouse through `Exec`, but still cached the `[]` it returned and coalesced identical calls in flight, so a repeat within the TTL answered `200` without executing and concurrent identical calls became one write — silently dropped writes, and with a shared cache ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)) on every instance. A pipe whose bound SQL `isMutation` classifies as a write — the same classifier that picks `Exec` — now skips the cache lookup, the fill and singleflight, and answers `X-Cache: BYPASS` with `Cache-Control: no-store`, so an HTTP cache in front of a `GET` cannot drop the write either. Classification stays automatic rather than a declared pipe property, so an operator cannot forget to mark one, and costs no ClickHouse round trip. A failed write answers with the status and `code` a failed read gets (see the ClickHouse-errors entry below), but always `retryable: false` and with no `Retry-After`, `503 clickhouse.unavailable` included: the statement may have run, so the SDK does not retry it. A write refused before it is sent, the tenant on no pool, keeps its `503` with `Retry-After: 30`. Read pipes are unchanged. Not in this fix: a write pipe still does not invalidate cached reads of the table it writes ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). -- **The write classifier skips whitespace, comments and quoted text the way ClickHouse's lexer does** (`internal/api/clickhouse_exec.go` (+ tests)): `isMutation` picks `Exec` for a write, and since [#386](https://github.com/Wave-RF/WaveHouse/issues/386) keeps a write pipe out of the cache. It missed a write behind a backslash-escaped quote (`'it\'s'`, and the same inside `"…"` and `` `…` ``), a heredoc (`$$ ( $$`, `$tag$ … $tag$`), a nested block comment (`/* a /* b */ SELECT */ INSERT …`), or behind leading whitespace other than space, tab, CR and LF: `\v`, `\f`, a no-break space, a byte-order mark, and the other Unicode spaces ClickHouse skips. A missed write went through `Query`, which ran it and then failed the call with a `5xx` the TypeScript SDK retries, so one call could write three times. The same gaps could make a read look like a write, which runs through `Exec` and answers `[]`. +- **A pipe that writes runs on every call instead of being answered from the cache** (`internal/api/{pipes,ch_errors}.go` (+ tests), `docs/src/content/docs/{pipes.mdx,api.md,architecture.md,configuration.mdx,ingest-pipeline.md,sdk/pipes.md,sdk/reference.md}`, `clients/ts/src/pipes.ts` (doc comment), `AGENTS.md`): fixes [#386](https://github.com/Wave-RF/WaveHouse/issues/386). `/v1/pipes/{name}` sent a write's SQL to ClickHouse through `Exec`, but still cached the `[]` it returned and coalesced identical calls in flight, so a repeat within the TTL answered `200` without executing and concurrent identical calls became one write — silently dropped writes, and with a shared cache ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)) on every instance. A pipe whose bound SQL `isMutation` classifies as a write — the same classifier that picks `Exec` — now skips the cache lookup, the fill and singleflight, and answers `X-Cache: BYPASS` with `Cache-Control: no-store`, so an HTTP cache in front of a `GET` cannot drop the write either. Classification stays automatic rather than a declared pipe property, so an operator cannot forget to mark one, and costs no ClickHouse round trip. A failed write answers with the status and `code` a failed read gets (see the ClickHouse-errors entry below), but always `retryable: false` and with no `Retry-After`, `503 clickhouse.unavailable` included: the statement may have run, so the SDK does not retry it. A write refused before it is sent, the tenant on no pool, keeps its `503` with `Retry-After: 30`. Read pipes are unchanged. Not in this fix: a write pipe still does not invalidate cached reads of the table it writes ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). +- **The write classifier skips whitespace, comments and quoted text the way ClickHouse's lexer does** (`internal/api/clickhouse_exec.go` (+ tests)): `isMutation` picks `Exec` for a write, and since [#386](https://github.com/Wave-RF/WaveHouse/issues/386) keeps a write pipe out of the cache. It missed a write behind a backslash-escaped quote (`'it\'s'`, and the same inside `"…"` and `` `…` ``), a heredoc (`$$ ( $$`, `$tag$ … $tag$`), a curly-quoted literal or identifier (`‘(’`, `“c(d”`), a `//` line comment, a nested block comment (`/* a /* b */ SELECT */ INSERT …`), or leading whitespace other than space, tab, CR and LF: `\v`, `\f`, a no-break space, a byte-order mark, and the other Unicode spaces ClickHouse skips. A missed write went through `Query`, which ran it and then failed the call with a `5xx` the TypeScript SDK retries, so one call could write three times. The same gaps, and a word led by `_` (`_delete`) whose tail was read as a verb, could make a read look like a write, which runs through `Exec` and answers `[]`. - **The pipes page no longer says a parameter can never break out of its literal** (`docs/src/content/docs/pipes.mdx`, `internal/pipes/pipes.go`): that holds only for a placeholder written bare. A string value brings its own quotes, so inside a quoted placeholder they close the template's: the body `{"id": " OR 1=1 OR id = "}` turns `WHERE id = '{{id}}'` into `WHERE id = '' OR 1=1 OR id = ''`, which matches every row. The page now says to write each placeholder bare, never inside quotes. - **A failed ClickHouse query answers by what went wrong, not a flat `500`/`502`** (`internal/api/ch_errors.go` (new, + tests), `internal/api/{errors,query,structured_query,pipes,schema,ch_settings}.go`, `internal/chconn/errclass.go` (`HTTPStatus` exported), `clients/ts/src/errors.ts` (+ tests), `tests/integration/query_errors_test.go` (new), `tests/integration/query_limits_test.go`, `internal/app/app_test.go`, `tests/e2e/sdk/{admin,query}.test.ts`, `AGENTS.md`, `docs/src/content/docs/{api,architecture}.md`, `docs/src/content/docs/{access-control,configuration}.mdx`, `docs/src/content/docs/sdk/{reference.md,index.mdx}`): fixes [#403](https://github.com/Wave-RF/WaveHouse/issues/403) and [#271](https://github.com/Wave-RF/WaveHouse/issues/271), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). ClickHouse answers a syntax error, a missing grant and an overloaded server alike with HTTP `500`, so `/v1/ops/query` turned a bad statement into a `502` and `/v1/query` and pipes into a `500` the SDK retried. All three now class the failure with `chconn.Classify` through one helper, `writeCHError`: a statement ClickHouse refused is `400 clickhouse.rejected`; a query over a rows/bytes limit, the role's own memory cap, or its time cap where that is no longer than `query_timeout` is `400 clickhouse.limit_exceeded`; `ACCESS_DENIED` is `403 clickhouse.access_denied`; credentials, user or database refused, or a redirect or `4xx` with no exception code from whatever fronts ClickHouse, is `502 clickhouse.misconfigured`; ClickHouse down, unreachable or overloaded is `503 clickhouse.unavailable` with `Retry-After: 5`; a failure with no verdict stays `500` (`502` on the proxy) as `clickhouse.unknown`. The error envelope gains `code` and `retryable` next to `error` on these responses — additive. A role with `max_execution_time` now queries with no context deadline and a cancel two seconds past the cap instead: clickhouse-go overwrote the cap's `max_execution_time` with deadline+5s for any deadline over 1s, so an overrun came back as a bare deadline, indistinguishable from waiting for a pooled connection; ClickHouse now enforces the cap itself and reports `TIMEOUT_EXCEEDED`. `POST /v1/ops/schema/refresh` against an unreachable ClickHouse is a `503` with `Retry-After` instead of a `500`. **SDK:** `WaveHouseError.code` and `retryable` now take the server's `code`/`retryable` when the body has them (`HTTP_` and "5xx retries" otherwise), so a rejected query is `clickhouse.rejected` rather than `HTTP_500`, and is not retried. - **An unavailable ClickHouse is retried with backoff instead of dead-lettering every row** (`internal/chconn/errclass.go` (new, + tests), `internal/ingest/{worker,backoff}.go` (`backoff.go` new, + tests), `internal/mq/{mq,embedded}.go`, `internal/testutil/mocks.go`, `tests/integration/ingest_outage_test.go` (new), `AGENTS.md`, `README.md`, `docs/src/content/docs/{ingest-pipeline,architecture,api,deployment,why-wavehouse}.md`, `docs/src/content/docs/{settings-directory,index,access-control}.mdx`): workstream A of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). A failed batch insert used to go through row-by-row isolation whatever the failure, so a ClickHouse that was down, overloaded or read-only failed every row twice and parked the whole batch on the DLQ. `chconn.Classify` now classes the failure first — `Rejected` (any ClickHouse exception code outside the availability and credential lists: the server read the row and refused it), `Unavailable` (connection refused/reset, timeouts, `TOO_MANY_SIMULTANEOUS_QUERIES`, `SERVER_OVERLOADED`, `MEMORY_LIMIT_EXCEEDED`, `TOO_MANY_PARTS`, `READONLY`, `TABLE_IS_READ_ONLY`, `KEEPER_EXCEPTION`, …), `Denied` (`AUTHENTICATION_FAILED`, `ACCESS_DENIED`, …) or `Unknown` (no code, no recognizable transport failure). Only `Rejected` is isolated and dead-lettered as before, and a multi-row batch refused with `TOO_MANY_PARTS` or `MEMORY_LIMIT_EXCEEDED` is split row by row first (`chconn.Splittable`), because a batch spanning too many partitions or too much memory can fail where each of its rows inserts; every other class hands the batch back to the queue with a delayed nak (`mq.Message.NakWithDelay`, new) under a jittered 1 s → 30 s backoff shared by every table on the same ClickHouse pool (a failure of one table — read-only, too many parts or mutations, a grant missing on it, `chconn.TableScoped` — backs off that table alone), which turns rows away without a request while it runs and probes once per window, and ClickHouse going away mid-isolation stops isolation and retries the rows it had not settled. Counted by the new `wavehouse_ingest_retries_total{table, reason}`; logged at `WARN` when an outage starts and at most every 30 s during it. A long outage now shows as a growing ingest stream and, at `mq.max_bytes_gb`, ingest `503`s — not as a full DLQ; a lasting failure of one table holds back its tenant's other tables once its waiting rows reach `maxAckPending`. Retried rows come back out of arrival order, which matters only to a `ReplacingMergeTree` without a version column or a `CollapsingMergeTree`. diff --git a/clients/ts/src/pipes.ts b/clients/ts/src/pipes.ts index 0297e8511..7fba4cc3a 100644 --- a/clients/ts/src/pipes.ts +++ b/clients/ts/src/pipes.ts @@ -67,8 +67,8 @@ export class PipeRef> implements PromiseLike ``` -What a caller sees when one of these trips: a row or byte limit is `400 clickhouse.limit_exceeded`; a server-wide time, memory or quota limit is `503 clickhouse.unavailable` (retryable, so the SDK retries it), since WaveHouse cannot tell it from server pressure — except on `/v1/query` for a role that sets its own `max_memory_usage`, where every memory-limit error is read as that cap and answered `400 clickhouse.limit_exceeded`. See [ClickHouse errors on the query paths](/api#clickhouse-errors-on-the-query-paths). +What a caller sees when one of these trips: a row or byte limit is `400 clickhouse.limit_exceeded`; a server-wide time, memory or quota limit is `503 clickhouse.unavailable` (retryable, so the SDK retries it — except on a [pipe that writes](/pipes#pipes-that-write), which is never retryable), since WaveHouse cannot tell it from server pressure — except on `/v1/query` for a role that sets its own `max_memory_usage`, where every memory-limit error is read as that cap and answered `400 clickhouse.limit_exceeded`. See [ClickHouse errors on the query paths](/api#clickhouse-errors-on-the-query-paths). :::caution[How the two layers compose] WaveHouse's per-role caps are sent as per-query `SETTINGS` on its connection, so they **compose** with the ClickHouse profile — a per-role cap *tightens* within the profile's ceiling, and a `` block bounds how far any setting can move. But if the profile marks a setting `readonly` (or `` disallows changing it), ClickHouse will **reject** WaveHouse's per-query override and the query fails. So keep the settings WaveHouse manages (`max_memory_usage`, `max_execution_time`, `max_rows_to_read`, `max_result_rows`) **changeable** for its user — use a `` constraint, not `readonly`, if you want a hard ceiling. diff --git a/docs/src/content/docs/ingest-pipeline.md b/docs/src/content/docs/ingest-pipeline.md index f27356c0e..77bb97d24 100644 --- a/docs/src/content/docs/ingest-pipeline.md +++ b/docs/src/content/docs/ingest-pipeline.md @@ -19,7 +19,7 @@ It is deliberately detailed: this is a hot, concurrency-heavy path, and the goro | `sweeper.go` | The **Active Sweeper** — every minute, asks the MQ to purge the events that are both written to ClickHouse and past the SSE gap window (the purge arithmetic below lives in `internal/mq/purge.go`) | | `types.go` | `EventMessage` wire format and the `BufferConsumerName` constant | -The pipeline is **insert-only**. (Upgrading across the v2 envelope? [Drain the queue first](/deployment#upgrading-across-the-v2-ingest-envelope).) The wire format carries `{table_name, scope, received_timestamp, format, columns, row}`: `row` is one `JSONCompactEachRow` line — a positional JSON array — and `columns` names its positions — the table's insertable columns, in declaration order (a `MATERIALIZED` or `ALIAS` column cannot be named in an `INSERT`, so it is not part of the row's contract). (`scope` is reserved and always `""` today.) Each NATS message is its own envelope, so the names ride along per record; where they are carried once is the `INSERT` the worker emits per group. The worker parses the envelope, groups a batch by column list, and bulk-`INSERT`s each group as `INSERT INTO … (cols) FORMAT JSONCompactEachRow` — schema validation already happened at the HTTP ingest handler, before publish. Non-insert mutations go through `POST /v1/ops/query` (admin-only). +The pipeline is **insert-only**. (Upgrading across the v2 envelope? [Drain the queue first](/deployment#upgrading-across-the-v2-ingest-envelope).) The wire format carries `{table_name, scope, received_timestamp, format, columns, row}`: `row` is one `JSONCompactEachRow` line — a positional JSON array — and `columns` names its positions — the table's insertable columns, in declaration order (a `MATERIALIZED` or `ALIAS` column cannot be named in an `INSERT`, so it is not part of the row's contract). (`scope` is reserved and always `""` today.) Each NATS message is its own envelope, so the names ride along per record; where they are carried once is the `INSERT` the worker emits per group. The worker parses the envelope, groups a batch by column list, and bulk-`INSERT`s each group as `INSERT INTO … (cols) FORMAT JSONCompactEachRow` — schema validation already happened at the HTTP ingest handler, before publish. Non-insert mutations go through `POST /v1/ops/query` (admin-only) or an operator-authored [pipe that writes](/pipes#pipes-that-write). ## High-level shape diff --git a/docs/src/content/docs/sdk/pipes.md b/docs/src/content/docs/sdk/pipes.md index 81bb77d22..685cfc34d 100644 --- a/docs/src/content/docs/sdk/pipes.md +++ b/docs/src/content/docs/sdk/pipes.md @@ -17,7 +17,7 @@ const { data } = await wh.pipe('top_pages', { start_date: '2026-01-01', limit: 5 ### `.fetch(opts?)` -Execute and return results. A [pipe that writes](/pipes#pipes-that-write) returns `[]`, and a failed one comes back `retryable: false`, which the SDK does not retry. Takes `PipeRequestOptions` — `{ signal }` only, narrower than the `.fetch(opts?)` on a [query builder](/sdk/queries), which also accepts `limit`. Passing a `limit` is a compile error rather than a silent no-op. +Execute and return results. A [pipe that writes](/pipes#pipes-that-write) returns `[]`. A ClickHouse failure on one comes back `retryable: false`, and the SDK does not retry it; it does still retry a request whose answer never arrives, such as one on a dropped connection, so a write can run twice. If that matters, give the client that runs write pipes [`options.maxRetries`](/sdk#clientconfigdb) `0`. Takes `PipeRequestOptions` — `{ signal }` only, narrower than the `.fetch(opts?)` on a [query builder](/sdk/queries), which also accepts `limit`. Passing a `limit` is a compile error rather than a silent no-op. `limit` is typed `never` rather than left out, so the rejection also catches a value passed in a variable — leaving it out would only reject an inline object. That cuts both ways: a value *declared* as `RequestOptions` is rejected whether or not it actually carries a limit, since the type permits one. If you share one options object across calls, type it as `PipeRequestOptions` — the table and query-builder `.fetch()` accept that too — or inline `{ signal }` at the pipe call. @@ -31,7 +31,7 @@ Open a live stream (see [Streaming](/sdk/streaming)). ## Pipes Admin — `wh.pipes` -Inspect the adopted named query pipes. Requires the admin gate — the admin role (`policy.admin_role`) or the [operator key](/api#authentication). Pipes are defined in the server's settings directory `pipes.json` — files are the only write path, so there is no `set` or `delete`: edit the file and let the server pick it up, or call [`wh.settings.reload()`](/sdk/admin#settings--whsettings). +Inspect the adopted named query pipes. Requires the admin gate — the admin role (`policy.admin_role`) or the [operator key](/api#authentication). Pipes are defined in the server's settings directory `pipes.json` — the files are the only way to define or change a pipe, so there is no `set` or `delete`: edit the file and let the server pick it up, or call [`wh.settings.reload()`](/sdk/admin#settings--whsettings). ```ts // List all pipes diff --git a/docs/src/content/docs/sdk/reference.md b/docs/src/content/docs/sdk/reference.md index fa6983da4..9db138f3f 100644 --- a/docs/src/content/docs/sdk/reference.md +++ b/docs/src/content/docs/sdk/reference.md @@ -25,7 +25,7 @@ if (error?.code === 'ABORTED') { The SDK **never throws** for anything the server returns — all API errors come back in `Result.error`. It does throw on caller and environment errors: a non-absolute `baseURL` (REST calls reject with a `TypeError`; streams report `SSE_CONNECT_ERROR` to the subscriber's `error` callback — see [Serving under a path prefix](/sdk#serving-under-a-path-prefix)), `.stream()` / `.liveQuery()` in a runtime with no global `fetch` and no `options.fetch` (see [Runtime support](/sdk#runtime-support)), and an `auth` callback that rejects — a token-refresh failure propagates out of the REST call, and on a stream is reported as a retryable `SSE_AUTH_ERROR`. One more exception escapes an SDK call synchronously, though it is yours rather than ours: your own `status` handler throwing on the first `.subscribe()` or `.liveQuery()`, described under *If your own callback throws* below. -`code` and `retryable` are the server's own when its error body carries them — a failed ClickHouse query does, with codes like `clickhouse.rejected` and `clickhouse.unavailable` ([the full list](/api#clickhouse-errors-on-the-query-paths)). Otherwise `code` is `HTTP_` and a `5xx` is retryable. +`code` and `retryable` are the server's own when its error body carries them — a failed ClickHouse query does, with codes like `clickhouse.rejected` and `clickhouse.unavailable` ([the full list](/api#clickhouse-errors-on-the-query-paths)). Otherwise `code` is `HTTP_` and a `5xx` is retryable. One exception to the table below: a [pipe that writes](/pipes#pipes-that-write) answers every ClickHouse failure `retryable: false` with no `Retry-After`, `clickhouse.unavailable` and `clickhouse.unknown` included, so the SDK returns it on the first attempt. | Status | Code | Retryable | Description | |--------|------|-----------|-------------| diff --git a/internal/api/clickhouse_exec.go b/internal/api/clickhouse_exec.go index c86d88a40..e77e8ac2b 100644 --- a/internal/api/clickhouse_exec.go +++ b/internal/api/clickhouse_exec.go @@ -164,12 +164,12 @@ var nonMutationVerbs = map[string]struct{}{ // containsMutationVerbAtTopLevel scans s for the statement-introducing // keyword at paren-depth 0, stepping over string literals and quoted -// identifiers (skipQuoted), heredocs (skipHeredoc), parenthesized CTE -// subqueries, and comments (skipComment). The CTE list contains ordinary -// identifiers (CTE names, table/database names) that must not be matched as -// mutation verbs — `system` would otherwise pattern-match `SYSTEM` and route a -// `WITH … SELECT * FROM system.tables` read through `Exec` (silent empty- -// array result instead of the actual rows). Two-part fix: +// identifiers (skipQuoted, skipCurlyQuoted), heredocs (skipHeredoc), +// parenthesized CTE subqueries, and comments (skipComment). The CTE list +// contains ordinary identifiers (CTE names, table/database names) that must +// not be matched as mutation verbs — `system` would otherwise pattern-match +// `SYSTEM` and route a `WITH … SELECT * FROM system.tables` read through +// `Exec` (silent empty-array result instead of the actual rows). Two-part fix: // // 1. Skip identifiers whose next non-whitespace, non-comment token is // `AS` (case-insensitive) or `(` — those are CTE definition names @@ -203,6 +203,14 @@ func containsMutationVerbAtTopLevel(s string) bool { i++ case c == '\'' || c == '"' || c == '`': i = skipQuoted(s, i) + case c == 0xE2: + // ‘…’ or “…”; any other character led by this byte is stepped + // over a byte at a time, like the default. + if j := skipCurlyQuoted(s, i); j > i { + i = j + } else { + i++ + } case c == '$': // A heredoc, else a bareword led by `$` (never a keyword) or a // lone `$`. @@ -211,7 +219,9 @@ func containsMutationVerbAtTopLevel(s string) bool { } else { i = skipWord(s, i+1) } - case (c >= 'A' && c <= 'Z') || (c >= 'a' && c <= 'z'): + case isWordByte(c): + // A word led by a digit or `_` is read whole, so its tail is + // never taken for a keyword (`_delete`). start := i i = skipWord(s, i) if depth == 0 { @@ -238,7 +248,7 @@ func containsMutationVerbAtTopLevel(s string) bool { return true } } - case c == '-' && i+1 < len(s) && s[i+1] == '-', c == '#', c == '/' && i+1 < len(s) && s[i+1] == '*': + case c == '-' && i+1 < len(s) && s[i+1] == '-', c == '#', c == '/' && i+1 < len(s) && (s[i+1] == '*' || s[i+1] == '/'): i = skipComment(s, i) default: i++ @@ -341,12 +351,12 @@ func sqlSpaceLen(s string, i int) int { } // skipComment returns the index just past the comment starting at s[i], or i -// if none starts there: `--` and MySQL-compat `#` to end of line, `/* … */` -// nesting as ClickHouse's do. An unclosed block comment runs to the end, as -// it does for ClickHouse, which then rejects the statement. +// if none starts there: `--`, `//` and MySQL-compat `#` to end of line, +// `/* … */` nesting as ClickHouse's do. An unclosed block comment runs to the +// end, as it does for ClickHouse, which then rejects the statement. func skipComment(s string, i int) int { switch { - case strings.HasPrefix(s[i:], "--"), s[i] == '#': + case strings.HasPrefix(s[i:], "--"), strings.HasPrefix(s[i:], "//"), s[i] == '#': if j := strings.IndexByte(s[i:], '\n'); j >= 0 { return i + j + 1 } @@ -373,6 +383,27 @@ func skipComment(s string, i int) int { return i } +// skipCurlyQuoted returns the index just past a string literal in ‘…’ or a +// quoted identifier in “…”, which ClickHouse reads so that SQL pasted from a +// word processor parses, or i if none opens at s[i]. Nothing escapes inside +// them; an unclosed one runs to the end. +func skipCurlyQuoted(s string, i int) int { + var closer string + switch { + case strings.HasPrefix(s[i:], "\u2018"): + closer = "\u2019" + case strings.HasPrefix(s[i:], "\u201c"): + closer = "\u201d" + default: + return i + } + start := i + len(closer) // the opener is as long as its closer + if k := strings.Index(s[start:], closer); k >= 0 { + return start + k + len(closer) + } + return len(s) +} + // skipQuoted returns the index just past the string literal or quoted // identifier opening at s[i] (`'`, `"` or backtick). As in ClickHouse's lexer, // a doubled quote or a backslash escapes the next byte; an unclosed one runs diff --git a/internal/api/clickhouse_exec_test.go b/internal/api/clickhouse_exec_test.go index 86ad6c6e7..7cd861421 100644 --- a/internal/api/clickhouse_exec_test.go +++ b/internal/api/clickhouse_exec_test.go @@ -128,6 +128,21 @@ func TestIsMutation(t *testing.T) { {"with tagged heredoc holding a paren and a verb then select", "WITH $x$ ) INSERT $x$ AS s SELECT s", false}, {"with CTE alias set$ (read)", "WITH set$ AS (SELECT 1 AS v) SELECT * FROM set$", false}, + // A word led by `_` is one bareword, never a keyword's tail. + {"with alias _delete (read)", "WITH 1 AS _delete SELECT _delete", false}, + {"with alias _set (read)", "WITH [1,2] AS _set SELECT has(_set, 1)", false}, + + // `//` starts a line comment. + {"slash comment then insert", "// note\nINSERT INTO t VALUES (1)", true}, + {"slash comment hiding insert then select", "// INSERT\nSELECT 1", false}, + {"with slash comment holding a paren then insert", "WITH x AS (SELECT 'a' AS s) // (\nINSERT INTO t SELECT * FROM x", true}, + + // ‘…’ is a string literal and “…” a quoted identifier; nothing escapes + // inside them. + {"with curly-quoted literal holding a paren then insert", "WITH x AS (SELECT \u2018(\u2019 AS s) INSERT INTO t SELECT s FROM x", true}, + {"with curly-quoted literal holding a verb then select", "WITH x AS (SELECT \u2018) INSERT\u2019 AS s) SELECT s FROM x", false}, + {"with curly-quoted identifier holding a paren then insert", "WITH x AS (SELECT 'q' AS \u201cc(d\u201d) INSERT INTO t SELECT * FROM x", true}, + // ClickHouse's lexer skips \v, \f and Unicode spaces as whitespace // (TestIsMutation_ClickHouseWhitespace covers the whole set). {"leading form feed then insert", "\fINSERT INTO t VALUES (1)", true}, From c427697ed6a716cccadd678172c64d0d1ad33e3b Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:45:45 -0400 Subject: [PATCH 47/79] fix(cache): replace connections so a failover reaches the new primary After a failover behind a stable address (a DNS name, a proxy), the connections a process holds stay on the demoted node, which answers but refuses writes, and rueidis does not redial on READONLY. The refusal opened the breaker, and the write probe kept it open for the life of the process, with its owed bumps never delivered. Each connection is now replaced after a minute (rueidis retries what was in flight), which re-resolves the address, so the process reaches the new primary and delivers what it owes within about that long. A test fails over from a primary to its replica behind a forwarder that moves new connections only. Review fixes: the CHANGELOG no longer says the cache.backend switch arrives with E4 (it exists and accepts only local), AGENTS.md scopes "every key leads with the tenant" to LocalCache, Invalidate's godoc says batches, and the refused-writes test holds the refusal until an unchecked drain would sit at its 10 s cap, bounding recovery at the probe period plus 3 s. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- AGENTS.md | 2 +- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- internal/cache/export_test.go | 10 ++- internal/cache/redis.go | 13 +++- internal/cache/redis_integration_test.go | 99 +++++++++++++++++++++++- internal/cache/redis_test.go | 1 + 7 files changed, 121 insertions(+), 8 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 9f7d9bcea..c4ee04283 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -31,7 +31,7 @@ Twenty internal packages under `internal/` (plus `internal/testutil/` for shared - **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers; `ch_errors.go` (`writeCHError`) is the one mapping from a failed ClickHouse query to status, `code` and `retryable` - **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` handing the sweeper each tenant's own gap window (a rejected tenant's as its folder last had it, unbounded for one rejected since boot) and the `mq.max_bytes_gb` reconcile each served tenant's byte budget, and `defaultPolicy` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, and the same hook's `Hub.Prune` ends the open streams of a tenant no longer served), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `New` wires only what the process's `roles` need (discovery, dedupe, auth verifiers, the hub bridge and keepalive per API process; the ingest worker per ingest process; the sweeper under its lease through `elected`); a process without `api` serves `api.NewOpsRouter` — probes, `/version`, metrics, and the settings reload behind the operator key alone. `cmd/wavehouse` and `tests/integration` both boot through it - **`auth/`** — JWT auth middleware: HMAC **or** JWKS verification with `alg` pinned to the active verifier, role extraction from a configurable claim path; always runs, never rejects (bad token → empty role + stashed reason). One verifier per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 9): `Authenticator` keys them by `tenant.ID` — the request store's `settings.Store.Tenant()`, through an injected `TenantSource`; `tenant.Default` on the tenant-exempt routes — built from each tenant's `auth` block by `Reconfigure`, dropped by `Prune` once the tenant stops being served (rejected or removed), released by `Close`; the secrets (`Config`) are boot-level and shared. A JWKS key set is fetched off the boot and reload paths: until one has been stored the verifier is pending and a token-bearing request gets `503` + `Retry-After` from `api.refuseUnverifiable` (`auth.ErrVerifierPending`), never a `default_role` evaluation; refresh is library-managed (Eric, 2026-09-22), response capped at 1 MiB; the operator key's admin role is the request tenant's -- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index), and `RedisCache`, the Redis-compatible shared backend (random version tokens under the tenant's hash tag, one-round-trip lookups, bypass on failure behind a circuit breaker, deferred invalidations retried; built and tested, not yet selectable by config — [#613](https://github.com/Wave-RF/WaveHouse/issues/613) E4). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for the caller's query key and its singleflight (escaped whole as the lead field of the stored key, `|.||…`), `..
.
.` for a namespace, each field escaped by `keyenc` where the key is built (a `Namespace` carries the raw table and scope, so no caller escapes) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version folded into every namespace and query key of one tenant, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read, taken before the handler chooses any input a bump invalidates — the tenant's connection included — and `Set` files the fill under it, so a write landing mid-query, or a reload moving the tenant to another address or database after the request took its connection, orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) +- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index), and `RedisCache`, the Redis-compatible shared backend (random version tokens under the tenant's hash tag, one-round-trip lookups, bypass on failure behind a circuit breaker, deferred invalidations retried; built and tested, not yet selectable by config — [#613](https://github.com/Wave-RF/WaveHouse/issues/613) E4). Every key carries the tenant (in `RedisCache`, right after the key prefix); in `LocalCache` and the version index it leads ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for the caller's query key and its singleflight (escaped whole as the lead field of the stored key, `|.||…`), `..
.
.` for a namespace, each field escaped by `keyenc` where the key is built (a `Namespace` carries the raw table and scope, so no caller escapes) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version folded into every namespace and query key of one tenant, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read, taken before the handler chooses any input a bump invalidates — the tenant's connection included — and `Set` files the fill under it, so a write landing mid-query, or a reload moving the tenant to another address or database after the request took its connection, orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config. `Classify` (`errclass.go`) says what a failed ClickHouse request means for the request — `Unavailable`, `Denied`, `Rejected` (any unlisted exception code: the server read it and refused it), or `Unknown` (no code, no recognizable transport failure) — over the driver's error types and the HTTP interface's `HTTPError`; the ingest worker and the query handlers (`api/ch_errors.go` `writeCHError`, [#403](https://github.com/Wave-RF/WaveHouse/issues/403), [#271](https://github.com/Wave-RF/WaveHouse/issues/271)) both use it - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) - **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability when a selected backend keeps state there (`NeedsDataDir`); `backends.go` holds each layer's `.backend` (only the in-process value today); `config.go` holds `roles` (`Has(Role)`) and `instance_id`, and `Validate` refuses a role split the backends cannot serve (any split over the embedded MQ; `api` without `ingest`, or the reverse, over a local cache) — boot is the validator, there is no dry run diff --git a/CHANGELOG.md b/CHANGELOG.md index b3553aebc..4cc140370 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A Redis-compatible shared cache backend, built but not yet selectable** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), open a circuit breaker that skips the server until a probe write succeeds, and deferred invalidations are retried until they land — the first at once, and at the probe's cadence while the breaker is open — collapsing to one tenant-wide bump per tenant past 100,000 keys; until one lands, the process that owes it bypasses the lookups it would orphan. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. Nothing selects it yet: the `cache.backend` switch and the `cache.redis.*` settings arrive with the wiring (E4), so every deployment still runs `LocalCache`. Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression, server-stops-answering, refused-writes (a demoted primary, a full `noeviction` server) and owed-invalidation cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. +- **A Redis-compatible shared cache backend, built but not yet selectable** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/cache.go`, `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`, `CONTRIBUTING.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), open a circuit breaker that skips the server until a probe write succeeds (connections are replaced every minute, so one that outlived a failover behind a stable address reaches the new primary), and deferred invalidations are retried until they land — the first at once, and at the probe's cadence while the breaker is open — collapsing to one tenant-wide bump per tenant past 100,000 keys; until one lands, the process that owes it bypasses the lookups it would orphan. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. Nothing selects it yet: `cache.backend` accepts only `local` until the wiring (E4) adds `redis` and the `cache.redis.*` settings, so every deployment still runs `LocalCache`. Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression, server-stops-answering, refused-writes (a demoted primary, a full `noeviction` server), failover-behind-a-stable-address and owed-invalidation cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. - **Leases for work that must run in one process at a time, and the sweeper runs under one** (`internal/coord/` (new: `coord.go`, `local.go`, `elect.go`, `coordtest/`, + tests), `internal/app/{app,wire}.go` (+ tests), `internal/config/backends.go`, `internal/ingest/sweeper.go`, `.github/labeler.yml`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development,ingest-pipeline}.md`, `docs/src/content/docs/configuration.mdx`, `config.yaml`, `AGENTS.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `coord.Coordinator` hands out named leases (`TryAcquire` → a `Term` with a strictly increasing fencing `Token`, a `Done` channel and `Resign`; `ErrHeld` while another holder's term is live), and `coord.RunElected` runs a loop only while its process holds the lease, resigning when the loop returns and campaigning again every 2s. `coord.Local` is the in-process implementation, and `coordtest.Conformance` is the suite every implementation runs — the NATS KV backend that lets several replicas share one queue comes next. The sweeper now runs through `RunElected` under the `sweeper` lease; with the in-process coordinator the one process always holds it, so nothing changes for a single-process deployment beyond one `coord: elected` log line at startup. `coord.backend` now selects the coordinator (`local`, the only value), so a `config.Config` built without `config.Load` must name it as well as the other three layers' backends. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 1d6a90356..ac449529b 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -117,7 +117,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: a namespace key is `..
.
.`, its fields joined with `keyenc.Join` so a dot in a table or scope name is escaped rather than read as a separator, and an entry's key, `|.||…` with the caller's query key escaped whole (its `:` become `%3A`), folds in the tenant's version and each dependency's namespace key and namespace version, so bumping a table (a scopeless write) or one scope — scope is reserved and empty today, so every write is the whole-table bump — orphans every dependent entry without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results, and the flat directory's single tenant simply carries the `0` prefix. `BumpTenant` (behind `InvalidateTenant`) advances the tenant version that leads every namespace key of one tenant, orphaning its every namespace and every cached query in one step (the tenant version is folded into every query key, so a pipe result with no dependencies is orphaned too) — a table no bump ever keyed included, which an enumeration of the index would miss — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. - **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET` and `MGET`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot; the table and scope are escaped with `keyenc` after the fixed prefix, so a `:` in a name is never read as the separator. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, the hash over their escaped forms, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token (a string of another length, a list, a hash) is replaced, so a token lost to eviction, expiry or a restart is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. While this process owes a bump on any of a lookup's tokens, that lookup is a bypass that files nothing: the bump would orphan whatever it found (other processes, which cannot know, serve those entries until it lands). `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing, each attempt bounded by `DialTimeout` (1 s), and a cluster client's topology read after the handshake by the larger of it and `Timeout`, which is also how long a connection waits on a silent server before it is redialed. Nothing selects this backend yet — the boot config and wiring arrive with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s E4. - **redis_codec.go** — the key schema (`tokenKeys`, `bumpKeys`, `valueKey`, which take a `Namespace`'s raw names and escape them) and the value frame: format, flags, expiry, the token list, the payload, zstd-compressed from `CompressMinBytes` when that is smaller. A stored value is capped at `MaxValueBytes` and a decoded one at eight times that, which refuses a zip bomb planted in a shared server; an unknown format, as a newer process writes during a rolling upgrade, is a miss. -- **breaker.go** — the circuit breaker: `BreakerThreshold` (5) consecutive failures open it, and while open every operation skips the server at once; after `BreakerOpenFor` (5 s) one background write (`SET :probe`) decides whether it closes, and only that probe's success closes it. A transport failure or a timeout counts against the server; an error reply saying it takes no writes right now — `READONLY` (a demoted primary), `MASTERDOWN`, `OOM` (full under `maxmemory-policy noeviction`, Redis's default, so give the cache an evicting policy), `NOREPLICAS`, `MISCONF`, `LOADING`, `BUSY`, `CLUSTERDOWN` — opens it at once, whatever the threshold, since no bump can land. Any other reply, an error reply about one key (`WRONGTYPE`, `NOPERM`, `TRYAGAIN`) or one the backend cannot use included, counts as a success, and a caller that gave up first counts as nothing. +- **breaker.go** — the circuit breaker: `BreakerThreshold` (5) consecutive failures open it, and while open every operation skips the server at once; after `BreakerOpenFor` (5 s) one background write (`SET :probe`) decides whether it closes, and only that probe's success closes it. Every connection is replaced after a minute (rueidis retries what was in flight), because one that outlives a failover behind a stable address stays on the demoted node, which answers but refuses writes; the replacement re-resolves the address, so a bypassed process reaches the new primary, and delivers the bumps it owes, within about that long. A transport failure or a timeout counts against the server; an error reply saying it takes no writes right now — `READONLY` (a demoted primary), `MASTERDOWN`, `OOM` (full under `maxmemory-policy noeviction`, Redis's default, so give the cache an evicting policy), `NOREPLICAS`, `MISCONF`, `LOADING`, `BUSY`, `CLUSTERDOWN` — opens it at once, whatever the threshold, since no bump can land. Any other reply, an error reply about one key (`WRONGTYPE`, `NOPERM`, `TRYAGAIN`) or one the backend cannot use included, counts as a success, and a caller that gave up first counts as nothing. - **pending.go** — invalidations the server did not take, kept per token key (repeats coalesce) and retried by a background loop until they land: the first one owed at once, then with backoff from 100 ms to 10 s, and while the breaker is open at each probe, which the loop starts when due, so a process that makes no lookups (ingest only) recovers as soon as one that does. `Invalidate` sends its bumps 1,000 to a round trip, as the retry does. Landing late is still correct: a fresh token orphans the pre-write entries and any fill made meanwhile. Past `PendingMax` (100,000) keys the set collapses to one tenant bump per affected tenant. `Close` makes one last attempt at the bumps still owed, past the breaker (an open one is why they are owed); what that attempt cannot deliver is lost, and the entries they would orphan are served until their TTL — the same failure as a worker stopping between an insert and its invalidation. - **metrics.go** — the shared backend's instruments (meter `wavehouse-cache`, `backend="redis"`, no tenant attribute): `wavehouse_cache_lookups_total` by `result` (`hit`, `miss`, `stale` — filed under since-bumped tokens, `bypass` — server skipped, or held by a bump this process owes, `error`), `wavehouse_cache_op_duration_seconds` by `op` (`lookup`, `set`, `invalidate`), `wavehouse_cache_breaker_open` (1 while bypassed, including before the first connection), `wavehouse_cache_invalidations_total` by `result` (`ok`, `deferred`), `wavehouse_cache_invalidations_pending` (bumps owed: entries they would orphan may be served stale meanwhile), `wavehouse_cache_value_bytes` (stored size), `wavehouse_cache_oversize_total` and `wavehouse_cache_set_failures_total` by `reason` (`oom`, `timeout`, `other`). diff --git a/internal/cache/export_test.go b/internal/cache/export_test.go index 0ef1f0b61..70b8544d2 100644 --- a/internal/cache/export_test.go +++ b/internal/cache/export_test.go @@ -1,6 +1,10 @@ package cache -import "github.com/redis/rueidis" +import ( + "time" + + "github.com/redis/rueidis" +) // Hooks for the integration tests in package cache_test. @@ -16,6 +20,10 @@ func Pending(r *RedisCache) int { return r.pending.len() } // Bypassed reports whether r is skipping the server. func Bypassed(r *RedisCache) bool { return r.bypassed() } +// SetConnLifetime shortens how long c's connections live before they are +// replaced, for a failover test. +func SetConnLifetime(c *RedisConfig, d time.Duration) { c.connLifetime = d } + // KeyPrefix is the prefix every key r writes leads with. func KeyPrefix(r *RedisCache) string { return r.cfg.KeyPrefix } diff --git a/internal/cache/redis.go b/internal/cache/redis.go index b8bd5e286..9d6733d4e 100644 --- a/internal/cache/redis.go +++ b/internal/cache/redis.go @@ -38,6 +38,7 @@ const ( DefaultRedisPendingMax = 100_000 defaultBreakerThreshold = 5 defaultBreakerOpenFor = 5 * time.Second + defaultConnLifetime = time.Minute ) const ( @@ -77,6 +78,8 @@ type RedisConfig struct { BreakerThreshold int // consecutive failures that open the breaker BreakerOpenFor time.Duration // how long it stays open before a probe + + connLifetime time.Duration // defaultConnLifetime; tests shorten it } func (c RedisConfig) withDefaults() (RedisConfig, error) { @@ -135,6 +138,7 @@ func (c RedisConfig) withDefaults() (RedisConfig, error) { c.PendingMax = cmpOr(c.PendingMax, DefaultRedisPendingMax) c.BreakerThreshold = cmpOr(c.BreakerThreshold, defaultBreakerThreshold) c.BreakerOpenFor = cmpOr(c.BreakerOpenFor, defaultBreakerOpenFor) + c.connLifetime = cmpOr(c.connLifetime, defaultConnLifetime) return c, nil } @@ -163,6 +167,13 @@ func (c RedisConfig) clientOption() rueidis.ClientOption { // Close would wait out; never under Timeout, so no connection is cut // while an operation may still wait on it. opt.ConnWriteTimeout = max(c.DialTimeout, c.Timeout) + // A connection outlives a failover behind a stable address: the demoted + // node still answers, refusing writes, and rueidis does not redial on + // READONLY. Replacing each connection this often (rueidis retries what + // was in flight) re-resolves the address, so a process the refusals + // bypass reaches the new primary, and delivers its owed bumps, within + // about this long. + opt.ConnLifetime = c.connLifetime if c.Mode == RedisSentinel { opt.Sentinel = rueidis.SentinelOption{MasterSet: c.SentinelMaster, TLSConfig: c.TLS, Dialer: opt.Dialer} } @@ -558,7 +569,7 @@ func setFailureReason(err error) string { } // Invalidate sets a fresh token for every token the namespaces' writes -// reach, in one pipelined round trip. Bumps the server does not take are +// reach, in pipelined batches of drainBatch, one round trip each. Bumps the server does not take are // kept and retried until it does, and reported as an error meanwhile. func (r *RedisCache) Invalidate(ctx context.Context, namespaces []Namespace) (uint64, error) { owner := map[string]tenant.ID{} diff --git a/internal/cache/redis_integration_test.go b/internal/cache/redis_integration_test.go index 68c3e71a2..3169e9230 100644 --- a/internal/cache/redis_integration_test.go +++ b/internal/cache/redis_integration_test.go @@ -7,6 +7,7 @@ import ( "context" "crypto/tls" "fmt" + "io" "net" "slices" "strconv" @@ -559,6 +560,98 @@ func (c unanswered) Write(b []byte) (int, error) { return c.Conn.Write(b) } +// forward proxies each connection it accepts to the address target holds +// at that moment, so a switch moves new connections only, as a stable DNS +// name or a proxy does after a failover. It returns the address to dial. +func forward(t *testing.T, target *atomic.Pointer[string]) string { + t.Helper() + var lc net.ListenConfig + ln, err := lc.Listen(context.Background(), "tcp", "127.0.0.1:0") + require.NoError(t, err) + t.Cleanup(func() { _ = ln.Close() }) + go func() { + for { + c, err := ln.Accept() + if err != nil { + return + } + go func() { + defer func() { _ = c.Close() }() + var d net.Dialer + u, err := d.DialContext(context.Background(), "tcp", *target.Load()) + if err != nil { + return + } + defer func() { _ = u.Close() }() + go func() { _, _ = io.Copy(u, c); _ = u.Close() }() + _, _ = io.Copy(c, u) + }() + } + }() + return ln.Addr().String() +} + +// A failover behind a stable address: the connections the process holds +// stay on the demoted node, which answers but refuses writes, while new ones +// reach the node promoted in its place. The refusals bypass the cache, and +// replacing connections carries it to the new primary, where the owed bump +// lands before anything it would orphan is served. +func TestRedis_FailoverBehindAStableAddress(t *testing.T) { + t.Parallel() + ctx := context.Background() + start := func() *server { // no delay before a full sync + return startStandalone(t, redisImage, "redis-server", "--save", "", "--appendonly", "no", "--repl-diskless-sync-delay", "0") + } + primary, replica := start(), start() + rPrimary, rReplica := raw(t, primary), raw(t, replica) + ip := func(s *server) string { + t.Helper() + ip, err := s.ctr.ContainerIP(ctx) + require.NoError(t, err) + return ip + } + command(t, rReplica, "REPLICAOF", ip(primary), "6379") + require.Eventually(t, func() bool { + info, err := rReplica.Do(ctx, rReplica.B().Info().Section("replication").Build()).ToString() + return err == nil && strings.Contains(info, "master_link_status:up") + }, 30*time.Second, 50*time.Millisecond, "the replica syncs") + + var target atomic.Pointer[string] + target.Store(&primary.addr) + stable := &server{addr: forward(t, &target), mode: cache.RedisStandalone} + prefix := uniquePrefix() + a := open(t, stable, prefix, func(c *cache.RedisConfig) { + c.BreakerThreshold, c.BreakerOpenFor = 1000, 200*time.Millisecond + cache.SetConnLifetime(c, time.Second) + }) + deps := []cache.Namespace{{Tenant: "acme", Table: "events"}} + _, snap, err := a.Lookup(ctx, "acme", "q", deps) + require.NoError(t, err) + require.NoError(t, a.Set(ctx, snap, []byte("pre-write rows"), time.Minute)) + acked, err := rPrimary.Do(ctx, rPrimary.B().Wait().Numreplicas(1).Timeout(5000).Build()).AsInt64() + require.NoError(t, err) + require.Equal(t, int64(1), acked, "the replica holds the fill") + + command(t, rReplica, "REPLICAOF", "NO", "ONE") + command(t, rPrimary, "REPLICAOF", ip(replica), "6379") + _, err = a.Invalidate(ctx, deps) + require.ErrorContains(t, err, "READONLY") + require.True(t, cache.Bypassed(a)) + target.Store(&replica.addr) + + deadline := time.Now().Add(15 * time.Second) + for cache.Pending(a) > 0 { + require.True(t, time.Now().Before(deadline), "the owed bump reaches the new primary") + e, _, err := a.Lookup(ctx, "acme", "q", deps) + require.NoError(t, err) + require.NotEqual(t, "pre-write rows", string(e.Value), "served before the owed bump landed") + time.Sleep(10 * time.Millisecond) + } + e, _, err := open(t, stable, prefix).Lookup(ctx, "acme", "q", deps) + require.NoError(t, err) + assert.Nil(t, e.Value, "the bump landed where every process now reads") +} + // command runs cmd on s through c, for the test to reconfigure the server. func command(t *testing.T, c rueidis.Client, cmd ...string) { t.Helper() @@ -693,8 +786,8 @@ func TestRedis_RefusedWrites(t *testing.T) { assert.True(t, cache.Bypassed(reader), "a refused fill opens the breaker too") // Long enough for many probes to be refused, and for a drain - // backing off unchecked to be seconds from its next attempt. - time.Sleep(3500 * time.Millisecond) + // backing off unchecked to be at its 10 s cap. + time.Sleep(7 * time.Second) assert.True(t, cache.Bypassed(reader), "owing nothing, it is held open by probes that write and are refused") assert.True(t, cache.Bypassed(a)) assert.Empty(t, lookup(a, events)) @@ -711,7 +804,7 @@ func TestRedis_RefusedWrites(t *testing.T) { time.Sleep(time.Millisecond) } require.Eventually(t, func() bool { return cache.Pending(ingest) == 0 }, 10*time.Second, 5*time.Millisecond) - assert.Less(t, time.Since(restored), openFor+time.Second, "an ingest-only process probes when due, not at the drain's backoff") + assert.Less(t, time.Since(restored), openFor+3*time.Second, "an ingest-only process probes when due, not at the drain's backoff") require.Eventually(t, func() bool { return !cache.Bypassed(reader) }, 10*time.Second, 5*time.Millisecond) fresh := open(t, s, prefix) diff --git a/internal/cache/redis_test.go b/internal/cache/redis_test.go index 17da785a6..a9ccc3711 100644 --- a/internal/cache/redis_test.go +++ b/internal/cache/redis_test.go @@ -61,6 +61,7 @@ func TestRedisConfig_Defaults(t *testing.T) { assert.Zero(t, c.CompressMinBytes, "0 means never compress, not the default") assert.True(t, c.clientOption().ForceSingleClient) assert.Equal(t, DefaultRedisDialTimeout, c.clientOption().ConnWriteTimeout) + assert.Equal(t, defaultConnLifetime, c.clientOption().ConnLifetime) slow := c slow.Timeout = 5 * time.Second assert.Equal(t, slow.Timeout, slow.clientOption().ConnWriteTimeout, "never under the op timeout") From 4df73f995fd14139088e310b2175faf27eebc69d Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:46:54 -0400 Subject: [PATCH 48/79] docs(cache): clarify which bumps a no-index tenant absorbs MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The version_manager.go bullet read "a bump of a tenant with no index is a no-op" right after introducing BumpTenant, so it looked scoped to that one call. It actually covers any bump — table, scope or tenant — against a tenant with no index, which is what lets Prune and a departed tenant's index stay gone (neither the sharedTables fan-out nor an insert still in flight for a just-pruned tenant can revive it). Found by pre-push review of #621 (origin/feat/cache-snapshot...HEAD). Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- docs/src/content/docs/architecture.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 4690761df..7aede5bad 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -114,7 +114,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on — one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)) — each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope by their raw names, which the cache escapes where it builds a key. `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The snapshot is taken before any input a bump invalidates is chosen, the tenant's connection included: a reload that moves the tenant to another address or database runs `Pools.Reconcile` and then `InvalidateTenant` (a repoint that keeps both, such as a username or `tls` change, reads the same tables and bumps nothing), so a request that took the old pool files the old database's rows under a version that bump orphans, whether its `Set` lands before the bump or after. The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. -- **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: one version per tenant, per (tenant, table) and per (tenant, table, scope), each keyed by its name alone and bumped in place, so the index holds one entry per live tenant, table and scope however often each is bumped ([#262](https://github.com/Wave-RF/WaveHouse/issues/262)). A query key folds the tenant's version and, for each dependency, its tenant's, table's and scope's, so bumping a table (a scopeless write) orphans every scope of it, and bumping one scope orphans that scope and the whole-table view — scope is reserved and empty today, so every write is the whole-table bump — all without touching the pool. Every field — the caller's query key, the tenant id, and each dependency's table and scope — is escaped and joined by `internal/keyenc` where the key is built, so a dot, a space or a `%` in a name is never read as a separator: each dependency renders as `..
.
..`, and the whole entry key is `|.||…`. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results. A tenant's version is a *generation*, unique within the process and handed out by the first key built for the tenant; `BumpTenant` (behind `InvalidateTenant`) drops the tenant's whole index, so the next key gets a fresh generation no cached entry folds, orphaning every cached result of the tenant in one step — a pipe result with no dependencies, and a table no bump ever keyed, included — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). `LocalCache.Prune` does the same for every tenant no longer served, which `internal/app` runs after each settings reload, so a tenant removed or rejected stops holding its index. A table bump drops the table's scope versions with it, since every key they were folded into also folds the old table version; and a bump of a tenant with no index is a no-op, since no key folds its next generation yet. The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. +- **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: one version per tenant, per (tenant, table) and per (tenant, table, scope), each keyed by its name alone and bumped in place, so the index holds one entry per live tenant, table and scope however often each is bumped ([#262](https://github.com/Wave-RF/WaveHouse/issues/262)). A query key folds the tenant's version and, for each dependency, its tenant's, table's and scope's, so bumping a table (a scopeless write) orphans every scope of it, and bumping one scope orphans that scope and the whole-table view — scope is reserved and empty today, so every write is the whole-table bump — all without touching the pool. Every field — the caller's query key, the tenant id, and each dependency's table and scope — is escaped and joined by `internal/keyenc` where the key is built, so a dot, a space or a `%` in a name is never read as a separator: each dependency renders as `..
.
..`, and the whole entry key is `|.||…`. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results. A tenant's version is a *generation*, unique within the process and handed out by the first key built for the tenant; `BumpTenant` (behind `InvalidateTenant`) drops the tenant's whole index, so the next key gets a fresh generation no cached entry folds, orphaning every cached result of the tenant in one step — a pipe result with no dependencies, and a table no bump ever keyed, included — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). `LocalCache.Prune` does the same for every tenant no longer served, which `internal/app` runs after each settings reload, so a tenant removed or rejected stops holding its index. A table bump drops the table's scope versions with it, since every key they were folded into also folds the old table version; and any bump (of a table, a scope or the tenant) under a tenant with no index is a no-op that records nothing, since the next key built for it gets a fresh generation no cached entry folds — so neither the `sharedTables` fan-out nor an insert still in flight for a tenant just pruned brings its index back. The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. ### `config/` — Configuration From 95575f27a49982cddd1c1febd72028ff669a901b Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:54:37 -0400 Subject: [PATCH 49/79] docs(cache): CHANGELOG says #262 is part-fixed, not fixed The #262 entry said "fixes #262 for the in-process cache", which would close the issue on merge (delete_branch_on_merge + squash_merge_commit_message: PR_BODY carry the PR body's own "Fixes #262" into main's history the same way). #262 has a residual left open on purpose: per-table scope cardinality isn't capped, since scope stays empty until #235 populates it. Reworded to "part of #262" and named the residual, matching the PR body and the issue comment that records the trigger. Found by pre-push review of #621 (origin/feat/cache-snapshot...HEAD). Co-Authored-By: Claude Sonnet 5 --- CHANGELOG.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index a89a0d2be..6631b3c4a 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -89,7 +89,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **An unavailable ClickHouse is retried with backoff instead of dead-lettering every row** (`internal/chconn/errclass.go` (new, + tests), `internal/ingest/{worker,backoff}.go` (`backoff.go` new, + tests), `internal/mq/{mq,embedded}.go`, `internal/testutil/mocks.go`, `tests/integration/ingest_outage_test.go` (new), `AGENTS.md`, `README.md`, `docs/src/content/docs/{ingest-pipeline,architecture,api,deployment,why-wavehouse}.md`, `docs/src/content/docs/{settings-directory,index,access-control}.mdx`): workstream A of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). A failed batch insert used to go through row-by-row isolation whatever the failure, so a ClickHouse that was down, overloaded or read-only failed every row twice and parked the whole batch on the DLQ. `chconn.Classify` now classes the failure first — `Rejected` (any ClickHouse exception code outside the availability and credential lists: the server read the row and refused it), `Unavailable` (connection refused/reset, timeouts, `TOO_MANY_SIMULTANEOUS_QUERIES`, `SERVER_OVERLOADED`, `MEMORY_LIMIT_EXCEEDED`, `TOO_MANY_PARTS`, `READONLY`, `TABLE_IS_READ_ONLY`, `KEEPER_EXCEPTION`, …), `Denied` (`AUTHENTICATION_FAILED`, `ACCESS_DENIED`, …) or `Unknown` (no code, no recognizable transport failure). Only `Rejected` is isolated and dead-lettered as before, and a multi-row batch refused with `TOO_MANY_PARTS` or `MEMORY_LIMIT_EXCEEDED` is split row by row first (`chconn.Splittable`), because a batch spanning too many partitions or too much memory can fail where each of its rows inserts; every other class hands the batch back to the queue with a delayed nak (`mq.Message.NakWithDelay`, new) under a jittered 1 s → 30 s backoff shared by every table on the same ClickHouse pool (a failure of one table — read-only, too many parts or mutations, a grant missing on it, `chconn.TableScoped` — backs off that table alone), which turns rows away without a request while it runs and probes once per window, and ClickHouse going away mid-isolation stops isolation and retries the rows it had not settled. Counted by the new `wavehouse_ingest_retries_total{table, reason}`; logged at `WARN` when an outage starts and at most every 30 s during it. A long outage now shows as a growing ingest stream and, at `mq.max_bytes_gb`, ingest `503`s — not as a full DLQ; a lasting failure of one table holds back its tenant's other tables once its waiting rows reach `maxAckPending`. Retried rows come back out of arrival order, which matters only to a `ReplacingMergeTree` without a version column or a `CollapsingMergeTree`. - **Schema discovery's retry loop jitters its backoff** (`internal/discovery/discovery.go` (+ tests), `internal/app/wire.go`, `internal/api/errors.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`): `RetryRefresh` slept exactly `2s * 2^n` capped at 60s, so instances retrying against one recovering ClickHouse fired in lockstep, every 60s on the same second. Each sleep is now drawn uniformly from below the backoff (full jitter), spreading the retries over the whole window and halving the mean wait — so a failing tenant's retries, their log lines and `wavehouse_schema_refresh_failures_total` come about twice as often ([#141](https://github.com/Wave-RF/WaveHouse/issues/141)). - **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/query/ident.go` (removed, + tests), `internal/app/wire.go`, `docs/src/content/docs/{api,architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL (a pipe result's key folded no version, so pipes were unaffected; with the tenant's version in every key they now take the same snapshot). The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. The snapshot is taken before any input a bump invalidates is chosen, the tenant's ClickHouse connection included: both handlers now look up before they resolve the tenant's pool, so a reload that moves the tenant to another address or database after a request took the old pool orphans that request's fill instead of filing the old database's rows as fresh under the new tenant version. A tenant on no pool is still a `503` before a cached result is served or a query runs. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. The cache now escapes table and scope names itself: a `Namespace` carries them raw and the version index builds its keys with `internal/keyenc`, so neither the structured-query read nor the ingest worker's invalidation escapes them (`query.SafeEncodeToken` is gone) and no name reaches a key unescaped. The suite pins that a name holding a dot, a space or a `%` is read and bumped under one key, and that names which would run together unescaped (`a.0.b` against `a` with scope `b.0.`) stay two entries. The version index's own entries are unaffected by the escaping; only the rendered key changes: it now carries the tenant's version and escapes the caller's query key whole. All of them live in the process, so nothing stored is orphaned. -- **The cache's version index no longer grows with every bump, and forgets a tenant no longer served** (`internal/cache/{local,version_manager}.go` (+ tests), `internal/app/wire.go` (+ tests), `docs/src/content/docs/architecture.md`, `AGENTS.md`): fixes [#262](https://github.com/Wave-RF/WaveHouse/issues/262) for the in-process cache, part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The index nested each table under its tenant's version and each scope under its table's, and never pruned, so every tenant invalidation left the tenant's whole index behind, and it grew with every tenant ever served. It now holds one version per tenant, per (tenant, table) and per (tenant, table, scope), bumped in place. A tenant invalidation drops the tenant's index and hands its next key a generation unique within the process, so nothing cached before it can match again, and a table bump drops the table's scope versions. After each settings reload the index of every tenant no longer served, removed or rejected, is dropped the same way; its cached results are orphaned with it, as they already were when such a tenant came back on a pool. No change to what is cached or served. The Redis-compatible backend (#613) will bound its versions with a TTL instead. +- **The cache's version index no longer grows with every bump, and forgets a tenant no longer served** (`internal/cache/{local,version_manager}.go` (+ tests), `internal/app/wire.go` (+ tests), `docs/src/content/docs/architecture.md`, `AGENTS.md`): part of [#262](https://github.com/Wave-RF/WaveHouse/issues/262) (growth across bumps and a departed tenant's memory; per-table scope cardinality is left open, see the issue) and of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The index nested each table under its tenant's version and each scope under its table's, and never pruned, so every tenant invalidation left the tenant's whole index behind, and it grew with every tenant ever served. It now holds one version per tenant, per (tenant, table) and per (tenant, table, scope), bumped in place. A tenant invalidation drops the tenant's index and hands its next key a generation unique within the process, so nothing cached before it can match again, and a table bump drops the table's scope versions. After each settings reload the index of every tenant no longer served, removed or rejected, is dropped the same way; its cached results are orphaned with it, as they already were when such a tenant came back on a pool. No change to what is cached or served. The Redis-compatible backend (#613) will bound its versions with a TTL instead. - **An explicit `false`, `0` or `""` in `config.yaml` is no longer replaced by the key's default** (`internal/config/config.go`, `internal/config/defaults_test.go` (new), `docs/src/content/docs/configuration.mdx`, `AGENTS.md`): [#631](https://github.com/Wave-RF/WaveHouse/issues/631). Defaults lived in cleanenv `env-default` tags, which cleanenv applies after the YAML decode to any field still at its zero value, so it could not tell a key the file set to its zero value from one the file left out. `otel.traces.enabled: false`, `otel.metrics.enabled: false` and `otel.logs.enabled: false` came back `true`; `otel.traces.sample_rate: 0` and `otel.logs.sample_rate: 0` came back `1.0`; `server.shutdown_timeout: 0` came back `10`; `cache.l1_max_cost: 0`, `prometheus.path: ""` and `data_dir: ""` came back as their defaults; `server.port: 0` came back `8080`. All of it was silent. Defaults now live in one Go function, `defaults()`, which `Load` starts from before decoding the file and then applying `WH_*` variables, so the order is env > YAML > default and a key the file sets always wins. **Behaviour change if your file relied on the bug:** a zero you wrote now takes effect. A file that says `sample_rate: 0` now exports no traces (or no DEBUG/INFO logs), where it silently exported everything; a signal set `enabled: false` is now off; `shutdown_timeout: 0` now skips the drain. `cache.l1_max_cost: 0`, `server.port: 0`, and `data_dir: ""` now refuse boot (`cache init: MaxCost can't be zero`, `server.port 0 out of range`, `data_dir (WH_DATA_DIR) is required`) instead of running on the default; an empty `prometheus.path` refuses boot when `prometheus.enabled` is true. Delete the key to get the default back. Env vars are unchanged: they already honoured an explicit zero. New tests load through `config.Load` for every affected key (a YAML zero is kept, an absent key gets the default, env wins in both directions), refuse an `env-default` tag on any field, and pin each documented default in `configuration.mdx` to `defaults()`. - **An embedded queue store that cannot be created fails boot at once, naming the cause** (`internal/mq/embedded.go` (+ tests)): part of [#617](https://github.com/Wave-RF/WaveHouse/issues/617). A regular file at `/nats`, or a `nats` directory that could not be created there, failed JetStream in the background, so boot waited out the server's 5s readiness check and reported only `nats server not ready`. `NewEmbedded` now creates the directory first (at `0700`, as the server does) and refuses boot with the mkdir error. An existing but unwritable `nats` directory still takes the old path. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its cached results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. From 9b8ed1a6db4726eecd23528cfb5209dba9506208 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:57:41 -0400 Subject: [PATCH 50/79] docs(pipes): a proxy's 5xx is retried too, so a write can run twice A 502/503/504 from a proxy in front of WaveHouse carries no retryable field, so the SDK retries it like a dropped connection; name both. The SDK pipes page keeps .fetch()'s options next to its opening sentence and moves the retry caveat into its own paragraph. Co-Authored-By: Claude Opus 5.5 (1M context) --- docs/src/content/docs/pipes.mdx | 2 +- docs/src/content/docs/sdk/pipes.md | 4 +++- 2 files changed, 4 insertions(+), 2 deletions(-) diff --git a/docs/src/content/docs/pipes.mdx b/docs/src/content/docs/pipes.mdx index f06d1ebe0..9a916730e 100644 --- a/docs/src/content/docs/pipes.mdx +++ b/docs/src/content/docs/pipes.mdx @@ -182,7 +182,7 @@ The response is a JSON array of rows. Results flow through the shared in-process A pipe's SQL may be a write: a statement led by a write verb WaveHouse recognizes — `INSERT`, `UPDATE`, `DELETE`, `ALTER` (so `ALTER … DELETE`), `CREATE`, `DROP`, `TRUNCATE`, `RENAME`, `EXCHANGE`, `REPLACE`, `OPTIMIZE`, `ATTACH`, `DETACH`, `GRANT`, `REVOKE`, `KILL`, `SET`, `USE` or `SYSTEM` — directly or after a `WITH` list, as in `WITH … INSERT`. Such a pipe runs on **every** call: it never reads or fills the cache and is never coalesced with an identical call in flight, so ten identical calls are ten writes. The response is `[]` with `X-Cache: BYPASS` and `Cache-Control: no-store`, so an HTTP cache in front of a `GET` does not answer a repeat either. WaveHouse classifies the statement from its leading keyword (after any `WITH` list), with the same classifier that sends it to ClickHouse as a write, so no pipe property marks it; a statement led by any other keyword runs as a read. -A failed write is not retried automatically, because it may have run. It answers with the status and `code` a failed read would ([ClickHouse errors on the query paths](/api#clickhouse-errors-on-the-query-paths)), but always with `retryable: false` and no `Retry-After`, `503 clickhouse.unavailable` included: once the statement is on its way to ClickHouse, WaveHouse cannot tell whether it ran. The [SDK](/sdk/pipes) does not retry such an answer, so check whether the write landed before you send it again. A call refused before anything is sent — the tenant on no ClickHouse pool, `503` with `Retry-After: 30` — cannot have run, and the SDK retries it. The SDK also retries a request whose answer never arrives, such as one on a dropped connection, so a write can still run twice that way; give a client that runs write pipes [`options.maxRetries`](/sdk#clientconfigdb) `0` if that matters. +A failed write is not retried automatically, because it may have run. It answers with the status and `code` a failed read would ([ClickHouse errors on the query paths](/api#clickhouse-errors-on-the-query-paths)), but always with `retryable: false` and no `Retry-After`, `503 clickhouse.unavailable` included: once the statement is on its way to ClickHouse, WaveHouse cannot tell whether it ran. The [SDK](/sdk/pipes) does not retry such an answer, so check whether the write landed before you send it again. A call refused before anything is sent — the tenant on no ClickHouse pool, `503` with `Retry-After: 30` — cannot have run, and the SDK retries it. The SDK also retries when WaveHouse's own answer never reaches it — a dropped connection, or a `502`/`503`/`504` from a proxy in front of WaveHouse that gave up waiting — so a write can still run twice that way; give a client that runs write pipes [`options.maxRetries`](/sdk#clientconfigdb) `0` if that matters. `allowed_roles` is a write pipe's only gate: the [policy engine](/access-control)'s insert rules do not apply to it, so any role you list — including a [`default_role`](/access-control#default_role--public-unauthenticated-access) that anonymous callers resolve to — can run the write. The operator fixes the statement and its predicate when authoring the pipe; callers supply only literal values, provided every placeholder is written bare ([how a value becomes SQL](#how-a-value-becomes-sql)). diff --git a/docs/src/content/docs/sdk/pipes.md b/docs/src/content/docs/sdk/pipes.md index 685cfc34d..04bab0c98 100644 --- a/docs/src/content/docs/sdk/pipes.md +++ b/docs/src/content/docs/sdk/pipes.md @@ -17,12 +17,14 @@ const { data } = await wh.pipe('top_pages', { start_date: '2026-01-01', limit: 5 ### `.fetch(opts?)` -Execute and return results. A [pipe that writes](/pipes#pipes-that-write) returns `[]`. A ClickHouse failure on one comes back `retryable: false`, and the SDK does not retry it; it does still retry a request whose answer never arrives, such as one on a dropped connection, so a write can run twice. If that matters, give the client that runs write pipes [`options.maxRetries`](/sdk#clientconfigdb) `0`. Takes `PipeRequestOptions` — `{ signal }` only, narrower than the `.fetch(opts?)` on a [query builder](/sdk/queries), which also accepts `limit`. Passing a `limit` is a compile error rather than a silent no-op. +Execute and return results. A [pipe that writes](/pipes#pipes-that-write) returns `[]`. `.fetch()` takes `PipeRequestOptions` — `{ signal }` only, narrower than the `.fetch(opts?)` on a [query builder](/sdk/queries), which also accepts `limit`. Passing a `limit` is a compile error rather than a silent no-op. `limit` is typed `never` rather than left out, so the rejection also catches a value passed in a variable — leaving it out would only reject an inline object. That cuts both ways: a value *declared* as `RequestOptions` is rejected whether or not it actually carries a limit, since the type permits one. If you share one options object across calls, type it as `PipeRequestOptions` — the table and query-builder `.fetch()` accept that too — or inline `{ signal }` at the pipe call. There is no per-call row cap here: the endpoint binds your `params` as the pipe's parameters, so a limit has to be declared in the pipe's SQL as `{{limit}}` (see [Named Pipes](/pipes)) and passed as `wh.pipe(name, { limit })`, as in the example above. +A ClickHouse failure on a [pipe that writes](/pipes#pipes-that-write) comes back `retryable: false`, and the SDK does not retry it. The SDK does still retry when WaveHouse's own answer never reaches it — a dropped connection, or a `502`/`503`/`504` from a proxy in front of WaveHouse that gave up waiting — so a write can run twice that way. If that matters, give the client that runs write pipes [`options.maxRetries`](/sdk#clientconfigdb) `0`. + ### `.stream(opts?)` Open a live stream (see [Streaming](/sdk/streaming)). From 47528116cf4e6cb0e02c55f27a2e68af2c4bd520 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 06:07:58 -0400 Subject: [PATCH 51/79] docs(cache): cache.backend exists; Redis key shapes; failover wording architecture.md no longer says the boot config arrives with E4: cache.backend exists and accepts only local until E4 adds redis. AGENTS.md gives both Redis key shapes (a token's tenant is a hash tag, a value's follows :q:), and the CHANGELOG's failover sentence says the process, not the old connection, reaches the new primary, within about a minute. Co-Authored-By: Claude Opus 5.5 (1M context) --- AGENTS.md | 2 +- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- 3 files changed, 3 insertions(+), 3 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index c4ee04283..81882f018 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -31,7 +31,7 @@ Twenty internal packages under `internal/` (plus `internal/testutil/` for shared - **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers; `ch_errors.go` (`writeCHError`) is the one mapping from a failed ClickHouse query to status, `code` and `retryable` - **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` handing the sweeper each tenant's own gap window (a rejected tenant's as its folder last had it, unbounded for one rejected since boot) and the `mq.max_bytes_gb` reconcile each served tenant's byte budget, and `defaultPolicy` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, and the same hook's `Hub.Prune` ends the open streams of a tenant no longer served), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `New` wires only what the process's `roles` need (discovery, dedupe, auth verifiers, the hub bridge and keepalive per API process; the ingest worker per ingest process; the sweeper under its lease through `elected`); a process without `api` serves `api.NewOpsRouter` — probes, `/version`, metrics, and the settings reload behind the operator key alone. `cmd/wavehouse` and `tests/integration` both boot through it - **`auth/`** — JWT auth middleware: HMAC **or** JWKS verification with `alg` pinned to the active verifier, role extraction from a configurable claim path; always runs, never rejects (bad token → empty role + stashed reason). One verifier per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 9): `Authenticator` keys them by `tenant.ID` — the request store's `settings.Store.Tenant()`, through an injected `TenantSource`; `tenant.Default` on the tenant-exempt routes — built from each tenant's `auth` block by `Reconfigure`, dropped by `Prune` once the tenant stops being served (rejected or removed), released by `Close`; the secrets (`Config`) are boot-level and shared. A JWKS key set is fetched off the boot and reload paths: until one has been stored the verifier is pending and a token-bearing request gets `503` + `Retry-After` from `api.refuseUnverifiable` (`auth.ErrVerifierPending`), never a `default_role` evaluation; refresh is library-managed (Eric, 2026-09-22), response capped at 1 MiB; the operator key's admin role is the request tenant's -- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index), and `RedisCache`, the Redis-compatible shared backend (random version tokens under the tenant's hash tag, one-round-trip lookups, bypass on failure behind a circuit breaker, deferred invalidations retried; built and tested, not yet selectable by config — [#613](https://github.com/Wave-RF/WaveHouse/issues/613) E4). Every key carries the tenant (in `RedisCache`, right after the key prefix); in `LocalCache` and the version index it leads ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for the caller's query key and its singleflight (escaped whole as the lead field of the stored key, `|.||…`), `..
.
.` for a namespace, each field escaped by `keyenc` where the key is built (a `Namespace` carries the raw table and scope, so no caller escapes) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version folded into every namespace and query key of one tenant, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read, taken before the handler chooses any input a bump invalidates — the tenant's connection included — and `Set` files the fill under it, so a write landing mid-query, or a reload moving the tenant to another address or database after the request took its connection, orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) +- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index), and `RedisCache`, the Redis-compatible shared backend (random version tokens under the tenant's hash tag, one-round-trip lookups, bypass on failure behind a circuit breaker, deferred invalidations retried; built and tested, not yet selectable by config — [#613](https://github.com/Wave-RF/WaveHouse/issues/613) E4). Every key carries the tenant (in `RedisCache`, after the key prefix: `:{}:…` for a version token, `:q::…` for a value); in `LocalCache` and the version index it leads ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for the caller's query key and its singleflight (escaped whole as the lead field of the stored key, `|.||…`), `..
.
.` for a namespace, each field escaped by `keyenc` where the key is built (a `Namespace` carries the raw table and scope, so no caller escapes) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version folded into every namespace and query key of one tenant, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read, taken before the handler chooses any input a bump invalidates — the tenant's connection included — and `Set` files the fill under it, so a write landing mid-query, or a reload moving the tenant to another address or database after the request took its connection, orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config. `Classify` (`errclass.go`) says what a failed ClickHouse request means for the request — `Unavailable`, `Denied`, `Rejected` (any unlisted exception code: the server read it and refused it), or `Unknown` (no code, no recognizable transport failure) — over the driver's error types and the HTTP interface's `HTTPError`; the ingest worker and the query handlers (`api/ch_errors.go` `writeCHError`, [#403](https://github.com/Wave-RF/WaveHouse/issues/403), [#271](https://github.com/Wave-RF/WaveHouse/issues/271)) both use it - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) - **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability when a selected backend keeps state there (`NeedsDataDir`); `backends.go` holds each layer's `.backend` (only the in-process value today); `config.go` holds `roles` (`Has(Role)`) and `instance_id`, and `Validate` refuses a role split the backends cannot serve (any split over the embedded MQ; `api` without `ingest`, or the reverse, over a local cache) — boot is the validator, there is no dry run diff --git a/CHANGELOG.md b/CHANGELOG.md index 4cc140370..2521ef894 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A Redis-compatible shared cache backend, built but not yet selectable** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/cache.go`, `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`, `CONTRIBUTING.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), open a circuit breaker that skips the server until a probe write succeeds (connections are replaced every minute, so one that outlived a failover behind a stable address reaches the new primary), and deferred invalidations are retried until they land — the first at once, and at the probe's cadence while the breaker is open — collapsing to one tenant-wide bump per tenant past 100,000 keys; until one lands, the process that owes it bypasses the lookups it would orphan. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. Nothing selects it yet: `cache.backend` accepts only `local` until the wiring (E4) adds `redis` and the `cache.redis.*` settings, so every deployment still runs `LocalCache`. Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression, server-stops-answering, refused-writes (a demoted primary, a full `noeviction` server), failover-behind-a-stable-address and owed-invalidation cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. +- **A Redis-compatible shared cache backend, built but not yet selectable** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/cache.go`, `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`, `CONTRIBUTING.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), open a circuit breaker that skips the server until a probe write succeeds (connections are replaced every minute, so after a failover behind a stable address the process reaches the new primary, and delivers the bumps it owes, within about a minute), and deferred invalidations are retried until they land — the first at once, and at the probe's cadence while the breaker is open — collapsing to one tenant-wide bump per tenant past 100,000 keys; until one lands, the process that owes it bypasses the lookups it would orphan. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. Nothing selects it yet: `cache.backend` accepts only `local` until the wiring (E4) adds `redis` and the `cache.redis.*` settings, so every deployment still runs `LocalCache`. Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression, server-stops-answering, refused-writes (a demoted primary, a full `noeviction` server), failover-behind-a-stable-address and owed-invalidation cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. - **Leases for work that must run in one process at a time, and the sweeper runs under one** (`internal/coord/` (new: `coord.go`, `local.go`, `elect.go`, `coordtest/`, + tests), `internal/app/{app,wire}.go` (+ tests), `internal/config/backends.go`, `internal/ingest/sweeper.go`, `.github/labeler.yml`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development,ingest-pipeline}.md`, `docs/src/content/docs/configuration.mdx`, `config.yaml`, `AGENTS.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `coord.Coordinator` hands out named leases (`TryAcquire` → a `Term` with a strictly increasing fencing `Token`, a `Done` channel and `Resign`; `ErrHeld` while another holder's term is live), and `coord.RunElected` runs a loop only while its process holds the lease, resigning when the loop returns and campaigning again every 2s. `coord.Local` is the in-process implementation, and `coordtest.Conformance` is the suite every implementation runs — the NATS KV backend that lets several replicas share one queue comes next. The sweeper now runs through `RunElected` under the `sweeper` lease; with the in-process coordinator the one process always holds it, so nothing changes for a single-process deployment beyond one `coord: elected` log line at startup. `coord.backend` now selects the coordinator (`local`, the only value), so a `config.Config` built without `config.Load` must name it as well as the other three layers' backends. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index ac449529b..dfd1b1400 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -115,7 +115,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on — one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)) — each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope by their raw names, which the cache escapes where it builds a key. `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The snapshot is taken before any input a bump invalidates is chosen, the tenant's connection included: a reload that moves the tenant to another address or database runs `Pools.Reconcile` and then `InvalidateTenant` (a repoint that keeps both, such as a username or `tls` change, reads the same tables and bumps nothing), so a request that took the old pool files the old database's rows under a version that bump orphans, whether its `Set` lands before the bump or after. The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. - **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: a namespace key is `..
.
.`, its fields joined with `keyenc.Join` so a dot in a table or scope name is escaped rather than read as a separator, and an entry's key, `|.||…` with the caller's query key escaped whole (its `:` become `%3A`), folds in the tenant's version and each dependency's namespace key and namespace version, so bumping a table (a scopeless write) or one scope — scope is reserved and empty today, so every write is the whole-table bump — orphans every dependent entry without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results, and the flat directory's single tenant simply carries the `0` prefix. `BumpTenant` (behind `InvalidateTenant`) advances the tenant version that leads every namespace key of one tenant, orphaning its every namespace and every cached query in one step (the tenant version is folded into every query key, so a pipe result with no dependencies is orphaned too) — a table no bump ever keyed included, which an enumeration of the index would miss — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. -- **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET` and `MGET`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot; the table and scope are escaped with `keyenc` after the fixed prefix, so a `:` in a name is never read as the separator. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, the hash over their escaped forms, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token (a string of another length, a list, a hash) is replaced, so a token lost to eviction, expiry or a restart is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. While this process owes a bump on any of a lookup's tokens, that lookup is a bypass that files nothing: the bump would orphan whatever it found (other processes, which cannot know, serve those entries until it lands). `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing, each attempt bounded by `DialTimeout` (1 s), and a cluster client's topology read after the handshake by the larger of it and `Timeout`, which is also how long a connection waits on a silent server before it is redialed. Nothing selects this backend yet — the boot config and wiring arrive with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s E4. +- **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET` and `MGET`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot; the table and scope are escaped with `keyenc` after the fixed prefix, so a `:` in a name is never read as the separator. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, the hash over their escaped forms, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token (a string of another length, a list, a hash) is replaced, so a token lost to eviction, expiry or a restart is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. While this process owes a bump on any of a lookup's tokens, that lookup is a bypass that files nothing: the bump would orphan whatever it found (other processes, which cannot know, serve those entries until it lands). `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing, each attempt bounded by `DialTimeout` (1 s), and a cluster client's topology read after the handshake by the larger of it and `Timeout`, which is also how long a connection waits on a silent server before it is redialed. Nothing selects this backend yet: `cache.backend` accepts only `local` until [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s E4 adds `redis`, the `cache.redis.*` settings and the wiring. - **redis_codec.go** — the key schema (`tokenKeys`, `bumpKeys`, `valueKey`, which take a `Namespace`'s raw names and escape them) and the value frame: format, flags, expiry, the token list, the payload, zstd-compressed from `CompressMinBytes` when that is smaller. A stored value is capped at `MaxValueBytes` and a decoded one at eight times that, which refuses a zip bomb planted in a shared server; an unknown format, as a newer process writes during a rolling upgrade, is a miss. - **breaker.go** — the circuit breaker: `BreakerThreshold` (5) consecutive failures open it, and while open every operation skips the server at once; after `BreakerOpenFor` (5 s) one background write (`SET :probe`) decides whether it closes, and only that probe's success closes it. Every connection is replaced after a minute (rueidis retries what was in flight), because one that outlives a failover behind a stable address stays on the demoted node, which answers but refuses writes; the replacement re-resolves the address, so a bypassed process reaches the new primary, and delivers the bumps it owes, within about that long. A transport failure or a timeout counts against the server; an error reply saying it takes no writes right now — `READONLY` (a demoted primary), `MASTERDOWN`, `OOM` (full under `maxmemory-policy noeviction`, Redis's default, so give the cache an evicting policy), `NOREPLICAS`, `MISCONF`, `LOADING`, `BUSY`, `CLUSTERDOWN` — opens it at once, whatever the threshold, since no bump can land. Any other reply, an error reply about one key (`WRONGTYPE`, `NOPERM`, `TRYAGAIN`) or one the backend cannot use included, counts as a success, and a caller that gave up first counts as nothing. - **pending.go** — invalidations the server did not take, kept per token key (repeats coalesce) and retried by a background loop until they land: the first one owed at once, then with backoff from 100 ms to 10 s, and while the breaker is open at each probe, which the loop starts when due, so a process that makes no lookups (ingest only) recovers as soon as one that does. `Invalidate` sends its bumps 1,000 to a round trip, as the retry does. Landing late is still correct: a fresh token orphans the pre-write entries and any fill made meanwhile. Past `PendingMax` (100,000) keys the set collapses to one tenant bump per affected tenant. `Close` makes one last attempt at the bumps still owed, past the breaker (an open one is why they are owed); what that attempt cannot deliver is lost, and the entries they would orphan are served until their TTL — the same failure as a worker stopping between an insert and its invalidation. From 1f24da5fff5b087112cc34355cbc12d7a8185a75 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 06:16:03 -0400 Subject: [PATCH 52/79] docs(cache): a reconnect must fit in the operation timeout rueidis dials a replacement connection lazily, under the timeout of the operation that needs it, not DialTimeout. Recovery after a failover within about a minute therefore assumes a fresh connection's dial and handshake, TLS included, fit in Timeout; say so, and that Timeout should be sized above it. Co-Authored-By: Claude Opus 5.5 (1M context) --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 2521ef894..563dd2e76 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A Redis-compatible shared cache backend, built but not yet selectable** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/cache.go`, `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`, `CONTRIBUTING.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), open a circuit breaker that skips the server until a probe write succeeds (connections are replaced every minute, so after a failover behind a stable address the process reaches the new primary, and delivers the bumps it owes, within about a minute), and deferred invalidations are retried until they land — the first at once, and at the probe's cadence while the breaker is open — collapsing to one tenant-wide bump per tenant past 100,000 keys; until one lands, the process that owes it bypasses the lookups it would orphan. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. Nothing selects it yet: `cache.backend` accepts only `local` until the wiring (E4) adds `redis` and the `cache.redis.*` settings, so every deployment still runs `LocalCache`. Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression, server-stops-answering, refused-writes (a demoted primary, a full `noeviction` server), failover-behind-a-stable-address and owed-invalidation cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. +- **A Redis-compatible shared cache backend, built but not yet selectable** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/cache.go`, `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`, `CONTRIBUTING.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), open a circuit breaker that skips the server until a probe write succeeds (connections are replaced every minute, so after a failover behind a stable address the process reaches the new primary, and delivers the bumps it owes, within about a minute, provided a fresh connection's dial and handshake fit in the per-operation timeout), and deferred invalidations are retried until they land — the first at once, and at the probe's cadence while the breaker is open — collapsing to one tenant-wide bump per tenant past 100,000 keys; until one lands, the process that owes it bypasses the lookups it would orphan. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. Nothing selects it yet: `cache.backend` accepts only `local` until the wiring (E4) adds `redis` and the `cache.redis.*` settings, so every deployment still runs `LocalCache`. Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression, server-stops-answering, refused-writes (a demoted primary, a full `noeviction` server), failover-behind-a-stable-address and owed-invalidation cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. - **Leases for work that must run in one process at a time, and the sweeper runs under one** (`internal/coord/` (new: `coord.go`, `local.go`, `elect.go`, `coordtest/`, + tests), `internal/app/{app,wire}.go` (+ tests), `internal/config/backends.go`, `internal/ingest/sweeper.go`, `.github/labeler.yml`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development,ingest-pipeline}.md`, `docs/src/content/docs/configuration.mdx`, `config.yaml`, `AGENTS.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `coord.Coordinator` hands out named leases (`TryAcquire` → a `Term` with a strictly increasing fencing `Token`, a `Done` channel and `Resign`; `ErrHeld` while another holder's term is live), and `coord.RunElected` runs a loop only while its process holds the lease, resigning when the loop returns and campaigning again every 2s. `coord.Local` is the in-process implementation, and `coordtest.Conformance` is the suite every implementation runs — the NATS KV backend that lets several replicas share one queue comes next. The sweeper now runs through `RunElected` under the `sweeper` lease; with the in-process coordinator the one process always holds it, so nothing changes for a single-process deployment beyond one `coord: elected` log line at startup. `coord.backend` now selects the coordinator (`local`, the only value), so a `config.Config` built without `config.Load` must name it as well as the other three layers' backends. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index dfd1b1400..51239a095 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -117,7 +117,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: a namespace key is `..
.
.`, its fields joined with `keyenc.Join` so a dot in a table or scope name is escaped rather than read as a separator, and an entry's key, `|.||…` with the caller's query key escaped whole (its `:` become `%3A`), folds in the tenant's version and each dependency's namespace key and namespace version, so bumping a table (a scopeless write) or one scope — scope is reserved and empty today, so every write is the whole-table bump — orphans every dependent entry without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results, and the flat directory's single tenant simply carries the `0` prefix. `BumpTenant` (behind `InvalidateTenant`) advances the tenant version that leads every namespace key of one tenant, orphaning its every namespace and every cached query in one step (the tenant version is folded into every query key, so a pipe result with no dependencies is orphaned too) — a table no bump ever keyed included, which an enumeration of the index would miss — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. - **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET` and `MGET`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot; the table and scope are escaped with `keyenc` after the fixed prefix, so a `:` in a name is never read as the separator. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, the hash over their escaped forms, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token (a string of another length, a list, a hash) is replaced, so a token lost to eviction, expiry or a restart is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. While this process owes a bump on any of a lookup's tokens, that lookup is a bypass that files nothing: the bump would orphan whatever it found (other processes, which cannot know, serve those entries until it lands). `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing, each attempt bounded by `DialTimeout` (1 s), and a cluster client's topology read after the handshake by the larger of it and `Timeout`, which is also how long a connection waits on a silent server before it is redialed. Nothing selects this backend yet: `cache.backend` accepts only `local` until [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s E4 adds `redis`, the `cache.redis.*` settings and the wiring. - **redis_codec.go** — the key schema (`tokenKeys`, `bumpKeys`, `valueKey`, which take a `Namespace`'s raw names and escape them) and the value frame: format, flags, expiry, the token list, the payload, zstd-compressed from `CompressMinBytes` when that is smaller. A stored value is capped at `MaxValueBytes` and a decoded one at eight times that, which refuses a zip bomb planted in a shared server; an unknown format, as a newer process writes during a rolling upgrade, is a miss. -- **breaker.go** — the circuit breaker: `BreakerThreshold` (5) consecutive failures open it, and while open every operation skips the server at once; after `BreakerOpenFor` (5 s) one background write (`SET :probe`) decides whether it closes, and only that probe's success closes it. Every connection is replaced after a minute (rueidis retries what was in flight), because one that outlives a failover behind a stable address stays on the demoted node, which answers but refuses writes; the replacement re-resolves the address, so a bypassed process reaches the new primary, and delivers the bumps it owes, within about that long. A transport failure or a timeout counts against the server; an error reply saying it takes no writes right now — `READONLY` (a demoted primary), `MASTERDOWN`, `OOM` (full under `maxmemory-policy noeviction`, Redis's default, so give the cache an evicting policy), `NOREPLICAS`, `MISCONF`, `LOADING`, `BUSY`, `CLUSTERDOWN` — opens it at once, whatever the threshold, since no bump can land. Any other reply, an error reply about one key (`WRONGTYPE`, `NOPERM`, `TRYAGAIN`) or one the backend cannot use included, counts as a success, and a caller that gave up first counts as nothing. +- **breaker.go** — the circuit breaker: `BreakerThreshold` (5) consecutive failures open it, and while open every operation skips the server at once; after `BreakerOpenFor` (5 s) one background write (`SET :probe`) decides whether it closes, and only that probe's success closes it. Every connection is replaced after a minute (rueidis retries what was in flight), because one that outlives a failover behind a stable address stays on the demoted node, which answers but refuses writes; the replacement re-resolves the address, so a bypassed process reaches the new primary, and delivers the bumps it owes, within about that long. rueidis dials a replacement lazily, under the `Timeout` of the operation that needs it rather than `DialTimeout`, so this holds only where a fresh connection's dial and handshake (TLS included) fit in `Timeout`; size `Timeout` above that, or every reconnect, at the lifetime or after a drop, fails and the next one starts over. A transport failure or a timeout counts against the server; an error reply saying it takes no writes right now — `READONLY` (a demoted primary), `MASTERDOWN`, `OOM` (full under `maxmemory-policy noeviction`, Redis's default, so give the cache an evicting policy), `NOREPLICAS`, `MISCONF`, `LOADING`, `BUSY`, `CLUSTERDOWN` — opens it at once, whatever the threshold, since no bump can land. Any other reply, an error reply about one key (`WRONGTYPE`, `NOPERM`, `TRYAGAIN`) or one the backend cannot use included, counts as a success, and a caller that gave up first counts as nothing. - **pending.go** — invalidations the server did not take, kept per token key (repeats coalesce) and retried by a background loop until they land: the first one owed at once, then with backoff from 100 ms to 10 s, and while the breaker is open at each probe, which the loop starts when due, so a process that makes no lookups (ingest only) recovers as soon as one that does. `Invalidate` sends its bumps 1,000 to a round trip, as the retry does. Landing late is still correct: a fresh token orphans the pre-write entries and any fill made meanwhile. Past `PendingMax` (100,000) keys the set collapses to one tenant bump per affected tenant. `Close` makes one last attempt at the bumps still owed, past the breaker (an open one is why they are owed); what that attempt cannot deliver is lost, and the entries they would orphan are served until their TTL — the same failure as a worker stopping between an insert and its invalidation. - **metrics.go** — the shared backend's instruments (meter `wavehouse-cache`, `backend="redis"`, no tenant attribute): `wavehouse_cache_lookups_total` by `result` (`hit`, `miss`, `stale` — filed under since-bumped tokens, `bypass` — server skipped, or held by a bump this process owes, `error`), `wavehouse_cache_op_duration_seconds` by `op` (`lookup`, `set`, `invalidate`), `wavehouse_cache_breaker_open` (1 while bypassed, including before the first connection), `wavehouse_cache_invalidations_total` by `result` (`ok`, `deferred`), `wavehouse_cache_invalidations_pending` (bumps owed: entries they would orphan may be served stale meanwhile), `wavehouse_cache_value_bytes` (stored size), `wavehouse_cache_oversize_total` and `wavehouse_cache_set_failures_total` by `reason` (`oom`, `timeout`, `other`). From 98f4a128e2ffb1b5ee8e7a34b7b4e1a5ec755135 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 06:18:25 -0400 Subject: [PATCH 53/79] fix(api): match a statement's leading bareword whole isMutation read the leading verb as a run of letters, so `insert_log` or `insert2` matched INSERT. ClickHouse reads a bareword whole (letters, digits, `_`, `$`), as the WITH scanner already does; the leading scan now uses the same skipWord. No valid statement starts with such a word, so this only aligns the classifier with ClickHouse's lexer. Co-Authored-By: Claude Opus 5.5 (1M context) --- internal/api/clickhouse_exec.go | 11 ++--------- internal/api/clickhouse_exec_test.go | 5 ++++- 2 files changed, 6 insertions(+), 10 deletions(-) diff --git a/internal/api/clickhouse_exec.go b/internal/api/clickhouse_exec.go index e77e8ac2b..e34de64e7 100644 --- a/internal/api/clickhouse_exec.go +++ b/internal/api/clickhouse_exec.go @@ -117,7 +117,7 @@ var mutationVerbs = map[string]struct{}{ // isMutation reports whether sql's leading statement is a non-SELECT — i.e. // one that returns no result set and must go through Exec, not Query. // Leading whitespace and comments are skipped as ClickHouse's lexer skips -// them, then the first alphabetic token is matched case-insensitively against +// them, then the first bareword is matched whole, case-insensitively, against // mutationVerbs. A leading WITH clause (CTE) routes through a paren-aware scan // because ClickHouse accepts `WITH cte AS (...) INSERT INTO t SELECT * FROM // cte` as equivalent to `INSERT INTO t WITH cte AS (...) SELECT * FROM cte` @@ -126,14 +126,7 @@ var mutationVerbs = map[string]struct{}{ // fails the call, so a client that retries the error writes again. func isMutation(sql string) bool { s := stripLeadingSQLComments(sql) - end := 0 - for end < len(s) { - c := s[end] - if (c < 'A' || c > 'Z') && (c < 'a' || c > 'z') { - break - } - end++ - } + end := skipWord(s, 0) if end == 0 { return false } diff --git a/internal/api/clickhouse_exec_test.go b/internal/api/clickhouse_exec_test.go index 7cd861421..55aacdd50 100644 --- a/internal/api/clickhouse_exec_test.go +++ b/internal/api/clickhouse_exec_test.go @@ -128,7 +128,10 @@ func TestIsMutation(t *testing.T) { {"with tagged heredoc holding a paren and a verb then select", "WITH $x$ ) INSERT $x$ AS s SELECT s", false}, {"with CTE alias set$ (read)", "WITH set$ AS (SELECT 1 AS v) SELECT * FROM set$", false}, - // A word led by `_` is one bareword, never a keyword's tail. + // A word led by `_` is one bareword, never a keyword's tail, and a + // leading bareword is matched whole, never by its first letters. + {"leading bareword insert_log", "insert_log VALUES (1)", false}, + {"leading bareword insert2", "insert2 INTO t VALUES (1)", false}, {"with alias _delete (read)", "WITH 1 AS _delete SELECT _delete", false}, {"with alias _set (read)", "WITH [1,2] AS _set SELECT has(_set, 1)", false}, From 5b31ce6d97bc0b176e172219e2ee9cf63a7c4db0 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 06:23:37 -0400 Subject: [PATCH 54/79] fix(config): cap cache.redis timeout and dial_timeout at 1s Since #626 bounds a cluster client's topology read by the larger of the two, a dial held by boot or Close is a connect and a handshake, each up to dial_timeout, plus that read. At a 2s dial cap it could reach 6s, past the 5s release budget. Both caps are now 1s, so at most 3s. The shared-cache docs now describe #626's behavior: a refused write opens the breaker, the owing instance bypasses the lookups it would orphan, and connections are replaced every minute after a failover behind a stable address. Co-Authored-By: Claude Opus 5.5 (1M context) --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/configuration.mdx | 6 +++--- docs/src/content/docs/deployment.md | 6 +++--- internal/config/cache_redis.go | 23 ++++++++++++++++------- internal/config/cache_redis_test.go | 9 +++++---- tests/integration/shared_cache_test.go | 6 +++--- 7 files changed, 32 insertions(+), 22 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index ae2b6ac50..51ed88656 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `internal/ingest/worker.go`, `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`), `dial_timeout` (`1s`, at most `2s`: boot and shutdown each wait out a dial), `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The ingest worker's log of an invalidation that did not land drops from `ERROR` to `WARN`, since the shared backend defers and retries it: an outage would otherwise log an `ERROR` for every batch. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits; a third runs the first one's hit, insert and fresh-miss lifecycle on the suite's own `cache.backend: local` app, since e2e no longer exercises that backend. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. +- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `internal/ingest/worker.go`, `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`) and `dial_timeout` (`1s`), each at most `1s` since boot and shutdown each wait out a connection attempt they bound, `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The ingest worker's log of an invalidation that did not land drops from `ERROR` to `WARN`, since the shared backend defers and retries it: an outage would otherwise log an `ERROR` for every batch. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits; a third runs the first one's hit, insert and fresh-miss lifecycle on the suite's own `cache.backend: local` app, since e2e no longer exercises that backend. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. - **A Redis-compatible shared cache backend** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/cache.go`, `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`, `CONTRIBUTING.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), open a circuit breaker that skips the server until a probe write succeeds (connections are replaced every minute, so after a failover behind a stable address the process reaches the new primary, and delivers the bumps it owes, within about a minute, provided a fresh connection's dial and handshake fit in the per-operation timeout), and deferred invalidations are retried until they land — the first at once, and at the probe's cadence while the breaker is open — collapsing to one tenant-wide bump per tenant past 100,000 keys; until one lands, the process that owes it bypasses the lookups it would orphan. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. `cache.backend: redis` selects it (the entry above). Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression, server-stops-answering, refused-writes (a demoted primary, a full `noeviction` server), failover-behind-a-stable-address and owed-invalidation cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index fdab8bbe1..d3e479627 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -126,7 +126,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — the implementation each layer runs on, the process's `roles`, resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `cache.redis.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). - **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load` when `Config.NeedsDataDir` says a selected backend keeps state there, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. - **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` in `Validate`, between `validateRoles` and `validateTopology`), which refuses a value not on the list and names the ones that are. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns what a valid configuration is still likely to get wrong — the combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), a `cache.redis` block that is not read, certificate verification turned off — which `app.New` logs at `WARN`. -- **cache_redis.go** — `CacheRedisConfig`, the `cache.redis` sub-block, and its checks: an address, each `host:port` (a URL or `user:password@` form refused without repeating it, since it may hold a password), a known mode (`sentinel` is refused until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `db` 0 in cluster mode, positive timeouts and sizes, a `dial_timeout` of at most 2 s (boot and `Close` each wait out a dial, a connect and a handshake bounded by it), a `version_ttl` of at least 2 s, and a `compress_min_bytes` that is not negative (`0` never compresses); its defaults are in `defaults()` with the rest. `CacheRedisTLS.Config` builds the `tls.Config`, reading the files; `Validate` calls it so an unreadable file refuses boot, and `internal/app` calls it again to build the connection. A TLS key set while `tls.enabled` is off is an error rather than a plaintext connection. +- **cache_redis.go** — `CacheRedisConfig`, the `cache.redis` sub-block, and its checks: an address, each `host:port` (a URL or `user:password@` form refused without repeating it, since it may hold a password), a known mode (`sentinel` is refused until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `db` 0 in cluster mode, positive timeouts and sizes, a `timeout` and `dial_timeout` of at most 1 s each (boot and `Close` each wait out a dial: a connect and a handshake bounded by `dial_timeout`, and a cluster's topology read bounded by the larger of the two), a `version_ttl` of at least 2 s, and a `compress_min_bytes` that is not negative (`0` never compresses); its defaults are in `defaults()` with the rest. `CacheRedisTLS.Config` builds the `tls.Config`, reading the files; `Validate` calls it so an unreadable file refuses boot, and `internal/app` calls it again to build the connection. A TLS key set while `tls.enabled` is off is an error rather than a plaintext connection. - **config.go**, roles — `roles` (`[]Role`: `api`, `ingest`, `sweeper`; `AllRoles` by default; `Has(Role)`) picks which components `internal/app` wires, and `instance_id` names the process (`-<8 hex>` when empty, resolved in `Load`; today only logged at boot, and a distributed coordinator will record it as a lease's holder). `validateRoles` refuses an empty list, an empty entry, an unknown or a repeated role; `validateTopology` refuses a role set the backends cannot serve: any split over the embedded MQ, and a process with exactly one of `api` and `ingest` over a local cache. `NeedsDataDir` counts Pebble only for a process running `api`, and `Warnings` is empty without `api`, since only that role opens a cache it reads or a dedupe store. - **strict.go** — `rejectUnknownKeys`, the YAML half: re-reads the file as a generic tree and walks it against the struct's `yaml` tags, listing every key the struct doesn't declare. cleanenv itself is lenient by design, which is exactly wrong for boot config once keys have moved to the settings directory. - **persistence.go** — `WarnIfFreshDataDir` logs the startup `WARN` when `data_dir` is missing or empty (on a redeploy, the sign that the volume didn't persist); `LogStorageInitError` attaches the UID-65532 `permissionHint` to a NATS or Pebble open failure that looks like a permission denial — the same hint string `CheckDataDir` uses. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index d55936e55..509088aee 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -162,13 +162,13 @@ The `redis` backend's settings, read only when `cache.backend` is `redis`. It is | `cache.redis.tls.server_name` | `WH_CACHE_REDIS_TLS_SERVER_NAME` | *(empty)* | Name to verify the server's certificate against, when it differs from the address. | | `cache.redis.tls.insecure_skip_verify` | `WH_CACHE_REDIS_TLS_INSECURE_SKIP_VERIFY` | `false` | Accept any server certificate. Logged at `WARN` at boot: whoever can intercept the connection can read and replace cached results. | | `cache.redis.key_prefix` | `WH_CACHE_REDIS_KEY_PREFIX` | `wh` | Leads every key, so several deployments can share one server, provided you trust each as much as the others: any of them can overwrite what the rest serve. No `{` or `}`. | -| `cache.redis.timeout` | `WH_CACHE_REDIS_TIMEOUT` | `100ms` | Per operation. A lookup or fill that takes longer is a miss or a skipped fill, never a failed query. | -| `cache.redis.dial_timeout` | `WH_CACHE_REDIS_DIAL_TIMEOUT` | `1s` | Per connection attempt, at most `2s`. An attempt is a connect and a handshake, each bounded by this, and both boot and shutdown wait out one in flight, so the cap keeps a stop inside the release's fixed 5s budget (see [Stopping](/deployment#stopping)). | +| `cache.redis.timeout` | `WH_CACHE_REDIS_TIMEOUT` | `100ms` | Per operation, at most `1s`. A lookup or fill that takes longer is a miss or a skipped fill, never a failed query. | +| `cache.redis.dial_timeout` | `WH_CACHE_REDIS_DIAL_TIMEOUT` | `1s` | Per connection attempt, at most `1s`. An attempt is a connect and a handshake, each bounded by this, and in `cluster` mode a topology read bounded by the larger of this and `timeout`. Boot and shutdown each wait out one in flight, so the two caps keep that under 3s, inside a stop's fixed 5s release budget (see [Stopping](/deployment#stopping)). | | `cache.redis.max_value_bytes` | `WH_CACHE_REDIS_MAX_VALUE_BYTES` | `1048576` | Largest result stored, after compression (1 MiB). A larger one is returned to the caller but not cached. | | `cache.redis.compress_min_bytes` | `WH_CACHE_REDIS_COMPRESS_MIN_BYTES` | `1024` | Results at least this large are zstd-compressed when that makes them smaller. `0` never compresses. | | `cache.redis.version_ttl` | `WH_CACHE_REDIS_VERSION_TTL` | `168h` | How long a table's or tenant's version token outlives its last write, so the tokens of dropped tables and removed tenants eventually expire. At least `2s`. An expired token only causes misses. | -**When the server is unreachable or misbehaves, the cache is bypassed; queries are not.** A failure or a timeout makes the lookup a miss and the fill a no-op, and five in a row open a circuit breaker that skips the server entirely until a probe, every 5 s, gets an answer. Queries then go straight to ClickHouse, still coalesced per instance by `singleflight`. An invalidation the server did not take is kept and retried until it lands. `/readyz` does not depend on the cache. +**When the server is unreachable or misbehaves, the cache is bypassed; queries are not.** A failure or a timeout makes the lookup a miss and the fill a no-op. Five in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), open a circuit breaker that skips the server entirely until a probe write, every 5 s, succeeds. Queries then go straight to ClickHouse, still coalesced per instance by `singleflight`. An invalidation the server did not take is kept and retried until it lands, and until then the instance that owes it bypasses the lookups it would orphan. `/readyz` does not depend on the cache. **Boot does not wait for the server.** A malformed block (an address without a port, `mode: cluster` with `db` other than `0`, an unreadable or unparsable TLS file) refuses boot. A server that cannot be reached, or that refuses the credentials, does not: the process boots with the cache bypassed and keeps reconnecting, with backoff up to 30 s. A rejected password (`WRONGPASS`, `NOAUTH`) is logged at `ERROR` on every attempt; any other failure at `WARN`. This is deliberate: a rotated Redis password must not crash-loop every instance at once. Watch `wavehouse_cache_breaker_open`, which reads `1` while the cache is bypassed, including before the first connection. diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index fc366405d..1e14399b2 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -432,12 +432,12 @@ The query-result cache is the layer that can be shared today. With the default ` **What another instance can see.** Ingest is already asynchronous: `/v1/ingest` answers before the batch is inserted. Once the inserting instance's worker has written the batch to ClickHouse, it replaces the table's version token in Redis, and from then on a lookup on any instance misses and reads the new rows. The cache adds no delay of its own beyond that single write. The exceptions: - **The server is unreachable from the inserting instance.** The invalidation is kept and retried until it lands (`wavehouse_cache_invalidations_pending` counts what is owed). Meanwhile other instances that can still reach the server keep serving the older results, for as long as the outage lasts and at most until each entry's TTL. An instance that stops while invalidations are still owed loses them, with the same bound. The same thing happens today when a process stops between an insert and its invalidation. -- **A failover to a replica that had not yet received the latest token writes** can bring back entries filed under the older tokens, bounded by the replication lag at the moment of failover and those entries' TTL. WaveHouse never reads from replicas. -- **The server is full and `maxmemory-policy` is `noeviction`.** It refuses the token writes, so invalidations are kept and retried, and until one lands every instance serves the results from before the insert, up to their TTL. +- **A failover to a replica that had not yet received the latest token writes** can bring back entries filed under the older tokens, bounded by the replication lag at the moment of failover and those entries' TTL. WaveHouse never reads from replicas. Behind a stable address (a managed primary endpoint), an instance still connected to the demoted node has its writes refused, which bypasses its cache; connections are replaced every minute, so it reaches the new primary and delivers the invalidations it owes within about that long. +- **The server is full and `maxmemory-policy` is `noeviction`.** It refuses the token writes. The inserting instance keeps its invalidations and retries them, bypassing its cache meanwhile, but every other instance serves the results from before the insert until one lands, up to their TTL. - **A pipe that writes** (an `INSERT` in `pipes.json`) has its result cached like a read, so a repeated identical call is answered from the cache and the write does not run again ([#386](https://github.com/Wave-RF/WaveHouse/issues/386)). With a shared cache that holds on every instance, until the entry's TTL. - **Admin writes through `POST /v1/ops/query`** do not invalidate the cache ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). With a shared cache, the stale results they leave are served by every instance, not only one. -**Sizing the server.** Every key WaveHouse writes has a TTL, and a version token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses, never bring back an entry it had invalidated. So set `maxmemory` and let the server evict: `maxmemory-policy allkeys-lru` (or `allkeys-lfu`, `volatile-lru`, `volatile-lfu`). Under `noeviction`, a full server refuses the writes. Fills then fail (counted by `wavehouse_cache_set_failures_total{reason="oom"}`), only lookups whose version tokens already exist keep working, and invalidations are kept and retried, so the pre-insert results above stay served: avoid `noeviction`. A stored result is capped at `cache.redis.max_value_bytes` (1 MiB compressed). A tenant's version tokens share one hash tag, so each lookup reads them in one `MGET` in cluster mode as well. The results themselves carry no hash tag and spread across shards. Persistence is not needed: an empty server after a restart is a cold cache, not a wrong one. +**Sizing the server.** Every key WaveHouse writes has a TTL, and a version token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses, never bring back an entry it had invalidated. So set `maxmemory` and let the server evict: `maxmemory-policy allkeys-lru` (or `allkeys-lfu`, `volatile-lru`, `volatile-lfu`). Under `noeviction`, a full server refuses the writes: each refusal (a fill's is counted by `wavehouse_cache_set_failures_total{reason="oom"}`) bypasses the cache of the instance that got it, and invalidations are kept and retried, so the pre-insert results above stay served by the others: avoid `noeviction`. A stored result is capped at `cache.redis.max_value_bytes` (1 MiB compressed). A tenant's version tokens share one hash tag, so each lookup reads them in one `MGET` in cluster mode as well. The results themselves carry no hash tag and spread across shards. Persistence is not needed: an empty server after a restart is a cold cache, not a wrong one. **The server is inside the trust boundary.** A cached result is served after the access policy has filtered it, so whoever can write to the server can change what any caller reads. Keep it on a private network, require a password or ACL user (`WH_CACHE_REDIS_PASSWORD`), use TLS across links you do not trust, and share it only with deployments you trust as much as this one. diff --git a/internal/config/cache_redis.go b/internal/config/cache_redis.go index 920630157..df8174353 100644 --- a/internal/config/cache_redis.go +++ b/internal/config/cache_redis.go @@ -19,11 +19,12 @@ const ( RedisSentinel = "sentinel" ) -// maxRedisDialTimeout caps cache.redis.dial_timeout. A dial is a connect -// and a handshake, each bounded by it, and both boot and the cache's Close -// wait out one in flight: at 2s that is at most 4s, inside the 5s budget -// Close shares with the stores released after the cache. -const maxRedisDialTimeout = 2 * time.Second +// maxRedisTimeout caps cache.redis.timeout and cache.redis.dial_timeout. +// Boot and the cache's Close each wait out a dial in flight: a connect and +// a handshake, each bounded by dial_timeout, then for a cluster a topology +// read bounded by the larger of the two. At the caps that is at most 3s, +// inside the 5s budget Close shares with the stores released after it. +const maxRedisTimeout = time.Second // CacheRedisConfig configures cache.backend=redis: one Redis-compatible server // (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB) shared by every process. @@ -108,8 +109,16 @@ func (r CacheRedisConfig) validate() error { return fmt.Errorf("%s %s must be positive", d.key, d.v) } } - if r.DialTimeout > maxRedisDialTimeout { - return fmt.Errorf("cache.redis.dial_timeout (WH_CACHE_REDIS_DIAL_TIMEOUT) %s is over %s: boot and shutdown each wait out a dial, up to twice this", r.DialTimeout, maxRedisDialTimeout) + for _, d := range []struct { + key string + v time.Duration + }{ + {"cache.redis.timeout (WH_CACHE_REDIS_TIMEOUT)", r.Timeout}, + {"cache.redis.dial_timeout (WH_CACHE_REDIS_DIAL_TIMEOUT)", r.DialTimeout}, + } { + if d.v > maxRedisTimeout { + return fmt.Errorf("%s %s is over %s: boot and shutdown each wait out a connection attempt, which both bound", d.key, d.v, maxRedisTimeout) + } } if r.VersionTTL < 2*time.Second { return fmt.Errorf("cache.redis.version_ttl (WH_CACHE_REDIS_VERSION_TTL) %s is under 2s", r.VersionTTL) diff --git a/internal/config/cache_redis_test.go b/internal/config/cache_redis_test.go index 7f91a0499..264ed31d9 100644 --- a/internal/config/cache_redis_test.go +++ b/internal/config/cache_redis_test.go @@ -57,7 +57,7 @@ func TestLoad_CacheRedisFromEnv(t *testing.T) { "WH_CACHE_REDIS_TLS_SERVER_NAME": "redis.internal", "WH_CACHE_REDIS_KEY_PREFIX": "staging", "WH_CACHE_REDIS_TIMEOUT": "250ms", - "WH_CACHE_REDIS_DIAL_TIMEOUT": "2s", + "WH_CACHE_REDIS_DIAL_TIMEOUT": "500ms", "WH_CACHE_REDIS_MAX_VALUE_BYTES": "2048", "WH_CACHE_REDIS_COMPRESS_MIN_BYTES": "0", "WH_CACHE_REDIS_VERSION_TTL": "24h", @@ -73,7 +73,7 @@ func TestLoad_CacheRedisFromEnv(t *testing.T) { TLS: CacheRedisTLS{ Enabled: true, CAFile: caFile, CertFile: certFile, KeyFile: keyFile, ServerName: "redis.internal", }, - KeyPrefix: "staging", Timeout: 250 * time.Millisecond, DialTimeout: 2 * time.Second, + KeyPrefix: "staging", Timeout: 250 * time.Millisecond, DialTimeout: 500 * time.Millisecond, MaxValueBytes: 2048, CompressMinBytes: 0, VersionTTL: 24 * time.Hour, }, cfg.Cache.Redis) tc, err := cfg.Cache.Redis.TLS.Config() @@ -221,8 +221,9 @@ func TestValidate_CacheRedis(t *testing.T) { {"brace prefix", func(r *CacheRedisConfig) { r.KeyPrefix = "a{b}" }, "hash-tag brace"}, {"zero timeout", func(r *CacheRedisConfig) { r.Timeout = 0 }, "cache.redis.timeout (WH_CACHE_REDIS_TIMEOUT) 0s must be positive"}, {"negative dial timeout", func(r *CacheRedisConfig) { r.DialTimeout = -time.Second }, "cache.redis.dial_timeout"}, - {"dial timeout at the cap", func(r *CacheRedisConfig) { r.DialTimeout = 2 * time.Second }, ""}, - {"dial timeout over the cap", func(r *CacheRedisConfig) { r.DialTimeout = 2*time.Second + time.Millisecond }, "cache.redis.dial_timeout (WH_CACHE_REDIS_DIAL_TIMEOUT) 2.001s is over 2s"}, + {"timeouts at the cap", func(r *CacheRedisConfig) { r.Timeout, r.DialTimeout = time.Second, time.Second }, ""}, + {"timeout over the cap", func(r *CacheRedisConfig) { r.Timeout = time.Second + time.Millisecond }, "cache.redis.timeout (WH_CACHE_REDIS_TIMEOUT) 1.001s is over 1s"}, + {"dial timeout over the cap", func(r *CacheRedisConfig) { r.DialTimeout = time.Second + time.Millisecond }, "cache.redis.dial_timeout (WH_CACHE_REDIS_DIAL_TIMEOUT) 1.001s is over 1s"}, {"short version ttl", func(r *CacheRedisConfig) { r.VersionTTL = time.Second }, "cache.redis.version_ttl (WH_CACHE_REDIS_VERSION_TTL) 1s is under 2s"}, {"zero max value", func(r *CacheRedisConfig) { r.MaxValueBytes = 0 }, "cache.redis.max_value_bytes"}, {"compress never", func(r *CacheRedisConfig) { r.CompressMinBytes = 0 }, ""}, diff --git a/tests/integration/shared_cache_test.go b/tests/integration/shared_cache_test.go index e9857b725..f120365c3 100644 --- a/tests/integration/shared_cache_test.go +++ b/tests/integration/shared_cache_test.go @@ -80,7 +80,7 @@ func bootRedisApp(t *testing.T, redisAddr, prefix string, timeout time.Duration) MQ: config.MQ{Backend: config.MQEmbedded}, Cache: config.Cache{Backend: config.CacheRedis, Redis: config.CacheRedisConfig{ Addrs: []string{redisAddr}, Mode: config.RedisStandalone, KeyPrefix: prefix, - Timeout: timeout, DialTimeout: 2 * time.Second, + Timeout: timeout, DialTimeout: time.Second, MaxValueBytes: 1 << 20, CompressMinBytes: 1 << 10, VersionTTL: time.Hour, }}, Dedupe: config.Dedupe{Backend: config.DedupePebble}, @@ -149,8 +149,8 @@ func TestSharedCache_IngestOnOneInstanceInvalidatesAnother(t *testing.T) { table := createTable(t, "user_id String, value Float64", "ORDER BY user_id") _, redisAddr := startRedis(t) prefix := fmt.Sprintf("it%d", cachePrefixes.Add(1)) - a := bootRedisApp(t, redisAddr, prefix, 2*time.Second) - b := bootRedisApp(t, redisAddr, prefix, 2*time.Second) + a := bootRedisApp(t, redisAddr, prefix, time.Second) + b := bootRedisApp(t, redisAddr, prefix, time.Second) rc, err := rueidis.NewClient(rueidis.ClientOption{InitAddress: []string{redisAddr}, DisableCache: true, ForceSingleClient: true}) require.NoError(t, err) From 41ebc664b561d06123093141b738a9b66ff96318 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 06:36:31 -0400 Subject: [PATCH 55/79] =?UTF-8?q?fix(cache):=20review=20round=201=20?= =?UTF-8?q?=E2=80=94=20one=20standalone=20address,=20e2e=20stack=20docs,?= =?UTF-8?q?=20e2e=20gate?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - Standalone mode dials only the first address, so a second one was silently ignored: it now refuses boot. A port must be 1-65535. - The manual e2e boot command needs a Redis now that the fixture sets cache.backend: redis; the e2e stack is described as ClickHouse and Redis wherever it was ClickHouse only. - The e2e gate (59.9%) excludes cache_redis.go's rejection paths and pending.go's outage-only retries, as it excludes internal/settings; unit and integration cover both, and the merged total still counts them. - The integration tests read the TTL floor from cache.QueryTimeToTTL. - A failed InvalidateTenant on reconcile logs at WARN, like the worker. - The overview pages mention the shared cache. Co-Authored-By: Claude Opus 5.5 (1M context) --- .github/workflows/README.md | 2 +- .github/workflows/ci.yml | 4 ++-- .testcoverage.yml | 6 ++++++ AGENTS.md | 6 +++--- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/configuration.mdx | 2 +- docs/src/content/docs/development.md | 11 ++++++----- docs/src/content/docs/index.mdx | 2 +- docs/src/content/docs/sdk/reference.md | 4 ++-- docs/src/content/docs/why-wavehouse.md | 2 +- internal/app/app_test.go | 6 ++++-- internal/app/wire.go | 2 +- internal/config/cache_redis.go | 20 +++++++++++--------- internal/config/cache_redis_test.go | 9 +++++++-- tests/integration/shared_cache_test.go | 3 ++- 16 files changed, 50 insertions(+), 33 deletions(-) diff --git a/.github/workflows/README.md b/.github/workflows/README.md index 575a83893..470f82f8b 100644 --- a/.github/workflows/README.md +++ b/.github/workflows/README.md @@ -44,7 +44,7 @@ Break one of these knowingly or not at all. 2. **A dedicated `coverage` job applies the consolidated gate, polling — not `needs`-ing — the suites.** Each suite (`unit`, `integration`, `e2e`) runs with `COV_DEFER=1` and uploads a `coverage-` fragment; the `coverage` job runs `make cov` (merge + every threshold gate) over all three — exactly like local `make ci`'s final step. Keeping it a separate job (not folded into e2e's tail) decouples the gate result from the e2e suite's pass/fail. Crucially it is `needs: changes` **only, not the suites**: a `needs` edge is a *scheduling* barrier — GitHub won't pick up a runner, check out, restore caches, or `pnpm install` until the needed jobs finish — so needing the suites would serialize this job's ~50s of setup onto the critical path after the last suite, for nothing (the setup doesn't depend on their results). Instead it starts at run creation, runs its setup in parallel with the suites, and blocks only at the merge by polling for the three fragments with [`scripts/ci/wait-artifact.sh`](../../scripts/ci/wait-artifact.sh) (fails fast if a producer concluded without producing). Tail on the critical path: ~10s, not ~50s. **The aggregator and `docs-deploy` must keep `coverage` *and* every suite in their `needs`** — the suites directly (a suite failure must red the gate even though `coverage` no longer needs them), and `coverage` (else a coverage-gate failure wouldn't block merge or a prod deploy). -3. **e2e builds its own inputs and mirrors local `make test-e2e`.** It compiles the SDK dist + cover binary itself (`make -j test-e2e`, warm per-suffix cache) rather than waiting on a builder job, and runs the suite exactly as a developer does — one orchestrator, one ClickHouse testcontainer, sequential files. The ClickHouse image pulls in the background while caches restore (also in the integration job). +3. **e2e builds its own inputs and mirrors local `make test-e2e`.** It compiles the SDK dist + cover binary itself (`make -j test-e2e`, warm per-suffix cache) rather than waiting on a builder job, and runs the suite exactly as a developer does — one orchestrator, one ClickHouse and one Redis testcontainer, sequential files. The ClickHouse image pulls in the background while caches restore (also in the integration job). 4. **One change classifier, split into a pure core + a CI wrapper.** The pure allowlist — file list on stdin ⇒ `code`/`docs` — lives in [`scripts/classify-paths.sh`](../../scripts/classify-paths.sh), dependency-free and unit-tested by [`scripts/classify-paths.test.sh`](../../scripts/classify-paths.test.sh) (`make test-classify-paths`, a `verify` leaf) so the allowlists can't silently regress. The `changes` job runs the thin wrapper [`scripts/ci/classify-changes.sh`](../../scripts/ci/classify-changes.sh), which adds the CI-only policy (API file-list fetch + fail-closed: pushes, dispatches, API hiccups ⇒ `code=true`) on top. Keeping the core pure means the local git hooks can share it (`git diff --name-only | scripts/classify-paths.sh`). The `code`/`docs` outputs gate the suites and docs jobs — gate on these, never on workflow-level `paths:` filters, which would orphan the required check (invariant 1). diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index f47619274..fa572e414 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -280,8 +280,8 @@ jobs: with: go-cache-suffix: "-e2e-cov" # `-j` builds the prereqs (build-ts ∥ build-cover) concurrently, - # then runs the orchestrator: ClickHouse testcontainer + the cover - # binary + the SDK vitest suite. + # then runs the orchestrator: ClickHouse and Redis testcontainers + + # the cover binary + the SDK vitest suite. - name: Build SDK dist + cover binary, run E2E suite run: make -j "$(nproc)" test-e2e COV_DEFER=1 - name: Upload coverage fragment diff --git a/.testcoverage.yml b/.testcoverage.yml index 84f028afe..451014a32 100644 --- a/.testcoverage.yml +++ b/.testcoverage.yml @@ -84,3 +84,9 @@ exclude: # never reaches them. The unit suite and the integration suite's main # app (cache.backend=local) cover them; the merged total still counts them. - ^internal/cache/(local|version_manager)\.go$ + # What e2e's Redis never makes it run: the cache.redis block's rejection + # paths (boot adopts a valid fixture, as with internal/settings above) + # and the retry of invalidations the server did not take, which needs an + # outage. The unit and integration suites cover both. + - ^internal/config/cache_redis\.go$ + - ^internal/cache/pending\.go$ diff --git a/AGENTS.md b/AGENTS.md index 8feaab737..a04c4a8be 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -152,7 +152,7 @@ If `make ci` passes locally, your commit has crossed the same gates CI will run ### Running `make ci` (for agents) -`make ci` is **self-contained**: the integration suite (`tests/integration/`) and the E2E orchestrator (`scripts/orchestrator/`) each boot ClickHouse via **testcontainers on random host ports**, and the shared cache backend's integration tests (`internal/cache/`) start their own Redis, Valkey, Dragonfly and one-node Redis Cluster containers the same way. The only prerequisite is a running **Docker daemon** — do **not** `make deps-up` or start ClickHouse first (`deps-up` is for `make dev` only). +`make ci` is **self-contained**: the integration suite (`tests/integration/`) and the E2E orchestrator (`scripts/orchestrator/`) each boot ClickHouse and a Redis via **testcontainers on random host ports**, and the shared cache backend's integration tests (`internal/cache/`) start their own Redis, Valkey, Dragonfly and one-node Redis Cluster containers the same way. The only prerequisite is a running **Docker daemon** — do **not** `make deps-up` or start ClickHouse first (`deps-up` is for `make dev` only). Run it via the **background Bash tool** (`run_in_background: true`) and wait for the completion notification; the harness re-invokes you on exit, so polling the log with `tail` only burns context: @@ -449,8 +449,8 @@ internal/stream/ → SSE fan-out (event Hub: project once per role, Subsc internal/tenant/ → Tenant id (type, grammar, reserved default, request header name) internal/testutil/ → Shared test helpers (mocks, JWT + schema helpers; logtest/ captures or silences the default logger; cachetest/ is the conformance suite every cache.Cache backend runs) tests/ → Integration & E2E tests -tests/integration/ → Go integration tests (//go:build integration; ClickHouse testcontainer). A package tested against its own external server keeps them beside it: internal/cache/redis_integration_test.go (Redis, Valkey, Dragonfly, Redis Cluster testcontainers) -tests/e2e/ → E2E test stack (scripts/orchestrator boots a ClickHouse testcontainer + the wavehouse-cov binary) +tests/integration/ → Go integration tests (//go:build integration; ClickHouse testcontainer, and Redis for shared_cache_test.go). A package tested against its own external server keeps them beside it: internal/cache/redis_integration_test.go (Redis, Valkey, Dragonfly, Redis Cluster testcontainers) +tests/e2e/ → E2E test stack (scripts/orchestrator boots ClickHouse and Redis testcontainers + the wavehouse-cov binary) tests/e2e/fixtures/ → Idempotent ClickHouse DDL scripts for test tables tests/e2e/sdk/ → E2E integration tests via TypeScript SDK (Vitest) deployments/compose/ → Docker Compose files (standalone.yaml, dependencies.yaml) diff --git a/CHANGELOG.md b/CHANGELOG.md index 51ed88656..c4e6e2934 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `internal/ingest/worker.go`, `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`) and `dial_timeout` (`1s`), each at most `1s` since boot and shutdown each wait out a connection attempt they bound, `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a port, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The ingest worker's log of an invalidation that did not land drops from `ERROR` to `WARN`, since the shared backend defers and retries it: an outage would otherwise log an `ERROR` for every batch. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits; a third runs the first one's hit, insert and fresh-miss lifecycle on the suite's own `cache.backend: local` app, since e2e no longer exercises that backend. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. +- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `internal/ingest/worker.go`, `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`) and `dial_timeout` (`1s`), each at most `1s` since boot and shutdown each wait out a connection attempt they bound, `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a valid port, more than one address in `standalone` mode, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The ingest worker's log of an invalidation that did not land drops from `ERROR` to `WARN`, since the shared backend defers and retries it: an outage would otherwise log an `ERROR` for every batch. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits; a third runs the first one's hit, insert and fresh-miss lifecycle on the suite's own `cache.backend: local` app, since e2e no longer exercises that backend. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. - **A Redis-compatible shared cache backend** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/cache.go`, `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`, `CONTRIBUTING.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), open a circuit breaker that skips the server until a probe write succeeds (connections are replaced every minute, so after a failover behind a stable address the process reaches the new primary, and delivers the bumps it owes, within about a minute, provided a fresh connection's dial and handshake fit in the per-operation timeout), and deferred invalidations are retried until they land — the first at once, and at the probe's cadence while the breaker is open — collapsing to one tenant-wide bump per tenant past 100,000 keys; until one lands, the process that owes it bypasses the lookups it would orphan. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. `cache.backend: redis` selects it (the entry above). Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression, server-stops-answering, refused-writes (a demoted primary, a full `noeviction` server), failover-behind-a-stable-address and owed-invalidation cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index d3e479627..7511a2a9b 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -126,7 +126,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — the implementation each layer runs on, the process's `roles`, resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `cache.redis.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). - **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load` when `Config.NeedsDataDir` says a selected backend keeps state there, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. - **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` in `Validate`, between `validateRoles` and `validateTopology`), which refuses a value not on the list and names the ones that are. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns what a valid configuration is still likely to get wrong — the combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), a `cache.redis` block that is not read, certificate verification turned off — which `app.New` logs at `WARN`. -- **cache_redis.go** — `CacheRedisConfig`, the `cache.redis` sub-block, and its checks: an address, each `host:port` (a URL or `user:password@` form refused without repeating it, since it may hold a password), a known mode (`sentinel` is refused until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `db` 0 in cluster mode, positive timeouts and sizes, a `timeout` and `dial_timeout` of at most 1 s each (boot and `Close` each wait out a dial: a connect and a handshake bounded by `dial_timeout`, and a cluster's topology read bounded by the larger of the two), a `version_ttl` of at least 2 s, and a `compress_min_bytes` that is not negative (`0` never compresses); its defaults are in `defaults()` with the rest. `CacheRedisTLS.Config` builds the `tls.Config`, reading the files; `Validate` calls it so an unreadable file refuses boot, and `internal/app` calls it again to build the connection. A TLS key set while `tls.enabled` is off is an error rather than a plaintext connection. +- **cache_redis.go** — `CacheRedisConfig`, the `cache.redis` sub-block, and its checks: an address (exactly one in `standalone` mode, which dials only the first), each `host:port` with a port from 1 to 65535 (a URL or `user:password@` form refused without repeating it, since it may hold a password), a known mode (`sentinel` is refused until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `db` 0 in cluster mode, positive timeouts and sizes, a `timeout` and `dial_timeout` of at most 1 s each (boot and `Close` each wait out a dial: a connect and a handshake bounded by `dial_timeout`, and a cluster's topology read bounded by the larger of the two), a `version_ttl` of at least 2 s, and a `compress_min_bytes` that is not negative (`0` never compresses); its defaults are in `defaults()` with the rest. `CacheRedisTLS.Config` builds the `tls.Config`, reading the files; `Validate` calls it so an unreadable file refuses boot, and `internal/app` calls it again to build the connection. A TLS key set while `tls.enabled` is off is an error rather than a plaintext connection. - **config.go**, roles — `roles` (`[]Role`: `api`, `ingest`, `sweeper`; `AllRoles` by default; `Has(Role)`) picks which components `internal/app` wires, and `instance_id` names the process (`-<8 hex>` when empty, resolved in `Load`; today only logged at boot, and a distributed coordinator will record it as a lease's holder). `validateRoles` refuses an empty list, an empty entry, an unknown or a repeated role; `validateTopology` refuses a role set the backends cannot serve: any split over the embedded MQ, and a process with exactly one of `api` and `ingest` over a local cache. `NeedsDataDir` counts Pebble only for a process running `api`, and `Warnings` is empty without `api`, since only that role opens a cache it reads or a dedupe store. - **strict.go** — `rejectUnknownKeys`, the YAML half: re-reads the file as a generic tree and walks it against the struct's `yaml` tags, listing every key the struct doesn't declare. cleanenv itself is lenient by design, which is exactly wrong for boot config once keys have moved to the settings directory. - **persistence.go** — `WarnIfFreshDataDir` logs the startup `WARN` when `data_dir` is missing or empty (on a redeploy, the sign that the volume didn't persist); `LogStorageInitError` attaches the UID-65532 `permissionHint` to a NATS or Pebble open failure that looks like a permission denial — the same hint string `CheckDataDir` uses. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 509088aee..d625b2e46 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -150,7 +150,7 @@ The `redis` backend's settings, read only when `cache.backend` is `redis`. It is | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | -| `cache.redis.addrs` | `WH_CACHE_REDIS_ADDRS` | *(empty)* | **Required** with `backend: redis`. `host:port` of the server; several are a cluster's seed nodes. Comma-separated in the env var. Not a URL: `redis://user:password@host:port` refuses boot, without repeating the value, so set the credentials through `username` and `password`, and `rediss://` through `tls.enabled`. | +| `cache.redis.addrs` | `WH_CACHE_REDIS_ADDRS` | *(empty)* | **Required** with `backend: redis`. `host:port` of the server: exactly one in `standalone` mode, where a second would be ignored and so refuses boot; several are a cluster's seed nodes. Comma-separated in the env var. Not a URL: `redis://user:password@host:port` refuses boot, without repeating the value, so set the credentials through `username` and `password`, and `rediss://` through `tls.enabled`. | | `cache.redis.mode` | `WH_CACHE_REDIS_MODE` | `standalone` | `standalone` or `cluster`. `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656): the cache does not yet authenticate to the sentinels or refresh their topology, so it could not be trusted to follow a failover. | | `cache.redis.username` | `WH_CACHE_REDIS_USERNAME` | *(empty)* | ACL user. Empty uses the server's `default` user. | | `cache.redis.password` | `WH_CACHE_REDIS_PASSWORD` | *(empty)* | A secret: set it through the environment (or your secret store's env injection), not in a tracked `config.yaml`. | diff --git a/docs/src/content/docs/development.md b/docs/src/content/docs/development.md index 7b3ec0e65..92785bda9 100644 --- a/docs/src/content/docs/development.md +++ b/docs/src/content/docs/development.md @@ -16,7 +16,7 @@ You need these on your `PATH` before any `make` recipe will work end-to-end: | **Go** | 1.26+ (matches `go.mod`) | Compiles `cmd/wavehouse`; also runs the pinned `tool` deps (`gotestsum`, `gofumpt`, `goimports`, `govulncheck`, `deadcode`, `gsa`, `goda`) via `go tool` | [go.dev/dl](https://go.dev/dl/) | | **GNU Make** | **4.0+** | The Makefile uses `--output-sync=target` (Make 4 only) and bash-pinned recipes. macOS ships with BSD Make 3.81, which **will not work** | macOS: `brew install make` then use `gmake` or put `$(brew --prefix make)/libexec/gnubin` on your PATH. Linux: usually already installed | | **bash** | 4+ recommended | Recipes are pinned to `bash`; the helper scripts under `scripts/` use `set -euo pipefail` and bash arrays | macOS default is bash 3.2 (works for current recipes, but `brew install bash` is safer); Linux distros ship 4+ | -| **Docker** *(or Podman)* | Engine 20.10+ with the Compose **v2** plugin (`docker compose`, no hyphen) | Compose stacks under `deployments/compose/`; the E2E and integration suites boot ClickHouse via testcontainers (no compose file), and the integration suite also starts Redis, Valkey, Dragonfly (pulled from `docker.dragonflydb.io`) and a one-node Redis Cluster for the shared cache backend | [Docker Desktop](https://docs.docker.com/get-docker/), [colima](https://github.com/abiosoft/colima), or [Podman](https://podman.io) with `podman-compose` / the `podman compose` plugin. The testcontainers Go library also honors `DOCKER_HOST` for rootless Podman setups | +| **Docker** *(or Podman)* | Engine 20.10+ with the Compose **v2** plugin (`docker compose`, no hyphen) | Compose stacks under `deployments/compose/`; the E2E and integration suites boot ClickHouse and a Redis via testcontainers (no compose file), and the integration suite also starts Valkey, Valkey, Dragonfly (pulled from `docker.dragonflydb.io`) and a one-node Redis Cluster for the shared cache backend | [Docker Desktop](https://docs.docker.com/get-docker/), [colima](https://github.com/abiosoft/colima), or [Podman](https://podman.io) with `podman-compose` / the `podman compose` plugin. The testcontainers Go library also honors `DOCKER_HOST` for rootless Podman setups | | **Node.js** | 22 LTS — pinned via `.nvmrc` at the repo root | Runtime for pnpm and the Vitest suites. Pinned to match CI (`setup-node` uses 22) and to avoid Node-major surprises; older Vitest versions in this repo were known to crash on Node 26 with a V8 heap-allocation abort | [nodejs.org](https://nodejs.org/) or `nvm use` / `fnm use` / `volta` (all read `.nvmrc`) | | **pnpm** | 11.21+ (pinned via `packageManager` in the root `package.json`) | Package manager for the TypeScript SDK, E2E test harness, and docs site (managed as a single pnpm workspace from the repo root); `make build-ts`, `make test-ts`, `make test-e2e`, `make build-docs`, `make dev-docs`, `make preview-docs` all shell out to `pnpm` | `corepack enable && corepack prepare pnpm@11.21.0 --activate` (recommended), or `npm i -g pnpm` | | **git** + **curl** | any recent | `git` for source + version metadata in builds; `curl` is used by the Makefile to fetch the pinned `golangci-lint` binary into `.bin/` | usually preinstalled | @@ -362,7 +362,7 @@ The primary E2E integration test suite lives in `tests/e2e/sdk/`. It uses the Ty **Architecture**: -- `scripts/orchestrator` — the E2E entrypoint behind `make test-e2e`: it starts a clean ClickHouse **testcontainer** per run, launches the `wavehouse-cov` binary on a random free port, runs the SDK suite against it, then SIGINTs the binary to flush coverage. No Compose file is involved. CI runs the exact same path. +- `scripts/orchestrator` — the E2E entrypoint behind `make test-e2e`: it starts a clean ClickHouse **testcontainer** and a Redis one (the fixture's shared cache, `cache.backend: redis`) per run, launches the `wavehouse-cov` binary on a random free port, runs the SDK suite against it, then SIGINTs the binary to flush coverage. No Compose file is involved. CI runs the exact same path. - `tests/e2e/sdk/setup.ts` — `globalSetup`. Probes the `CLICKHOUSE_URL` / `WAVEHOUSE_URL` the orchestrator injects, creates the per-suite tables, refreshes the schema, and writes the baseline policy into the run's settings directory (adopted via `POST /v1/ops/settings/reload` — files are the only write path). It starts nothing itself and fails fast if either URL isn't up. It also prints the active Node/undici version, warning when the local Node major differs from `.nvmrc` — a runtime-specific transport bug is otherwise indistinguishable from a code failure (see [#440](https://github.com/Wave-RF/WaveHouse/issues/440)). - `tests/e2e/sdk/helpers.ts` — JWT factories, typed client constructors, async wait helpers, direct ClickHouse query helper. @@ -375,10 +375,11 @@ make test-e2e `make test-e2e` builds `bin/wavehouse-cov` (coverage-instrumented) and runs the orchestrator under `scripts/orchestrator/` to wire ClickHouse + the cover binary into the suite. covdata flushes on SIGINT into `tmp/coverage/e2e/data/`. -The orchestrator always provisions its own stack — a fresh ClickHouse testcontainer plus `wavehouse-cov` on a random free port — so a running `make dev` on `:8080` is neither detected nor reused, and the two don't collide. To run vitest against a stack you manage yourself, start the server from the **repo root** with the E2E fixture config: +The orchestrator always provisions its own stack — fresh ClickHouse and Redis testcontainers plus `wavehouse-cov` on a random free port — so a running `make dev` on `:8080` is neither detected nor reused, and the two don't collide. To run vitest against a stack you manage yourself, start a Redis for the fixture's `cache.backend: redis`, then the server from the **repo root** with the E2E fixture config: ```bash -WH_CONFIG=tests/e2e/fixtures/config.yaml go run ./cmd/wavehouse +docker compose -f deployments/compose/dependencies.yaml --profile redis up -d +WH_CONFIG=tests/e2e/fixtures/config.yaml WH_CACHE_REDIS_ADDRS=localhost:6379 go run ./cmd/wavehouse ``` The fixture matters: the suite signs its tokens with its `sdk-dev-secret` and depends on its dedupe, DLQ, and 5s schema-refresh settings. Point the suite at a default `make dev` server (`jwt_secret: change-me-in-production`) and setup's schema calls are rejected, then global setup dies 30s later on a misleading `schema not refreshed within 30s`. The repo root matters too — the fixture's `settings.dir` is relative to the working directory. The fixture's settings directory (policy, pipes, and tunables) points at ClickHouse on `localhost:9000`; if yours isn't there, edit `clickhouse.addr` / `http_port` in `tests/e2e/fixtures/settings/config.json` (the orchestrator patches them itself for its testcontainer). @@ -474,7 +475,7 @@ WaveHouse/ │ └── testutil/ # Shared test helpers and mocks ├── tests/ # Integration & E2E tests │ ├── integration/ # Go integration tests (//go:build integration) -│ └── e2e/ # E2E suite (orchestrator + ClickHouse testcontainer) +│ └── e2e/ # E2E suite (orchestrator + ClickHouse and Redis testcontainers) │ ├── fixtures/ # ClickHouse DDL + config and settings-directory fixtures │ └── sdk/ # E2E specs driven through the TypeScript SDK (Vitest) ├── clients/ # Client SDKs diff --git a/docs/src/content/docs/index.mdx b/docs/src/content/docs/index.mdx index 3a6143ca5..205a6f845 100644 --- a/docs/src/content/docs/index.mdx +++ b/docs/src/content/docs/index.mdx @@ -97,7 +97,7 @@ If you're building user-facing analytics, **WaveHouse is like Supabase for Click Every event is broadcast to SSE subscribers **before** it's flushed to ClickHouse. Gap-fill from JetStream history for late-connecting clients. - Ristretto cache plus Go `singleflight` coalesces identical concurrent queries — dashboards survive thundering herds without an extra cache tier to operate. + Ristretto cache plus Go `singleflight` coalesces identical concurrent queries — dashboards survive thundering herds without an extra cache tier to operate. Running several instances? Share one Redis-compatible cache with `cache.backend: redis`. Per-table, per-role column and row-level policies with JWT claim templating, defined in the hot-reloadable settings directory. diff --git a/docs/src/content/docs/sdk/reference.md b/docs/src/content/docs/sdk/reference.md index fa6983da4..91df43af6 100644 --- a/docs/src/content/docs/sdk/reference.md +++ b/docs/src/content/docs/sdk/reference.md @@ -184,8 +184,8 @@ export interface ClicksRow { The SDK doubles as the E2E integration test harness. Tests in `tests/e2e/sdk/` exercise the full pipeline (ingest → ClickHouse → query) through the SDK, validating both the backend and the client library in one pass. ```bash -# Run all E2E tests: the orchestrator boots a ClickHouse testcontainer + -# the wavehouse-cov binary, then runs the SDK suite +# Run all E2E tests: the orchestrator boots ClickHouse and Redis +# testcontainers + the wavehouse-cov binary, then runs the SDK suite make test-e2e ``` diff --git a/docs/src/content/docs/why-wavehouse.md b/docs/src/content/docs/why-wavehouse.md index 766d62144..883d19284 100644 --- a/docs/src/content/docs/why-wavehouse.md +++ b/docs/src/content/docs/why-wavehouse.md @@ -152,7 +152,7 @@ flowchart TB | ---------- | --------- | --------- | | Durable ingest buffer | Kafka / Redpanda cluster (3+ brokers, Zookeeper/KRaft) | Embedded NATS JetStream | | Batch consumer | Custom Go/Rust/Java service you write and operate | Built in | -| Query cache | Redis + singleflight middleware you write | Built in (Ristretto + singleflight) | +| Query cache | Redis + singleflight middleware you write | Built in (Ristretto + singleflight; or one Redis shared by every instance, `cache.backend: redis`) | | Real-time push | WebSocket service + bridge from Kafka | Built in (`/v1/stream`) | | Schema validation | Custom code in ingest API | Built in (discovers `system.columns`) | | Row/column access control | Custom middleware or a dedicated service | Built in (Hasura-style, JWT-driven) | diff --git a/internal/app/app_test.go b/internal/app/app_test.go index 4b7d400b5..7a23330ed 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -806,14 +806,14 @@ func TestNew_RedisCacheRefusesAnUnreadableTLSFile(t *testing.T) { func TestRedisConfig_FromLoadedDefaults(t *testing.T) { t.Setenv("WH_SETTINGS_DIR", t.TempDir()) t.Setenv("WH_CACHE_BACKEND", "redis") - t.Setenv("WH_CACHE_REDIS_ADDRS", "a:6379,b:6379") + t.Setenv("WH_CACHE_REDIS_ADDRS", "a:6379") t.Setenv("WH_CACHE_REDIS_PASSWORD", "pw") loaded, err := config.Load(filepath.Join(t.TempDir(), "none.yaml")) require.NoError(t, err) got, err := redisConfig(loaded.Cache.Redis) require.NoError(t, err) assert.Equal(t, cache.RedisConfig{ - Addrs: []string{"a:6379", "b:6379"}, Mode: cache.RedisStandalone, Password: "pw", + Addrs: []string{"a:6379"}, Mode: cache.RedisStandalone, Password: "pw", KeyPrefix: cache.DefaultRedisKeyPrefix, Timeout: cache.DefaultRedisTimeout, DialTimeout: cache.DefaultRedisDialTimeout, MaxValueBytes: cache.DefaultRedisMaxValueBytes, CompressMinBytes: cache.DefaultRedisCompressMinBytes, VersionTTL: cache.DefaultRedisVersionTTL, @@ -821,12 +821,14 @@ func TestRedisConfig_FromLoadedDefaults(t *testing.T) { t.Setenv("WH_CACHE_REDIS_COMPRESS_MIN_BYTES", "0") t.Setenv("WH_CACHE_REDIS_MODE", "cluster") + t.Setenv("WH_CACHE_REDIS_ADDRS", "a:6379,b:6379") loaded, err = config.Load(filepath.Join(t.TempDir(), "none.yaml")) require.NoError(t, err) got, err = redisConfig(loaded.Cache.Redis) require.NoError(t, err) assert.Zero(t, got.CompressMinBytes, "the backend's never, not its default") assert.Equal(t, cache.RedisCluster, got.Mode) + assert.Equal(t, []string{"a:6379", "b:6379"}, got.Addrs, "a cluster's seeds") assert.Equal(t, cache.RedisSentinel, config.RedisSentinel) } diff --git a/internal/app/wire.go b/internal/app/wire.go index a7c82a972..fd1eb2e12 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -322,7 +322,7 @@ func (a *App) wireClickHouse() error { // cached is stale, so all of it is orphaned at once. for _, id := range stale { if err := a.cache.InvalidateTenant(a.stopCtx, id); err != nil { - slog.Error("cache invalidation of a stale tenant failed; it may serve stale rows until they expire", "tenant", id, "error", err) + slog.Warn("cache invalidation of a stale tenant did not land; it may serve stale rows until it does", "tenant", id, "error", err) } } }) diff --git a/internal/config/cache_redis.go b/internal/config/cache_redis.go index df8174353..be065266f 100644 --- a/internal/config/cache_redis.go +++ b/internal/config/cache_redis.go @@ -7,6 +7,7 @@ import ( "fmt" "net" "os" + "strconv" "strings" "time" ) @@ -78,9 +79,13 @@ func (r CacheRedisConfig) validate() error { if strings.TrimSpace(a) != a { return fmt.Errorf("cache.redis.addrs (WH_CACHE_REDIS_ADDRS) %q: no spaces around an address", a) } - if _, _, err := net.SplitHostPort(a); err != nil { + _, port, err := net.SplitHostPort(a) + if err != nil { return fmt.Errorf("cache.redis.addrs (WH_CACHE_REDIS_ADDRS) %q: want host:port: %w", a, err) } + if n, err := strconv.Atoi(port); err != nil || n < 1 || n > 65535 { + return fmt.Errorf("cache.redis.addrs (WH_CACHE_REDIS_ADDRS) %q: want host:port with a port from 1 to 65535", a) + } } switch r.Mode { case RedisStandalone, RedisCluster: @@ -89,6 +94,11 @@ func (r CacheRedisConfig) validate() error { default: return fmt.Errorf("cache.redis.mode (WH_CACHE_REDIS_MODE) %q: valid: %s, %s", r.Mode, RedisStandalone, RedisCluster) } + // Standalone dials the first address only, so a second one (a replica, + // say) would be silently ignored rather than failed over to. + if r.Mode == RedisStandalone && len(r.Addrs) > 1 { + return fmt.Errorf("cache.redis.addrs (WH_CACHE_REDIS_ADDRS) has %d addresses: mode standalone connects to one server; several are a cluster's seeds (mode cluster)", len(r.Addrs)) + } if r.DB < 0 { return fmt.Errorf("cache.redis.db (WH_CACHE_REDIS_DB) %d is negative", r.DB) } @@ -108,14 +118,6 @@ func (r CacheRedisConfig) validate() error { if d.v <= 0 { return fmt.Errorf("%s %s must be positive", d.key, d.v) } - } - for _, d := range []struct { - key string - v time.Duration - }{ - {"cache.redis.timeout (WH_CACHE_REDIS_TIMEOUT)", r.Timeout}, - {"cache.redis.dial_timeout (WH_CACHE_REDIS_DIAL_TIMEOUT)", r.DialTimeout}, - } { if d.v > maxRedisTimeout { return fmt.Errorf("%s %s is over %s: boot and shutdown each wait out a connection attempt, which both bound", d.key, d.v, maxRedisTimeout) } diff --git a/internal/config/cache_redis_test.go b/internal/config/cache_redis_test.go index 264ed31d9..3c1f7c05d 100644 --- a/internal/config/cache_redis_test.go +++ b/internal/config/cache_redis_test.go @@ -45,7 +45,7 @@ func TestLoad_CacheRedisFromEnv(t *testing.T) { caFile, certFile, keyFile := writeTestPKI(t, dir) for k, v := range map[string]string{ "WH_CACHE_BACKEND": "redis", - "WH_CACHE_REDIS_ADDRS": "r1:6379,r2:6379", + "WH_CACHE_REDIS_ADDRS": "r1:6379", "WH_CACHE_REDIS_MODE": "standalone", "WH_CACHE_REDIS_USERNAME": "wavehouse", "WH_CACHE_REDIS_PASSWORD": "s3cret", @@ -68,7 +68,7 @@ func TestLoad_CacheRedisFromEnv(t *testing.T) { cfg, err := Load("nonexistent.yaml") require.NoError(t, err) assert.Equal(t, CacheRedisConfig{ - Addrs: []string{"r1:6379", "r2:6379"}, Mode: RedisStandalone, + Addrs: []string{"r1:6379"}, Mode: RedisStandalone, Username: "wavehouse", Password: "s3cret", DB: 2, TLS: CacheRedisTLS{ Enabled: true, CAFile: caFile, CertFile: certFile, KeyFile: keyFile, ServerName: "redis.internal", @@ -211,6 +211,11 @@ func TestValidate_CacheRedis(t *testing.T) { {"no addrs", func(r *CacheRedisConfig) { r.Addrs = nil }, "cache.backend=redis needs cache.redis.addrs (WH_CACHE_REDIS_ADDRS)"}, {"addr with space", func(r *CacheRedisConfig) { r.Addrs = []string{"a:6379", " b:6379"} }, "no spaces around an address"}, {"addr without port", func(r *CacheRedisConfig) { r.Addrs = []string{"redis"} }, `cache.redis.addrs (WH_CACHE_REDIS_ADDRS) "redis": want host:port`}, + {"addr empty port", func(r *CacheRedisConfig) { r.Addrs = []string{"redis:"} }, `"redis:": want host:port with a port from 1 to 65535`}, + {"addr port not a number", func(r *CacheRedisConfig) { r.Addrs = []string{"redis:637x"} }, "a port from 1 to 65535"}, + {"addr port out of range", func(r *CacheRedisConfig) { r.Addrs = []string{"redis:99999"} }, "a port from 1 to 65535"}, + {"standalone with two addrs", func(r *CacheRedisConfig) { r.Addrs = []string{"a:6379", "b:6379"} }, "has 2 addresses: mode standalone connects to one server"}, + {"cluster with two seeds", func(r *CacheRedisConfig) { r.Mode, r.Addrs = RedisCluster, []string{"a:6379", "b:6379"} }, ""}, {"mode", func(r *CacheRedisConfig) { r.Mode = "replica" }, `cache.redis.mode (WH_CACHE_REDIS_MODE) "replica": valid: standalone, cluster`}, {"sentinel refused", func(r *CacheRedisConfig) { r.Mode = RedisSentinel }, `cache.redis.mode (WH_CACHE_REDIS_MODE) "sentinel" is not supported yet: the cache neither authenticates to the sentinels nor refreshes their topology (https://github.com/Wave-RF/WaveHouse/issues/656)`}, {"cluster", func(r *CacheRedisConfig) { r.Mode = RedisCluster }, ""}, diff --git a/tests/integration/shared_cache_test.go b/tests/integration/shared_cache_test.go index f120365c3..6ace9a764 100644 --- a/tests/integration/shared_cache_test.go +++ b/tests/integration/shared_cache_test.go @@ -23,6 +23,7 @@ import ( "github.com/testcontainers/testcontainers-go/wait" "github.com/Wave-RF/WaveHouse/internal/app" + "github.com/Wave-RF/WaveHouse/internal/cache" "github.com/Wave-RF/WaveHouse/internal/config" ) @@ -31,7 +32,7 @@ const redisImage = "redis:8.10.2-alpine" // minCacheTTL is cache.QueryTimeToTTL's floor: a fill made less than this // long ago cannot have expired, so a miss inside it is an invalidation. -const minCacheTTL = 10 * time.Second +var minCacheTTL = cache.QueryTimeToTTL(0) var cachePrefixes atomic.Uint64 From 64f09e8968194757cab996583a31a046bb4d6418 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 06:42:03 -0400 Subject: [PATCH 56/79] docs(dev): list Redis among the cache backend's integration servers Co-Authored-By: Claude Opus 5.5 (1M context) --- docs/src/content/docs/development.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/src/content/docs/development.md b/docs/src/content/docs/development.md index 92785bda9..581df72d1 100644 --- a/docs/src/content/docs/development.md +++ b/docs/src/content/docs/development.md @@ -16,7 +16,7 @@ You need these on your `PATH` before any `make` recipe will work end-to-end: | **Go** | 1.26+ (matches `go.mod`) | Compiles `cmd/wavehouse`; also runs the pinned `tool` deps (`gotestsum`, `gofumpt`, `goimports`, `govulncheck`, `deadcode`, `gsa`, `goda`) via `go tool` | [go.dev/dl](https://go.dev/dl/) | | **GNU Make** | **4.0+** | The Makefile uses `--output-sync=target` (Make 4 only) and bash-pinned recipes. macOS ships with BSD Make 3.81, which **will not work** | macOS: `brew install make` then use `gmake` or put `$(brew --prefix make)/libexec/gnubin` on your PATH. Linux: usually already installed | | **bash** | 4+ recommended | Recipes are pinned to `bash`; the helper scripts under `scripts/` use `set -euo pipefail` and bash arrays | macOS default is bash 3.2 (works for current recipes, but `brew install bash` is safer); Linux distros ship 4+ | -| **Docker** *(or Podman)* | Engine 20.10+ with the Compose **v2** plugin (`docker compose`, no hyphen) | Compose stacks under `deployments/compose/`; the E2E and integration suites boot ClickHouse and a Redis via testcontainers (no compose file), and the integration suite also starts Valkey, Valkey, Dragonfly (pulled from `docker.dragonflydb.io`) and a one-node Redis Cluster for the shared cache backend | [Docker Desktop](https://docs.docker.com/get-docker/), [colima](https://github.com/abiosoft/colima), or [Podman](https://podman.io) with `podman-compose` / the `podman compose` plugin. The testcontainers Go library also honors `DOCKER_HOST` for rootless Podman setups | +| **Docker** *(or Podman)* | Engine 20.10+ with the Compose **v2** plugin (`docker compose`, no hyphen) | Compose stacks under `deployments/compose/`; the E2E and integration suites boot ClickHouse and a Redis via testcontainers (no compose file), and the integration suite also runs the shared cache backend against Redis, Valkey, Dragonfly (pulled from `docker.dragonflydb.io`) and a one-node Redis Cluster | [Docker Desktop](https://docs.docker.com/get-docker/), [colima](https://github.com/abiosoft/colima), or [Podman](https://podman.io) with `podman-compose` / the `podman compose` plugin. The testcontainers Go library also honors `DOCKER_HOST` for rootless Podman setups | | **Node.js** | 22 LTS — pinned via `.nvmrc` at the repo root | Runtime for pnpm and the Vitest suites. Pinned to match CI (`setup-node` uses 22) and to avoid Node-major surprises; older Vitest versions in this repo were known to crash on Node 26 with a V8 heap-allocation abort | [nodejs.org](https://nodejs.org/) or `nvm use` / `fnm use` / `volta` (all read `.nvmrc`) | | **pnpm** | 11.21+ (pinned via `packageManager` in the root `package.json`) | Package manager for the TypeScript SDK, E2E test harness, and docs site (managed as a single pnpm workspace from the repo root); `make build-ts`, `make test-ts`, `make test-e2e`, `make build-docs`, `make dev-docs`, `make preview-docs` all shell out to `pnpm` | `corepack enable && corepack prepare pnpm@11.21.0 --activate` (recommended), or `npm i -g pnpm` | | **git** + **curl** | any recent | `git` for source + version metadata in builds; `curl` is used by the Makefile to fetch the pinned `golangci-lint` binary into `.bin/` | usually preinstalled | From 558bb23835691a2ca43c4b3b7afb85ce3d755972 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:30:43 -0400 Subject: [PATCH 57/79] fix(cache): pin the no-index bump test and fix a "query key" mixup MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit TestVersionManager_BumpWithoutIndex asserted only on vm.size() after its final BumpTenant call, which unconditionally deletes the tenant's index regardless of what the earlier no-index bumps did — so it could not fail against a tableLocked that wrongly creates an index instead of returning nil. Assert after each bump instead, and add the pruned-tenant case a write racing Prune must not revive. Also reword two docs passages that overloaded or misstated a term: architecture.md used "query key" for both the caller's input and the rendered entry key in the same paragraph; AGENTS.md's cache bullet read as if the index maps were keyed by escaped names; CHANGELOG.md's #382 bullet said the version index builds its keys with internal/keyenc, contradicting its own closing sentence that the index's entries are unaffected by the escaping. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- AGENTS.md | 2 +- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- internal/cache/version_manager_test.go | 18 +++++++++++++++++- 4 files changed, 20 insertions(+), 4 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index a504ffd11..a74dc6bb8 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -31,7 +31,7 @@ Twenty internal packages under `internal/` (plus `internal/testutil/` for shared - **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers; `ch_errors.go` (`writeCHError`) is the one mapping from a failed ClickHouse query to status, `code` and `retryable` - **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` handing the sweeper each tenant's own gap window (a rejected tenant's as its folder last had it, unbounded for one rejected since boot) and the `mq.max_bytes_gb` reconcile each served tenant's byte budget, and `defaultPolicy` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, the same hook's `Hub.Prune` ends the open streams of a tenant no longer served, and `wireCache`'s hook drops, through `LocalCache.Prune`, the cache version index of a tenant no longer served ([#262](https://github.com/Wave-RF/WaveHouse/issues/262))), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `New` wires only what the process's `roles` need (discovery, dedupe, auth verifiers, the hub bridge and keepalive per API process; the ingest worker per ingest process; the sweeper under its lease through `elected`); a process without `api` serves `api.NewOpsRouter` — probes, `/version`, metrics, and the settings reload behind the operator key alone. `cmd/wavehouse` and `tests/integration` both boot through it - **`auth/`** — JWT auth middleware: HMAC **or** JWKS verification with `alg` pinned to the active verifier, role extraction from a configurable claim path; always runs, never rejects (bad token → empty role + stashed reason). One verifier per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 9): `Authenticator` keys them by `tenant.ID` — the request store's `settings.Store.Tenant()`, through an injected `TenantSource`; `tenant.Default` on the tenant-exempt routes — built from each tenant's `auth` block by `Reconfigure`, dropped by `Prune` once the tenant stops being served (rejected or removed), released by `Close`; the secrets (`Config`) are boot-level and shared. A JWKS key set is fetched off the boot and reload paths: until one has been stored the verifier is pending and a token-bearing request gets `503` + `Retry-After` from `api.refuseUnverifiable` (`auth.ErrVerifierPending`), never a `default_role` evaluation; refresh is library-managed (Eric, 2026-09-22), response capped at 1 MiB; the operator key's admin role is the request tenant's -- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, a version per tenant, per (tenant, table) and per (tenant, table, scope), keyed by name and bumped in place (one entry per live namespace however often it is bumped, [#262](https://github.com/Wave-RF/WaveHouse/issues/262)), each raw table and scope name escaped by `keyenc` where the key is built (a `Namespace` carries them raw, so no caller escapes) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` drops the tenant's index so its next key gets a process-unique generation, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read, taken before the handler chooses any input a bump invalidates — the tenant's connection included — and `Set` files the fill under it, so a write landing mid-query, or a reload moving the tenant to another address or database after the request took its connection, orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) +- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for the caller's query key and its singleflight, escaped whole as the lead field of the stored key `|.||…`, where each raw table and scope name is escaped by `keyenc` (a `Namespace` carries them raw, so no caller escapes); the index holds a version per tenant, per (tenant, table) and per (tenant, table, scope), keyed by raw name and bumped in place (one entry per live namespace however often it is bumped, [#262](https://github.com/Wave-RF/WaveHouse/issues/262)) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` drops the tenant's index so its next key gets a process-unique generation, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read, taken before the handler chooses any input a bump invalidates — the tenant's connection included — and `Set` files the fill under it, so a write landing mid-query, or a reload moving the tenant to another address or database after the request took its connection, orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config. `Classify` (`errclass.go`) says what a failed ClickHouse request means for the request — `Unavailable`, `Denied`, `Rejected` (any unlisted exception code: the server read it and refused it), or `Unknown` (no code, no recognizable transport failure) — over the driver's error types and the HTTP interface's `HTTPError`; the ingest worker and the query handlers (`api/ch_errors.go` `writeCHError`, [#403](https://github.com/Wave-RF/WaveHouse/issues/403), [#271](https://github.com/Wave-RF/WaveHouse/issues/271)) both use it - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) - **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability when a selected backend keeps state there (`NeedsDataDir`); `backends.go` holds each layer's `.backend` (only the in-process value today); `config.go` holds `roles` (`Has(Role)`) and `instance_id`, and `Validate` refuses a role split the backends cannot serve (any split over the embedded MQ; `api` without `ingest`, or the reverse, over a local cache) — boot is the validator, there is no dry run diff --git a/CHANGELOG.md b/CHANGELOG.md index 6631b3c4a..14de673f7 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -88,7 +88,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **A failed ClickHouse query answers by what went wrong, not a flat `500`/`502`** (`internal/api/ch_errors.go` (new, + tests), `internal/api/{errors,query,structured_query,pipes,schema,ch_settings}.go`, `internal/chconn/errclass.go` (`HTTPStatus` exported), `clients/ts/src/errors.ts` (+ tests), `tests/integration/query_errors_test.go` (new), `tests/integration/query_limits_test.go`, `internal/app/app_test.go`, `tests/e2e/sdk/{admin,query}.test.ts`, `AGENTS.md`, `docs/src/content/docs/{api,architecture}.md`, `docs/src/content/docs/{access-control,configuration}.mdx`, `docs/src/content/docs/sdk/{reference.md,index.mdx}`): fixes [#403](https://github.com/Wave-RF/WaveHouse/issues/403) and [#271](https://github.com/Wave-RF/WaveHouse/issues/271), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). ClickHouse answers a syntax error, a missing grant and an overloaded server alike with HTTP `500`, so `/v1/ops/query` turned a bad statement into a `502` and `/v1/query` and pipes into a `500` the SDK retried. All three now class the failure with `chconn.Classify` through one helper, `writeCHError`: a statement ClickHouse refused is `400 clickhouse.rejected`; a query over a rows/bytes limit, the role's own memory cap, or its time cap where that is no longer than `query_timeout` is `400 clickhouse.limit_exceeded`; `ACCESS_DENIED` is `403 clickhouse.access_denied`; credentials, user or database refused, or a redirect or `4xx` with no exception code from whatever fronts ClickHouse, is `502 clickhouse.misconfigured`; ClickHouse down, unreachable or overloaded is `503 clickhouse.unavailable` with `Retry-After: 5`; a failure with no verdict stays `500` (`502` on the proxy) as `clickhouse.unknown`. The error envelope gains `code` and `retryable` next to `error` on these responses — additive. A role with `max_execution_time` now queries with no context deadline and a cancel two seconds past the cap instead: clickhouse-go overwrote the cap's `max_execution_time` with deadline+5s for any deadline over 1s, so an overrun came back as a bare deadline, indistinguishable from waiting for a pooled connection; ClickHouse now enforces the cap itself and reports `TIMEOUT_EXCEEDED`. `POST /v1/ops/schema/refresh` against an unreachable ClickHouse is a `503` with `Retry-After` instead of a `500`. **SDK:** `WaveHouseError.code` and `retryable` now take the server's `code`/`retryable` when the body has them (`HTTP_` and "5xx retries" otherwise), so a rejected query is `clickhouse.rejected` rather than `HTTP_500`, and is not retried. - **An unavailable ClickHouse is retried with backoff instead of dead-lettering every row** (`internal/chconn/errclass.go` (new, + tests), `internal/ingest/{worker,backoff}.go` (`backoff.go` new, + tests), `internal/mq/{mq,embedded}.go`, `internal/testutil/mocks.go`, `tests/integration/ingest_outage_test.go` (new), `AGENTS.md`, `README.md`, `docs/src/content/docs/{ingest-pipeline,architecture,api,deployment,why-wavehouse}.md`, `docs/src/content/docs/{settings-directory,index,access-control}.mdx`): workstream A of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). A failed batch insert used to go through row-by-row isolation whatever the failure, so a ClickHouse that was down, overloaded or read-only failed every row twice and parked the whole batch on the DLQ. `chconn.Classify` now classes the failure first — `Rejected` (any ClickHouse exception code outside the availability and credential lists: the server read the row and refused it), `Unavailable` (connection refused/reset, timeouts, `TOO_MANY_SIMULTANEOUS_QUERIES`, `SERVER_OVERLOADED`, `MEMORY_LIMIT_EXCEEDED`, `TOO_MANY_PARTS`, `READONLY`, `TABLE_IS_READ_ONLY`, `KEEPER_EXCEPTION`, …), `Denied` (`AUTHENTICATION_FAILED`, `ACCESS_DENIED`, …) or `Unknown` (no code, no recognizable transport failure). Only `Rejected` is isolated and dead-lettered as before, and a multi-row batch refused with `TOO_MANY_PARTS` or `MEMORY_LIMIT_EXCEEDED` is split row by row first (`chconn.Splittable`), because a batch spanning too many partitions or too much memory can fail where each of its rows inserts; every other class hands the batch back to the queue with a delayed nak (`mq.Message.NakWithDelay`, new) under a jittered 1 s → 30 s backoff shared by every table on the same ClickHouse pool (a failure of one table — read-only, too many parts or mutations, a grant missing on it, `chconn.TableScoped` — backs off that table alone), which turns rows away without a request while it runs and probes once per window, and ClickHouse going away mid-isolation stops isolation and retries the rows it had not settled. Counted by the new `wavehouse_ingest_retries_total{table, reason}`; logged at `WARN` when an outage starts and at most every 30 s during it. A long outage now shows as a growing ingest stream and, at `mq.max_bytes_gb`, ingest `503`s — not as a full DLQ; a lasting failure of one table holds back its tenant's other tables once its waiting rows reach `maxAckPending`. Retried rows come back out of arrival order, which matters only to a `ReplacingMergeTree` without a version column or a `CollapsingMergeTree`. - **Schema discovery's retry loop jitters its backoff** (`internal/discovery/discovery.go` (+ tests), `internal/app/wire.go`, `internal/api/errors.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`): `RetryRefresh` slept exactly `2s * 2^n` capped at 60s, so instances retrying against one recovering ClickHouse fired in lockstep, every 60s on the same second. Each sleep is now drawn uniformly from below the backoff (full jitter), spreading the retries over the whole window and halving the mean wait — so a failing tenant's retries, their log lines and `wavehouse_schema_refresh_failures_total` come about twice as often ([#141](https://github.com/Wave-RF/WaveHouse/issues/141)). -- **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/query/ident.go` (removed, + tests), `internal/app/wire.go`, `docs/src/content/docs/{api,architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL (a pipe result's key folded no version, so pipes were unaffected; with the tenant's version in every key they now take the same snapshot). The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. The snapshot is taken before any input a bump invalidates is chosen, the tenant's ClickHouse connection included: both handlers now look up before they resolve the tenant's pool, so a reload that moves the tenant to another address or database after a request took the old pool orphans that request's fill instead of filing the old database's rows as fresh under the new tenant version. A tenant on no pool is still a `503` before a cached result is served or a query runs. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. The cache now escapes table and scope names itself: a `Namespace` carries them raw and the version index builds its keys with `internal/keyenc`, so neither the structured-query read nor the ingest worker's invalidation escapes them (`query.SafeEncodeToken` is gone) and no name reaches a key unescaped. The suite pins that a name holding a dot, a space or a `%` is read and bumped under one key, and that names which would run together unescaped (`a.0.b` against `a` with scope `b.0.`) stay two entries. The version index's own entries are unaffected by the escaping; only the rendered key changes: it now carries the tenant's version and escapes the caller's query key whole. All of them live in the process, so nothing stored is orphaned. +- **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/query/ident.go` (removed, + tests), `internal/app/wire.go`, `docs/src/content/docs/{api,architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL (a pipe result's key folded no version, so pipes were unaffected; with the tenant's version in every key they now take the same snapshot). The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. The snapshot is taken before any input a bump invalidates is chosen, the tenant's ClickHouse connection included: both handlers now look up before they resolve the tenant's pool, so a reload that moves the tenant to another address or database after a request took the old pool orphans that request's fill instead of filing the old database's rows as fresh under the new tenant version. A tenant on no pool is still a `503` before a cached result is served or a query runs. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. The cache now escapes table and scope names itself: a `Namespace` carries them raw and the cache renders each result's key with `internal/keyenc`, so neither the structured-query read nor the ingest worker's invalidation escapes them (`query.SafeEncodeToken` is gone) and no name reaches a key unescaped. The suite pins that a name holding a dot, a space or a `%` is read and bumped under one key, and that names which would run together unescaped (`a.0.b` against `a` with scope `b.0.`) stay two entries. The version index's own entries are unaffected by the escaping; only the rendered key changes: it now carries the tenant's version and escapes the caller's query key whole. All of them live in the process, so nothing stored is orphaned. - **The cache's version index no longer grows with every bump, and forgets a tenant no longer served** (`internal/cache/{local,version_manager}.go` (+ tests), `internal/app/wire.go` (+ tests), `docs/src/content/docs/architecture.md`, `AGENTS.md`): part of [#262](https://github.com/Wave-RF/WaveHouse/issues/262) (growth across bumps and a departed tenant's memory; per-table scope cardinality is left open, see the issue) and of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The index nested each table under its tenant's version and each scope under its table's, and never pruned, so every tenant invalidation left the tenant's whole index behind, and it grew with every tenant ever served. It now holds one version per tenant, per (tenant, table) and per (tenant, table, scope), bumped in place. A tenant invalidation drops the tenant's index and hands its next key a generation unique within the process, so nothing cached before it can match again, and a table bump drops the table's scope versions. After each settings reload the index of every tenant no longer served, removed or rejected, is dropped the same way; its cached results are orphaned with it, as they already were when such a tenant came back on a pool. No change to what is cached or served. The Redis-compatible backend (#613) will bound its versions with a TTL instead. - **An explicit `false`, `0` or `""` in `config.yaml` is no longer replaced by the key's default** (`internal/config/config.go`, `internal/config/defaults_test.go` (new), `docs/src/content/docs/configuration.mdx`, `AGENTS.md`): [#631](https://github.com/Wave-RF/WaveHouse/issues/631). Defaults lived in cleanenv `env-default` tags, which cleanenv applies after the YAML decode to any field still at its zero value, so it could not tell a key the file set to its zero value from one the file left out. `otel.traces.enabled: false`, `otel.metrics.enabled: false` and `otel.logs.enabled: false` came back `true`; `otel.traces.sample_rate: 0` and `otel.logs.sample_rate: 0` came back `1.0`; `server.shutdown_timeout: 0` came back `10`; `cache.l1_max_cost: 0`, `prometheus.path: ""` and `data_dir: ""` came back as their defaults; `server.port: 0` came back `8080`. All of it was silent. Defaults now live in one Go function, `defaults()`, which `Load` starts from before decoding the file and then applying `WH_*` variables, so the order is env > YAML > default and a key the file sets always wins. **Behaviour change if your file relied on the bug:** a zero you wrote now takes effect. A file that says `sample_rate: 0` now exports no traces (or no DEBUG/INFO logs), where it silently exported everything; a signal set `enabled: false` is now off; `shutdown_timeout: 0` now skips the drain. `cache.l1_max_cost: 0`, `server.port: 0`, and `data_dir: ""` now refuse boot (`cache init: MaxCost can't be zero`, `server.port 0 out of range`, `data_dir (WH_DATA_DIR) is required`) instead of running on the default; an empty `prometheus.path` refuses boot when `prometheus.enabled` is true. Delete the key to get the default back. Env vars are unchanged: they already honoured an explicit zero. New tests load through `config.Load` for every affected key (a YAML zero is kept, an absent key gets the default, env wins in both directions), refuse an `env-default` tag on any field, and pin each documented default in `configuration.mdx` to `defaults()`. - **An embedded queue store that cannot be created fails boot at once, naming the cause** (`internal/mq/embedded.go` (+ tests)): part of [#617](https://github.com/Wave-RF/WaveHouse/issues/617). A regular file at `/nats`, or a `nats` directory that could not be created there, failed JetStream in the background, so boot waited out the server's 5s readiness check and reported only `nats server not ready`. `NewEmbedded` now creates the directory first (at `0700`, as the server does) and refuses boot with the mkdir error. An existing but unwritable `nats` directory still takes the old path. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 7aede5bad..678c52a3f 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -114,7 +114,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on — one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)) — each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope by their raw names, which the cache escapes where it builds a key. `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The snapshot is taken before any input a bump invalidates is chosen, the tenant's connection included: a reload that moves the tenant to another address or database runs `Pools.Reconcile` and then `InvalidateTenant` (a repoint that keeps both, such as a username or `tls` change, reads the same tables and bumps nothing), so a request that took the old pool files the old database's rows under a version that bump orphans, whether its `Set` lands before the bump or after. The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. -- **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: one version per tenant, per (tenant, table) and per (tenant, table, scope), each keyed by its name alone and bumped in place, so the index holds one entry per live tenant, table and scope however often each is bumped ([#262](https://github.com/Wave-RF/WaveHouse/issues/262)). A query key folds the tenant's version and, for each dependency, its tenant's, table's and scope's, so bumping a table (a scopeless write) orphans every scope of it, and bumping one scope orphans that scope and the whole-table view — scope is reserved and empty today, so every write is the whole-table bump — all without touching the pool. Every field — the caller's query key, the tenant id, and each dependency's table and scope — is escaped and joined by `internal/keyenc` where the key is built, so a dot, a space or a `%` in a name is never read as a separator: each dependency renders as `..
.
..`, and the whole entry key is `|.||…`. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results. A tenant's version is a *generation*, unique within the process and handed out by the first key built for the tenant; `BumpTenant` (behind `InvalidateTenant`) drops the tenant's whole index, so the next key gets a fresh generation no cached entry folds, orphaning every cached result of the tenant in one step — a pipe result with no dependencies, and a table no bump ever keyed, included — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). `LocalCache.Prune` does the same for every tenant no longer served, which `internal/app` runs after each settings reload, so a tenant removed or rejected stops holding its index. A table bump drops the table's scope versions with it, since every key they were folded into also folds the old table version; and any bump (of a table, a scope or the tenant) under a tenant with no index is a no-op that records nothing, since the next key built for it gets a fresh generation no cached entry folds — so neither the `sharedTables` fan-out nor an insert still in flight for a tenant just pruned brings its index back. The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. +- **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: one version per tenant, per (tenant, table) and per (tenant, table, scope), each keyed by its name alone and bumped in place, so the index holds one entry per live tenant, table and scope however often each is bumped ([#262](https://github.com/Wave-RF/WaveHouse/issues/262)). An entry's key (`QueryKey`) folds the tenant's version and, for each dependency, its tenant's, table's and scope's, so bumping a table (a scopeless write) orphans every scope of it, and bumping one scope orphans that scope and the whole-table view — scope is reserved and empty today, so every write is the whole-table bump — all without touching the pool. Every field — the caller's query key, the tenant id, and each dependency's table and scope — is escaped and joined by `internal/keyenc` where the key is built, so a dot, a space or a `%` in a name is never read as a separator: each dependency renders as `..
.
..`, and the whole entry key is `|.||…`. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results. A tenant's version is a *generation*, unique within the process and handed out by the first key built for the tenant; `BumpTenant` (behind `InvalidateTenant`) drops the tenant's whole index, so the next key gets a fresh generation no cached entry folds, orphaning every cached result of the tenant in one step — a pipe result with no dependencies, and a table no bump ever keyed, included — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). `LocalCache.Prune` does the same for every tenant no longer served, which `internal/app` runs after each settings reload, so a tenant removed or rejected stops holding its index. A table bump drops the table's scope versions with it, since every key they were folded into also folds the old table version; and any bump (of a table, a scope or the tenant) under a tenant with no index is a no-op that records nothing, since the next key built for it gets a fresh generation no cached entry folds — so neither the `sharedTables` fan-out nor an insert still in flight for a tenant just pruned brings its index back. The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. ### `config/` — Configuration diff --git a/internal/cache/version_manager_test.go b/internal/cache/version_manager_test.go index 847920542..079453060 100644 --- a/internal/cache/version_manager_test.go +++ b/internal/cache/version_manager_test.go @@ -164,10 +164,26 @@ func TestVersionManager_GenerationsNeverRepeat(t *testing.T) { func TestVersionManager_BumpWithoutIndex(t *testing.T) { t.Parallel() vm := NewVersionManager() + vm.BumpTable("acme", "users") + assert.Zero(t, vm.size(), "a table bump for a tenant with no index creates nothing") + vm.BumpNamespace(Namespace{Tenant: "acme", Table: "users", Scope: "org_1"}) + assert.Zero(t, vm.size(), "a namespace bump for a tenant with no index creates nothing") + vm.BumpTenant("acme") - assert.Zero(t, vm.size()) + assert.Zero(t, vm.size(), "bumping a tenant with no index is a no-op") + + // An insert still in flight for a tenant just pruned must not bring its + // index back: a write racing the prune sees the tenant gone and bumps + // blind, same as above. + vm.QueryKey("acme", "h", nil) + vm.Prune(func(tenant.ID) bool { return false }) + assert.Zero(t, vm.size(), "prune released the tenant's index") + + vm.BumpTable("acme", "users") + vm.BumpNamespace(Namespace{Tenant: "acme", Table: "users", Scope: "org_1"}) + assert.Zero(t, vm.size(), "a bump for a tenant just pruned must not recreate its index") } func TestVersionManager_Prune(t *testing.T) { From 1febf32a49c6d6a31b5714aa3bc28ddb0d6cbf3e Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:34:44 -0400 Subject: [PATCH 58/79] docs(cache): fix entry-vs-caller query key terms and cachetest coverage Review round found "query key" naming both the caller's cache key and the version-folded entry key it seeds; disambiguate in architecture.md and AGENTS.md, untangle a muddled deployment.md paragraph on which cached results a tenant move drops, document the cachetest conformance suite's Options in architecture.md (with a pointer from development.md's tree), and fix a CHANGELOG.md "pool" that meant the cache's capacity. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- AGENTS.md | 2 +- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 3 ++- docs/src/content/docs/deployment.md | 2 +- docs/src/content/docs/development.md | 2 +- 5 files changed, 6 insertions(+), 5 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 12805f8b8..5a55f6e1b 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -31,7 +31,7 @@ Twenty internal packages under `internal/` (plus `internal/testutil/` for shared - **`api/`** — Chi HTTP router, JWT/JWKS middleware (from `auth/`), ingest/query/structured-query/SSE/schema/DLQ/pipes handlers; `ch_errors.go` (`writeCHError`) is the one mapping from a failed ClickHouse query to status, `code` and `retryable` - **`app/`** — the process wiring: `New` builds every component from the boot config and the settings directory (each one wired in one place — what it opens, what it loops, what it releases — with the settings registry handed to its wiring function whole, the injection point of the per-tenant registry of #583: store-keyed getters for the handlers, `perTenant` for the async paths (with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker), the `chconn.Pools` and the per-tenant `discoveries` reconciled from `AfterAdopt`, `shortestKeepalive` for the one setting folded over every tenant served, `gapWindows` handing the sweeper each tenant's own gap window (a rejected tenant's as its folder last had it, unbounded for one rejected since boot) and the `mq.max_bytes_gb` reconcile each served tenant's byte budget, and `defaultPolicy` for the one setting that still follows tenant `0`, a flat directory's ops-gate admin role; the auth verifiers are per tenant, reconfigured (rebuilt only on changed wiring) and pruned from `AfterAdopt`, and the same hook's `Hub.Prune` ends the open streams of a tenant no longer served), `Run` drives the long-lived ones under one `errgroup` until the context is cancelled or one fails, `Close` releases them in reverse order. `New` wires only what the process's `roles` need (discovery, dedupe, auth verifiers, the hub bridge and keepalive per API process; the ingest worker per ingest process; the sweeper under its lease through `elected`); a process without `api` serves `api.NewOpsRouter` — probes, `/version`, metrics, and the settings reload behind the operator key alone. `cmd/wavehouse` and `tests/integration` both boot through it - **`auth/`** — JWT auth middleware: HMAC **or** JWKS verification with `alg` pinned to the active verifier, role extraction from a configurable claim path; always runs, never rejects (bad token → empty role + stashed reason). One verifier per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 9): `Authenticator` keys them by `tenant.ID` — the request store's `settings.Store.Tenant()`, through an injected `TenantSource`; `tenant.Default` on the tenant-exempt routes — built from each tenant's `auth` block by `Reconfigure`, dropped by `Prune` once the tenant stops being served (rejected or removed), released by `Close`; the secrets (`Config`) are boot-level and shared. A JWKS key set is fetched off the boot and reload paths: until one has been stored the verifier is pending and a token-bearing request gets `503` + `Retry-After` from `api.refuseUnverifiable` (`auth.ErrVerifierPending`), never a `default_role` evaluation; refresh is library-managed (Eric, 2026-09-22), response capped at 1 MiB; the operator key's admin role is the request tenant's -- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for the caller's query key and its singleflight (escaped whole as the lead field of the stored key, `|.||…`), `..
.
.` for a namespace, each field escaped by `keyenc` where the key is built (a `Namespace` carries the raw table and scope, so no caller escapes) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version folded into every namespace and query key of one tenant, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read, taken before the handler chooses any input a bump invalidates — the tenant's connection included — and `Set` files the fill under it, so a write landing mid-query, or a reload moving the tenant to another address or database after the request took its connection, orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) +- **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for the caller's query key and its singleflight (escaped whole as the lead field of the stored key, `|.||…`), `..
.
.` for a namespace, each field escaped by `keyenc` where the key is built (a `Namespace` carries the raw table and scope, so no caller escapes) — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version folded into every namespace key and entry key of one tenant, orphaning its every cached result in one step, pipe results included (no insert reaches a pipe result until [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); `Lookup` returns a `Snapshot` of the versions it read, taken before the handler chooses any input a bump invalidates — the tenant's connection included — and `Set` files the fill under it, so a write landing mid-query, or a reload moving the tenant to another address or database after the request took its connection, orphans the fill ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)), and every backend runs the conformance suite `internal/testutil/cachetest`; the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the whole cache of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config. `Classify` (`errclass.go`) says what a failed ClickHouse request means for the request — `Unavailable`, `Denied`, `Rejected` (any unlisted exception code: the server read it and refused it), or `Unknown` (no code, no recognizable transport failure) — over the driver's error types and the HTTP interface's `HTTPError`; the ingest worker and the query handlers (`api/ch_errors.go` `writeCHError`, [#403](https://github.com/Wave-RF/WaveHouse/issues/403), [#271](https://github.com/Wave-RF/WaveHouse/issues/271)) both use it - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) - **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability when a selected backend keeps state there (`NeedsDataDir`); `backends.go` holds each layer's `.backend` (only the in-process value today); `config.go` holds `roles` (`Has(Role)`) and `instance_id`, and `Validate` refuses a role split the backends cannot serve (any split over the embedded MQ; `api` without `ingest`, or the reverse, over a local cache) — boot is the validator, there is no dry run diff --git a/CHANGELOG.md b/CHANGELOG.md index 6baafa281..8492d1576 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -88,7 +88,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **A failed ClickHouse query answers by what went wrong, not a flat `500`/`502`** (`internal/api/ch_errors.go` (new, + tests), `internal/api/{errors,query,structured_query,pipes,schema,ch_settings}.go`, `internal/chconn/errclass.go` (`HTTPStatus` exported), `clients/ts/src/errors.ts` (+ tests), `tests/integration/query_errors_test.go` (new), `tests/integration/query_limits_test.go`, `internal/app/app_test.go`, `tests/e2e/sdk/{admin,query}.test.ts`, `AGENTS.md`, `docs/src/content/docs/{api,architecture}.md`, `docs/src/content/docs/{access-control,configuration}.mdx`, `docs/src/content/docs/sdk/{reference.md,index.mdx}`): fixes [#403](https://github.com/Wave-RF/WaveHouse/issues/403) and [#271](https://github.com/Wave-RF/WaveHouse/issues/271), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). ClickHouse answers a syntax error, a missing grant and an overloaded server alike with HTTP `500`, so `/v1/ops/query` turned a bad statement into a `502` and `/v1/query` and pipes into a `500` the SDK retried. All three now class the failure with `chconn.Classify` through one helper, `writeCHError`: a statement ClickHouse refused is `400 clickhouse.rejected`; a query over a rows/bytes limit, the role's own memory cap, or its time cap where that is no longer than `query_timeout` is `400 clickhouse.limit_exceeded`; `ACCESS_DENIED` is `403 clickhouse.access_denied`; credentials, user or database refused, or a redirect or `4xx` with no exception code from whatever fronts ClickHouse, is `502 clickhouse.misconfigured`; ClickHouse down, unreachable or overloaded is `503 clickhouse.unavailable` with `Retry-After: 5`; a failure with no verdict stays `500` (`502` on the proxy) as `clickhouse.unknown`. The error envelope gains `code` and `retryable` next to `error` on these responses — additive. A role with `max_execution_time` now queries with no context deadline and a cancel two seconds past the cap instead: clickhouse-go overwrote the cap's `max_execution_time` with deadline+5s for any deadline over 1s, so an overrun came back as a bare deadline, indistinguishable from waiting for a pooled connection; ClickHouse now enforces the cap itself and reports `TIMEOUT_EXCEEDED`. `POST /v1/ops/schema/refresh` against an unreachable ClickHouse is a `503` with `Retry-After` instead of a `500`. **SDK:** `WaveHouseError.code` and `retryable` now take the server's `code`/`retryable` when the body has them (`HTTP_` and "5xx retries" otherwise), so a rejected query is `clickhouse.rejected` rather than `HTTP_500`, and is not retried. - **An unavailable ClickHouse is retried with backoff instead of dead-lettering every row** (`internal/chconn/errclass.go` (new, + tests), `internal/ingest/{worker,backoff}.go` (`backoff.go` new, + tests), `internal/mq/{mq,embedded}.go`, `internal/testutil/mocks.go`, `tests/integration/ingest_outage_test.go` (new), `AGENTS.md`, `README.md`, `docs/src/content/docs/{ingest-pipeline,architecture,api,deployment,why-wavehouse}.md`, `docs/src/content/docs/{settings-directory,index,access-control}.mdx`): workstream A of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). A failed batch insert used to go through row-by-row isolation whatever the failure, so a ClickHouse that was down, overloaded or read-only failed every row twice and parked the whole batch on the DLQ. `chconn.Classify` now classes the failure first — `Rejected` (any ClickHouse exception code outside the availability and credential lists: the server read the row and refused it), `Unavailable` (connection refused/reset, timeouts, `TOO_MANY_SIMULTANEOUS_QUERIES`, `SERVER_OVERLOADED`, `MEMORY_LIMIT_EXCEEDED`, `TOO_MANY_PARTS`, `READONLY`, `TABLE_IS_READ_ONLY`, `KEEPER_EXCEPTION`, …), `Denied` (`AUTHENTICATION_FAILED`, `ACCESS_DENIED`, …) or `Unknown` (no code, no recognizable transport failure). Only `Rejected` is isolated and dead-lettered as before, and a multi-row batch refused with `TOO_MANY_PARTS` or `MEMORY_LIMIT_EXCEEDED` is split row by row first (`chconn.Splittable`), because a batch spanning too many partitions or too much memory can fail where each of its rows inserts; every other class hands the batch back to the queue with a delayed nak (`mq.Message.NakWithDelay`, new) under a jittered 1 s → 30 s backoff shared by every table on the same ClickHouse pool (a failure of one table — read-only, too many parts or mutations, a grant missing on it, `chconn.TableScoped` — backs off that table alone), which turns rows away without a request while it runs and probes once per window, and ClickHouse going away mid-isolation stops isolation and retries the rows it had not settled. Counted by the new `wavehouse_ingest_retries_total{table, reason}`; logged at `WARN` when an outage starts and at most every 30 s during it. A long outage now shows as a growing ingest stream and, at `mq.max_bytes_gb`, ingest `503`s — not as a full DLQ; a lasting failure of one table holds back its tenant's other tables once its waiting rows reach `maxAckPending`. Retried rows come back out of arrival order, which matters only to a `ReplacingMergeTree` without a version column or a `CollapsingMergeTree`. - **Schema discovery's retry loop jitters its backoff** (`internal/discovery/discovery.go` (+ tests), `internal/app/wire.go`, `internal/api/errors.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`): `RetryRefresh` slept exactly `2s * 2^n` capped at 60s, so instances retrying against one recovering ClickHouse fired in lockstep, every 60s on the same second. Each sleep is now drawn uniformly from below the backoff (full jitter), spreading the retries over the whole window and halving the mean wait — so a failing tenant's retries, their log lines and `wavehouse_schema_refresh_failures_total` come about twice as often ([#141](https://github.com/Wave-RF/WaveHouse/issues/141)). -- **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/query/ident.go` (removed, + tests), `internal/app/wire.go`, `docs/src/content/docs/{api,architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL (a pipe result's key folded no version, so pipes were unaffected; with the tenant's version in every key they now take the same snapshot). The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. The snapshot is taken before any input a bump invalidates is chosen, the tenant's ClickHouse connection included: both handlers now look up before they resolve the tenant's pool, so a reload that moves the tenant to another address or database after a request took the old pool orphans that request's fill instead of filing the old database's rows as fresh under the new tenant version. A tenant on no pool is still a `503` before a cached result is served or a query runs. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the pool, a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. The cache now escapes table and scope names itself: a `Namespace` carries them raw and the version index builds its keys with `internal/keyenc`, so neither the structured-query read nor the ingest worker's invalidation escapes them (`query.SafeEncodeToken` is gone) and no name reaches a key unescaped. The suite pins that a name holding a dot, a space or a `%` is read and bumped under one key, and that names which would run together unescaped (`a.0.b` against `a` with scope `b.0.`) stay two entries. Namespace keys are unchanged; an entry's key now carries the tenant's version and escapes the caller's query key whole. All of them live in the process, so nothing stored is orphaned. +- **A write that lands while a cached read is running no longer re-homes the pre-write rows under the post-write key** (`internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/cachetest` (new), `internal/api/{structured_query,pipes}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/query/ident.go` (removed, + tests), `internal/app/wire.go`, `docs/src/content/docs/{api,architecture,deployment}.md`, `AGENTS.md`): fixes [#382](https://github.com/Wave-RF/WaveHouse/issues/382), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `POST /v1/query` rebuilt the version-folded cache key after the query ran, so an insert invalidating the table mid-query filed the rows read before it under the new versions, and they were served as fresh until their TTL (a pipe result's key folded no version, so pipes were unaffected; with the tenant's version in every key they now take the same snapshot). The `Cache` interface now snapshots at lookup: `Lookup(ctx, tenant, sha, deps)` returns the `Entry` and a `Snapshot` of the versions it read, and `Set(ctx, snapshot, value, ttl)` stores under that snapshot, so such a fill is orphaned and the next request reads the post-write rows. The singleflight leader's snapshot is the one used; coalescing is unchanged. The snapshot is taken before any input a bump invalidates is chosen, the tenant's ClickHouse connection included: both handlers now look up before they resolve the tenant's pool, so a reload that moves the tenant to another address or database after a request took the old pool orphans that request's fill instead of filing the old database's rows as fresh under the new tenant version. A tenant on no pool is still a `503` before a cached result is served or a query runs. A `Lookup` whose dependencies name another tenant is refused (`ErrForeignDependency`). `Set` now errors only when the backend failed: a value the cache declines (larger than the cache holds (`cache.l1_max_cost`), a non-positive TTL) is not an error. One behavior change: the tenant's version is folded into every key, a pipe result's included, so `InvalidateTenant` (a tenant back on a pool after an absence, or moved to another ClickHouse address or database) now drops that tenant's cached pipe results as well as its query results; before, a pipe result stayed until its TTL. Inserts still do not invalidate pipe results ([#343](https://github.com/Wave-RF/WaveHouse/pull/343)). A backend-agnostic conformance suite, `cachetest.Run`, pins what a hit, a miss and each kind of bump mean, and `LocalCache` runs it; the Redis-compatible backend will run the same suite. The cache now escapes table and scope names itself: a `Namespace` carries them raw and the version index builds its keys with `internal/keyenc`, so neither the structured-query read nor the ingest worker's invalidation escapes them (`query.SafeEncodeToken` is gone) and no name reaches a key unescaped. The suite pins that a name holding a dot, a space or a `%` is read and bumped under one key, and that names which would run together unescaped (`a.0.b` against `a` with scope `b.0.`) stay two entries. Namespace keys are unchanged; an entry's key now carries the tenant's version and escapes the caller's query key whole. All of them live in the process, so nothing stored is orphaned. - **An explicit `false`, `0` or `""` in `config.yaml` is no longer replaced by the key's default** (`internal/config/config.go`, `internal/config/defaults_test.go` (new), `docs/src/content/docs/configuration.mdx`, `AGENTS.md`): [#631](https://github.com/Wave-RF/WaveHouse/issues/631). Defaults lived in cleanenv `env-default` tags, which cleanenv applies after the YAML decode to any field still at its zero value, so it could not tell a key the file set to its zero value from one the file left out. `otel.traces.enabled: false`, `otel.metrics.enabled: false` and `otel.logs.enabled: false` came back `true`; `otel.traces.sample_rate: 0` and `otel.logs.sample_rate: 0` came back `1.0`; `server.shutdown_timeout: 0` came back `10`; `cache.l1_max_cost: 0`, `prometheus.path: ""` and `data_dir: ""` came back as their defaults; `server.port: 0` came back `8080`. All of it was silent. Defaults now live in one Go function, `defaults()`, which `Load` starts from before decoding the file and then applying `WH_*` variables, so the order is env > YAML > default and a key the file sets always wins. **Behaviour change if your file relied on the bug:** a zero you wrote now takes effect. A file that says `sample_rate: 0` now exports no traces (or no DEBUG/INFO logs), where it silently exported everything; a signal set `enabled: false` is now off; `shutdown_timeout: 0` now skips the drain. `cache.l1_max_cost: 0`, `server.port: 0`, and `data_dir: ""` now refuse boot (`cache init: MaxCost can't be zero`, `server.port 0 out of range`, `data_dir (WH_DATA_DIR) is required`) instead of running on the default; an empty `prometheus.path` refuses boot when `prometheus.enabled` is true. Delete the key to get the default back. Env vars are unchanged: they already honoured an explicit zero. New tests load through `config.Load` for every affected key (a YAML zero is kept, an absent key gets the default, env wins in both directions), refuse an `env-default` tag on any field, and pin each documented default in `configuration.mdx` to `defaults()`. - **An embedded queue store that cannot be created fails boot at once, naming the cause** (`internal/mq/embedded.go` (+ tests)): part of [#617](https://github.com/Wave-RF/WaveHouse/issues/617). A regular file at `/nats`, or a `nats` directory that could not be created there, failed JetStream in the background, so boot waited out the server's 5s readiness check and reported only `nats server not ready`. `NewEmbedded` now creates the directory first (at `0700`, as the server does) and refuses boot with the mkdir error. An existing but unwritable `nats` directory still takes the old path. - **An insert invalidates a table's cached results under every tenant the directory holds** (`internal/app/wire.go` (+ tests), `internal/settings/registry.go` (+ tests), `AGENTS.md`, `docs/src/content/docs/{deployment,architecture,ingest-pipeline}.md`): until [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6 gives each tenant its own ClickHouse, every tenant reads the same tables, but the ingest worker — which writes every event as tenant `0`'s until story 5 — bumped only tenant `0`'s cache namespaces after an insert, so another tenant's cached query could answer stale rows for up to its TTL (an hour at most). The cache the worker invalidates through now fans each bumped namespace out to every tenant the registry knows (the new `Registry.Known`), the named one and a rejected one included — a rejected tenant comes back into service with the entries it has, so leaving it out would let a folder repaired inside a TTL serve pre-insert rows; reads are untouched, so a tenant is still never served another's cached rows. The residual, a folder removed and restored inside a TTL, is closed since #610: a tenant back on a pool after an absence has its cached results orphaned at once (`Cache.InvalidateTenant`). Raised by CodeRabbit on #602. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 202111750..5672ddd16 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -114,7 +114,8 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on — one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)) — each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope by their raw names, which the cache escapes where it builds a key. `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The snapshot is taken before any input a bump invalidates is chosen, the tenant's connection included: a reload that moves the tenant to another address or database runs `Pools.Reconcile` and then `InvalidateTenant` (a repoint that keeps both, such as a username or `tls` change, reads the same tables and bumps nothing), so a request that took the old pool files the old database's rows under a version that bump orphans, whether its `Set` lands before the bump or after. The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. -- **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: a namespace key is `..
.
.`, its fields joined with `keyenc.Join` so a dot in a table or scope name is escaped rather than read as a separator, and an entry's key, `|.||…` with the caller's query key escaped whole (its `:` become `%3A`), folds in the tenant's version and each dependency's namespace key and namespace version, so bumping a table (a scopeless write) or one scope — scope is reserved and empty today, so every write is the whole-table bump — orphans every dependent entry without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results, and the flat directory's single tenant simply carries the `0` prefix. `BumpTenant` (behind `InvalidateTenant`) advances the tenant version that leads every namespace key of one tenant, orphaning its every namespace and every cached query in one step (the tenant version is folded into every query key, so a pipe result with no dependencies is orphaned too) — a table no bump ever keyed included, which an enumeration of the index would miss — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. +- **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: a namespace key is `..
.
.`, its fields joined with `keyenc.Join` so a dot in a table or scope name is escaped rather than read as a separator, and an entry's key, `|.||…` with the caller's query key escaped whole (its `:` become `%3A`), folds in the tenant's version and each dependency's namespace key and namespace version, so bumping a table (a scopeless write) or one scope — scope is reserved and empty today, so every write is the whole-table bump — orphans every dependent entry without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results, and the flat directory's single tenant simply carries the `0` prefix. `BumpTenant` (behind `InvalidateTenant`) advances the tenant version that leads every namespace key of one tenant, orphaning its every namespace and every cached query in one step (the tenant version is folded into every entry's key, so a pipe result with no dependencies is orphaned too) — a table no bump ever keyed included, which an enumeration of the index would miss — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. +- **cachetest** (`internal/testutil/cachetest`) — `Run(t, factory, Options)`, the conformance suite every `Cache` backend runs, each case on a fresh cache from the factory: what a hit, a miss and each kind of bump mean, independent of where entries and versions live. `Options` describes what a backend can do beyond the `Cache` contract, and leaving one unset skips the cases it enables: `MaxValueBytes` opens the oversize case; `NewPair` returns two instances over one shared store, for the cross-instance cases a shared backend must run; `Entries` counts a cache's entries, and without it the zero-snapshot case — no `Lookup` reads the key a zero `Snapshot` would land under, so only a count shows that one stored nothing — is skipped. ### `config/` — Configuration diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index f36b37ac8..08f538c65 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -409,7 +409,7 @@ The folder name is the tenant id, and each folder is a complete settings directo **The admin routes take the operator key only.** `/v1/ops/*` reaches every tenant, so over a nested directory no tenant's admin role opens it: the [operator key](/api#authentication) alone does, and a token carrying an admin role gets `403`. Boot a nested directory without `auth.operator_key` and no caller can reach these routes at all, which leaves `SIGHUP` as the only reload; the server warns about it at boot. `GET /v1/ops/pipes`, `GET /v1/ops/pipes/{name}`, `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the same `?tenant=`, and address tenant `0` without it; `GET /v1/ops/dlq/stats` takes it too, and reads a rejected or removed tenant's dead-letter queue like a served one's, since the queue is kept; a tenant that has none is a `404`. On the routes that take it the parameter is parsed strictly — a query string that does not parse, an empty or repeated `tenant`, or a malformed id is a `400`, never a silent read of the default tenant or, on the reload route, a reload of every tenant. The SDK sends it as the [`tenant` option](/sdk/admin#settings--whsettings). -**What a tenant's folder decides.** A request is evaluated against its own tenant's `policies.json` and `pipes.json` (ingest, structured queries, pipes), its `query.*` keys, its `cors.allowed_origins`, and its `dedupe` block: whether its records are deduplicated, by which id, against the tenant's own store, which that folder's `dedupe.enabled` opens and closes on reload exactly as [the single-tenant one](/settings-directory#deduplication) does (every tenant's store is a share of the one Pebble instance at `/pebble`, each key led by its tenant), so `wavehouse_ingest_dedupe_disabled_total` ticks only across a tenant's own reload, whatever the other tenants' switches say. A tenant's seen ids are its own: the same event id is first seen under each tenant that sends it. Its `auth` block is its own too: each tenant's folder wires that tenant's token verifier (`jwks_url`, `role_claim`), built when the folder is adopted and rebuilt when its wiring changes, so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its provider's key set. Under another tenant's header a token is treated as invalid, and the request falls back to that tenant's `default_role` like any other unverifiable token, possibly after a rate-limited key refetch (see [Authentication](/settings-directory#authentication)). Keep `X-Tenant-ID` pinned at the proxy so a token is never presented under the wrong tenant. Tenants can still accept each other's tokens: those that leave `jwks_url` empty share the boot HMAC secret when `auth.jwt_secret` is set, so a token verifies under any of them (with no secret they validate no token at all), and those whose `jwks_url` names the same key set accept each other's tokens; isolate them by provider, or scope rows by a signed claim ([row-level security](/access-control#row-level-security)). A tenant whose `jwks_url` has not been fetched yet answers `503` with `Retry-After` to its token-bearing requests alone. A tenant that stops being served — its folder rejected or removed — loses its verifier and the JWKS refresh with it, and gets a fresh one when its folder is adopted again. The HMAC secret and the operator key stay boot config, shared by every tenant; the operator key is stamped with the request tenant's `admin_role`. A tenant's `clickhouse` and `schema` blocks are its own as well: each tenant reads and writes its own ClickHouse — one native pool per distinct address, database, user, password and `tls` tuple, shared by the tenants naming it, under the process-wide [connection ceiling](/settings-directory#clickhouse) — and discovers its own tables from its own database on its own `schema.refresh_interval`. Its message queue is its own as well: its events are queued on a stream of their own, capped at its own `mq.max_bytes_gb` — at that budget its ingest answers `503` while every other tenant's keeps publishing — beside a dead-letter stream of its own at a tenth of it, and the history that gap-fill replays from it is kept for its own `stream.gap_window_minutes`. Nothing checks what the tenants' budgets add up to against the disk, so size them together ([Message Queue](/settings-directory#message-queue)). An event is published on its tenant's subject (`ingest.{tenant}.{table}`), so a `GET /v1/stream` connection is authorized by its own tenant's `policies.json` and receives its own tenant's rows alone, the ingest worker inserts a row into its own tenant's ClickHouse, a rejected row is parked under its own tenant's `dlq.enabled` and subject (`dlq.{tenant}.{table}`), and two tenants' tables of one name never share a batch. The query cache is one pool, but its entries are keyed by tenant: identical `POST /v1/query` and pipe requests from two tenants are two entries and two queries to ClickHouse, and a tenant is never served another's cached rows. An insert invalidates the table's cached results under every tenant on the same ClickHouse address and database as the tenant it was ingested for, whatever their user or `tls` block, since they read the same tables; a tenant on no pool — its folder rejected or removed, or no pool could be opened for it, such as by the ceiling — is out of that fan-out while it is, and has its cached `POST /v1/query` results dropped the moment it is back on one, so a repaired or restored folder never serves query rows cached before the inserts it missed, and so does a tenant whose folder moves it to another address or database, whose cached rows came from other tables. Both also drop the tenant's cached pipe results; apart from them a cached pipe result stays until its TTL expires, since a pipe names no table and no insert invalidates it. One setting weighs every tenant: the SSE keepalive, where the wheel runs at the shortest `stream.keepalive_interval` among the tenants being served, with that tenant's `stream.keepalive_buckets`. +**What a tenant's folder decides.** A request is evaluated against its own tenant's `policies.json` and `pipes.json` (ingest, structured queries, pipes), its `query.*` keys, its `cors.allowed_origins`, and its `dedupe` block: whether its records are deduplicated, by which id, against the tenant's own store, which that folder's `dedupe.enabled` opens and closes on reload exactly as [the single-tenant one](/settings-directory#deduplication) does (every tenant's store is a share of the one Pebble instance at `/pebble`, each key led by its tenant), so `wavehouse_ingest_dedupe_disabled_total` ticks only across a tenant's own reload, whatever the other tenants' switches say. A tenant's seen ids are its own: the same event id is first seen under each tenant that sends it. Its `auth` block is its own too: each tenant's folder wires that tenant's token verifier (`jwks_url`, `role_claim`), built when the folder is adopted and rebuilt when its wiring changes, so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its provider's key set. Under another tenant's header a token is treated as invalid, and the request falls back to that tenant's `default_role` like any other unverifiable token, possibly after a rate-limited key refetch (see [Authentication](/settings-directory#authentication)). Keep `X-Tenant-ID` pinned at the proxy so a token is never presented under the wrong tenant. Tenants can still accept each other's tokens: those that leave `jwks_url` empty share the boot HMAC secret when `auth.jwt_secret` is set, so a token verifies under any of them (with no secret they validate no token at all), and those whose `jwks_url` names the same key set accept each other's tokens; isolate them by provider, or scope rows by a signed claim ([row-level security](/access-control#row-level-security)). A tenant whose `jwks_url` has not been fetched yet answers `503` with `Retry-After` to its token-bearing requests alone. A tenant that stops being served — its folder rejected or removed — loses its verifier and the JWKS refresh with it, and gets a fresh one when its folder is adopted again. The HMAC secret and the operator key stay boot config, shared by every tenant; the operator key is stamped with the request tenant's `admin_role`. A tenant's `clickhouse` and `schema` blocks are its own as well: each tenant reads and writes its own ClickHouse — one native pool per distinct address, database, user, password and `tls` tuple, shared by the tenants naming it, under the process-wide [connection ceiling](/settings-directory#clickhouse) — and discovers its own tables from its own database on its own `schema.refresh_interval`. Its message queue is its own as well: its events are queued on a stream of their own, capped at its own `mq.max_bytes_gb` — at that budget its ingest answers `503` while every other tenant's keeps publishing — beside a dead-letter stream of its own at a tenth of it, and the history that gap-fill replays from it is kept for its own `stream.gap_window_minutes`. Nothing checks what the tenants' budgets add up to against the disk, so size them together ([Message Queue](/settings-directory#message-queue)). An event is published on its tenant's subject (`ingest.{tenant}.{table}`), so a `GET /v1/stream` connection is authorized by its own tenant's `policies.json` and receives its own tenant's rows alone, the ingest worker inserts a row into its own tenant's ClickHouse, a rejected row is parked under its own tenant's `dlq.enabled` and subject (`dlq.{tenant}.{table}`), and two tenants' tables of one name never share a batch. The query cache is one pool, but its entries are keyed by tenant: identical `POST /v1/query` and pipe requests from two tenants are two entries and two queries to ClickHouse, and a tenant is never served another's cached rows. An insert invalidates the table's cached results under every tenant on the same ClickHouse address and database as the tenant it was ingested for, whatever their user or `tls` block, since they read the same tables; a tenant on no pool — its folder rejected or removed, or no pool could be opened for it, such as by the ceiling — is out of that fan-out while it is, and has its cached results (`POST /v1/query` and pipe alike) dropped the moment it is back on one, so a repaired or restored folder never serves rows cached before the inserts it missed. A tenant whose folder moves it to another address or database has them dropped too, since they came from other tables. Apart from those two drops, a cached pipe result stays until its TTL expires, since a pipe names no table and no insert invalidates it. One setting weighs every tenant: the SSE keepalive, where the wheel runs at the shortest `stream.keepalive_interval` among the tenants being served, with that tenant's `stream.keepalive_buckets`. **What a lost tenant `0` costs.** A `0` folder that a reload rejects or removes stops tenant `0` being served like any other, and what becomes of the shared settings depends on how they are read. Tenant `0` leaves its ClickHouse pool (closed only once no served tenant names its tuple), and its schema registry and verifier are released with the folder, like any other tenant's; the `/v1/ops/*` routes, which resolve no tenant, verify against it, so a token there reads as invalid (`401`) rather than merely non-admin (`403`) until tenant `0` is served again — the operator key, which never consults a verifier, is unaffected. CORS does not stay either: the responses that read tenant `0`'s list — the tenant-exempt routes, the refusals, a preflight naming no tenant — carry no CORS headers until the folder is served again, while every other tenant's routes keep their own list. Tenant `0`'s own dedupe store closes, as any rejected or removed tenant's does, its seen ids kept for the folder that restores it. What is read per event follows the event's tenant, so tenant `0`'s events are the ones affected: with no ClickHouse to insert into, its rows fail and are parked on the DLQ whatever its switch said, and its open `GET /v1/stream` connections are ended, as any tenant's are when it stops being served — the other tenants' events are untouched. A nested directory that has never served a tenant `0` — no `0` folder, or one rejected at boot — serves every other tenant from its own ClickHouse. Outside `/v1/ops/*`, a `/v1` request that sends no `X-Tenant-ID` resolves to tenant `0`, so with no `0` folder it answers `404 unknown tenant: 0` (`503` with a rejected one) — the SDK's `/v1/health` reachability ping included. diff --git a/docs/src/content/docs/development.md b/docs/src/content/docs/development.md index 04aac7490..4d91f6710 100644 --- a/docs/src/content/docs/development.md +++ b/docs/src/content/docs/development.md @@ -471,7 +471,7 @@ WaveHouse/ │ ├── settings/ # Settings directory: validate, adopted snapshot, reload │ ├── stream/ # SSE fan-out: Hub, Subscriber queue, Bucket, keepalive wheel │ ├── tenant/ # Tenant id: type, grammar, reserved default, request header name -│ └── testutil/ # Shared test helpers and mocks +│ └── testutil/ # Shared test helpers and mocks (cachetest suite) ├── tests/ # Integration & E2E tests │ ├── integration/ # Go integration tests (//go:build integration) │ └── e2e/ # E2E suite (orchestrator + ClickHouse testcontainer) From 5a72724b83afe4b14bcf9f0f43f66ccd105f9048 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:40:39 -0400 Subject: [PATCH 59/79] =?UTF-8?q?fix(cache):=20review=20round=202=20?= =?UTF-8?q?=E2=80=94=20stale=20docs=20claims,=20redisConfig=20field=20pin?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - configuration.mdx: isAuthError also matches NOPERM (an ACL user missing a connection command); the boot-error-logging line only named WRONGPASS/NOAUTH. - CHANGELOG.md: the E4 entry claimed the e2e coverage exclude for the cache files was gone; .testcoverage.yml still excludes internal/cache/pending.go, internal/config/cache_redis.go and internal/cache/(local|version_manager).go from e2e, for reasons the entry now states. Also added the docs/.github files this PR touched that the entry's file list was missing. - pipes.mdx, api.md: "L1" claims that are wrong once cache.backend is redis (there is no L1) — reworded to "the query cache"/"cache". - deployment.md: "Persistence is not needed" was incomplete — a restart *with* persistence reloads a stale snapshot and can serve invalidated results as hits until their TTL. Now says to run without persistence, and that restoring a snapshot is a rollback. - development.md: note that shared_cache_test.go starts its own Redis per test and boots extra cache.backend=redis instances over the integration suite's ClickHouse. - app_test.go: TestRedisConfig_FromLoadedDefaults left Username, DB and TLS at their zero value on both sides of its assert.Equal, so deleting any of their three mapping lines in redisConfig wouldn't fail it. Added TestRedisConfig_UsernameDBTLSMapped, which drives all three to a non-zero value through config.Load. Mutation-checked: deleting each of the three lines in redisConfig fails this test (the TLS line fails to compile instead, since dropping it leaves the local `t` unused — still a build failure). Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- CHANGELOG.md | 2 +- docs/src/content/docs/api.md | 2 +- docs/src/content/docs/configuration.mdx | 2 +- docs/src/content/docs/deployment.md | 2 +- docs/src/content/docs/development.md | 2 +- docs/src/content/docs/pipes.mdx | 2 +- internal/app/app_test.go | 22 ++++++++++++++++++++++ 7 files changed, 28 insertions(+), 6 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index c4e6e2934..472f94af1 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `internal/ingest/worker.go`, `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`) and `dial_timeout` (`1s`), each at most `1s` since boot and shutdown each wait out a connection attempt they bound, `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a valid port, more than one address in `standalone` mode, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The ingest worker's log of an invalidation that did not land drops from `ERROR` to `WARN`, since the shared backend defers and retries it: an outage would otherwise log an `ERROR` for every batch. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end, and the e2e coverage exclude for its files is gone; the unit and integration suites keep covering `local`. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits; a third runs the first one's hit, insert and fresh-miss lifecycle on the suite's own `cache.backend: local` app, since e2e no longer exercises that backend. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. +- **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `internal/ingest/worker.go`, `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `.github/workflows/{ci.yml,README.md}`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md,index.mdx,why-wavehouse.md,sdk/reference.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`) and `dial_timeout` (`1s`), each at most `1s` since boot and shutdown each wait out a connection attempt they bound, `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a valid port, more than one address in `standalone` mode, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The ingest worker's log of an invalidation that did not land drops from `ERROR` to `WARN`, since the shared backend defers and retries it: an outage would otherwise log an `ERROR` for every batch. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end; the e2e per-suite exclude now names what its run still can't reach instead — `internal/cache/pending.go` (the retry of an invalidation the server did not take, which needs an outage), `internal/config/cache_redis.go` (the block's own rejection paths) and `internal/cache/(local|version_manager).go` (the `local` backend, which e2e no longer runs) — and the unit and integration suites keep covering them. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits; a third runs the first one's hit, insert and fresh-miss lifecycle on the suite's own `cache.backend: local` app, since e2e no longer exercises that backend. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. - **A Redis-compatible shared cache backend** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/cache.go`, `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`, `CONTRIBUTING.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), open a circuit breaker that skips the server until a probe write succeeds (connections are replaced every minute, so after a failover behind a stable address the process reaches the new primary, and delivers the bumps it owes, within about a minute, provided a fresh connection's dial and handshake fit in the per-operation timeout), and deferred invalidations are retried until they land — the first at once, and at the probe's cadence while the breaker is open — collapsing to one tenant-wide bump per tenant past 100,000 keys; until one lands, the process that owes it bypasses the lookups it would orphan. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. `cache.backend: redis` selects it (the entry above). Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression, server-stops-answering, refused-writes (a demoted primary, a full `noeviction` server), failover-behind-a-stable-address and owed-invalidation cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index 0ed66f231..77bbe7921 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -445,7 +445,7 @@ ClickHouse's inline `FORMAT` clause (e.g. `SELECT 1 FORMAT CSV` or `… FORMAT P The proxy buffers the upstream response in memory before forwarding (no row-streaming yet), so a `SELECT *` from a large table can pin RAM on the API server. To avoid an admin OOMing themselves, responses larger than 64 MiB return 502 with a `clickhouse response exceeded N bytes` error. Narrow the query with `LIMIT`, or use a streaming client outside WaveHouse that talks to ClickHouse directly (the standard escape hatch — the same admin credentials work). ::: -This endpoint **does not cache, does not singleflight, and emits `Cache-Control: no-store`** — every request goes straight to ClickHouse, mutation or read, and downstream HTTP caches are explicitly told not to store the response. Raw SQL is an admin escape hatch with infrequent, ad-hoc traffic, so the L1/singleflight machinery would only add complexity without a real hit-rate win. Use [`POST /v1/query?table={table}`](#post-v1querytabletable--structured-query) or [`GET/POST /v1/pipes/{name}`](#getpost-v1pipesname--execute-named-pipe) for the cached read paths (dashboards, high-QPS clients, etc.) — both go through the query cache ([`cache.backend`](/configuration#backends): in-process, or a Redis shared by every instance) with singleflight coalescing. +This endpoint **does not cache, does not singleflight, and emits `Cache-Control: no-store`** — every request goes straight to ClickHouse, mutation or read, and downstream HTTP caches are explicitly told not to store the response. Raw SQL is an admin escape hatch with infrequent, ad-hoc traffic, so the cache/singleflight machinery would only add complexity without a real hit-rate win. Use [`POST /v1/query?table={table}`](#post-v1querytabletable--structured-query) or [`GET/POST /v1/pipes/{name}`](#getpost-v1pipesname--execute-named-pipe) for the cached read paths (dashboards, high-QPS clients, etc.) — both go through the query cache ([`cache.backend`](/configuration#backends): in-process, or a Redis shared by every instance) with singleflight coalescing. :::note[Admin only] The route is mounted under `/v1/ops/*`, behind the `RequireAdmin` gate: only a caller whose JWT role equals the policy `admin_role` (`"admin"` by default) — or who presents the non-JWT [operator key](#authentication) — may use it. A tokenless request (or a valid token without a role claim) resolves to the `default_role` (not the admin role unless `default_role` is deliberately set to it — a loudly-warned dev-only setting) and is rejected with `403`; a present-but-invalid token — expired, malformed, bad signature — keeps its stashed verification error and fails loud with `401` instead. Raw SQL has no per-statement scope check (a full SQL parser would be needed to authorize predicates), so the role gate is the entire authorization story, shared with the rest of `/v1/ops/*` (see [Admin Endpoints](#admin-endpoints)). The normal surfaces for non-admin callers are `POST /v1/ingest?table={table}` for writes, `POST /v1/query?table={table}` for structured reads, and `GET/POST /v1/pipes/{name}` for pre-defined queries — none of which expose raw SQL. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index d625b2e46..de970487c 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -170,7 +170,7 @@ The `redis` backend's settings, read only when `cache.backend` is `redis`. It is **When the server is unreachable or misbehaves, the cache is bypassed; queries are not.** A failure or a timeout makes the lookup a miss and the fill a no-op. Five in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), open a circuit breaker that skips the server entirely until a probe write, every 5 s, succeeds. Queries then go straight to ClickHouse, still coalesced per instance by `singleflight`. An invalidation the server did not take is kept and retried until it lands, and until then the instance that owes it bypasses the lookups it would orphan. `/readyz` does not depend on the cache. -**Boot does not wait for the server.** A malformed block (an address without a port, `mode: cluster` with `db` other than `0`, an unreadable or unparsable TLS file) refuses boot. A server that cannot be reached, or that refuses the credentials, does not: the process boots with the cache bypassed and keeps reconnecting, with backoff up to 30 s. A rejected password (`WRONGPASS`, `NOAUTH`) is logged at `ERROR` on every attempt; any other failure at `WARN`. This is deliberate: a rotated Redis password must not crash-loop every instance at once. Watch `wavehouse_cache_breaker_open`, which reads `1` while the cache is bypassed, including before the first connection. +**Boot does not wait for the server.** A malformed block (an address without a port, `mode: cluster` with `db` other than `0`, an unreadable or unparsable TLS file) refuses boot. A server that cannot be reached, or that refuses the credentials, does not: the process boots with the cache bypassed and keeps reconnecting, with backoff up to 30 s. A rejected credential (`WRONGPASS`, `NOAUTH`, or `NOPERM` for an ACL user missing a connection command) is logged at `ERROR` on every attempt; any other failure at `WARN`. This is deliberate: a rotated Redis password must not crash-loop every instance at once. Watch `wavehouse_cache_breaker_open`, which reads `1` while the cache is bypassed, including before the first connection. ### Authentication diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 1e14399b2..aa29bb6b6 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -437,7 +437,7 @@ The query-result cache is the layer that can be shared today. With the default ` - **A pipe that writes** (an `INSERT` in `pipes.json`) has its result cached like a read, so a repeated identical call is answered from the cache and the write does not run again ([#386](https://github.com/Wave-RF/WaveHouse/issues/386)). With a shared cache that holds on every instance, until the entry's TTL. - **Admin writes through `POST /v1/ops/query`** do not invalidate the cache ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). With a shared cache, the stale results they leave are served by every instance, not only one. -**Sizing the server.** Every key WaveHouse writes has a TTL, and a version token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses, never bring back an entry it had invalidated. So set `maxmemory` and let the server evict: `maxmemory-policy allkeys-lru` (or `allkeys-lfu`, `volatile-lru`, `volatile-lfu`). Under `noeviction`, a full server refuses the writes: each refusal (a fill's is counted by `wavehouse_cache_set_failures_total{reason="oom"}`) bypasses the cache of the instance that got it, and invalidations are kept and retried, so the pre-insert results above stay served by the others: avoid `noeviction`. A stored result is capped at `cache.redis.max_value_bytes` (1 MiB compressed). A tenant's version tokens share one hash tag, so each lookup reads them in one `MGET` in cluster mode as well. The results themselves carry no hash tag and spread across shards. Persistence is not needed: an empty server after a restart is a cold cache, not a wrong one. +**Sizing the server.** Every key WaveHouse writes has a TTL, and a version token lost to eviction, expiry or `FLUSHALL` can only cause misses, never bring back an entry it had invalidated. So set `maxmemory` and let the server evict: `maxmemory-policy allkeys-lru` (or `allkeys-lfu`, `volatile-lru`, `volatile-lfu`). Under `noeviction`, a full server refuses the writes: each refusal (a fill's is counted by `wavehouse_cache_set_failures_total{reason="oom"}`) bypasses the cache of the instance that got it, and invalidations are kept and retried, so the pre-insert results above stay served by the others: avoid `noeviction`. A stored result is capped at `cache.redis.max_value_bytes` (1 MiB compressed). A tenant's version tokens share one hash tag, so each lookup reads them in one `MGET` in cluster mode as well. The results themselves carry no hash tag and spread across shards. **Run it without persistence** (`save ""` and `appendonly no`): a restart without persistence can only cause misses, the same as any other token loss. With persistence on, a restart is not that — it reloads whatever snapshot or AOF it last wrote, tokens and values it had already invalidated included, so a fresh instance can serve the pre-write rows filed under them as hits until their TTL (up to 1 h) expires. Treat restoring a snapshot as a rollback, not a resume. **The server is inside the trust boundary.** A cached result is served after the access policy has filtered it, so whoever can write to the server can change what any caller reads. Keep it on a private network, require a password or ACL user (`WH_CACHE_REDIS_PASSWORD`), use TLS across links you do not trust, and share it only with deployments you trust as much as this one. diff --git a/docs/src/content/docs/development.md b/docs/src/content/docs/development.md index 581df72d1..11fcb82d0 100644 --- a/docs/src/content/docs/development.md +++ b/docs/src/content/docs/development.md @@ -345,7 +345,7 @@ Each test target writes `covdata` to `tmp/coverage//data/`, renders a tex | E2E tests (SDK) | `tests/e2e/sdk/*.test.ts` | Yes | `make test-e2e` | - **Unit tests** live beside the code they test (e.g., `internal/discovery/discovery_test.go`). They use mocks or embedded NATS (in-process, no Docker needed). -- **Integration tests** use the `//go:build integration` build tag. In `tests/integration`, `TestMain` starts one ClickHouse testcontainer and boots the production wiring against it through `app.New` (embedded NATS, ingest worker, sweeper, hub, the API server on a random loopback port); tests reach it via `env(t)` and create their own tables. DLQ tests use `assert.Eventually` with a 30-second timeout for the 5-second ingest worker batch window. `internal/cache`'s integration tests start their own containers instead — Redis, Valkey, Dragonfly and a one-node Redis Cluster — for the shared backend. +- **Integration tests** use the `//go:build integration` build tag. In `tests/integration`, `TestMain` starts one ClickHouse testcontainer and boots the production wiring against it through `app.New` (embedded NATS, ingest worker, sweeper, hub, the API server on a random loopback port); tests reach it via `env(t)` and create their own tables. DLQ tests use `assert.Eventually` with a 30-second timeout for the 5-second ingest worker batch window. `internal/cache`'s integration tests start their own containers instead — Redis, Valkey, Dragonfly and a one-node Redis Cluster — for the shared backend. `shared_cache_test.go` starts its own Redis testcontainer per test (`startRedis`) and boots extra, independent `cache.backend: redis` instances over that same ClickHouse (`bootRedisApp`), to exercise the cache shared across processes rather than one package in isolation. Shared test utilities live in `internal/testutil/`. The packages log through `slog.Default()`, so tests reach log output through `internal/testutil/logtest`: `logtest.Silence()` in a package's `TestMain` discards it, and `logtest.Capture(t, level)` routes it to a buffer for a test that asserts on log lines — such a test must not call `t.Parallel()`, because the default logger is process-wide. diff --git a/docs/src/content/docs/pipes.mdx b/docs/src/content/docs/pipes.mdx index b64129225..b6ded2005 100644 --- a/docs/src/content/docs/pipes.mdx +++ b/docs/src/content/docs/pipes.mdx @@ -7,7 +7,7 @@ sidebar: A **named pipe** is a saved SQL query, registered under a name, that callers run by name with parameters — without ever sending raw SQL. They turn an ad-hoc query into a stable, cached, access-controlled endpoint: you write the SQL once as an operator in the settings directory's [`pipes.json`](/settings-directory#pipesjson), expose it at `GET/POST /v1/pipes/{name}`, and clients supply only the declared parameters. -Pipes are the right tool when a query is reusable and shouldn't live in client code — dashboards, reports, public APIs over curated slices of data. They sit on the **cached read path** (shared L1 + singleflight, same as structured queries), and authorize through a simple per-pipe allowlist rather than the full [policy engine](/access-control). +Pipes are the right tool when a query is reusable and shouldn't live in client code — dashboards, reports, public APIs over curated slices of data. They sit on the **cached read path** (the query cache + singleflight, same as structured queries), and authorize through a simple per-pipe allowlist rather than the full [policy engine](/access-control). ## Anatomy of a pipe diff --git a/internal/app/app_test.go b/internal/app/app_test.go index 7a23330ed..945f7eb96 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -832,6 +832,28 @@ func TestRedisConfig_FromLoadedDefaults(t *testing.T) { assert.Equal(t, cache.RedisSentinel, config.RedisSentinel) } +// Username, DB and TLS are zero on both sides of TestRedisConfig_FromLoadedDefaults' +// assert.Equal, so deleting any of their three mapping lines in redisConfig +// would pass it anyway. Drive all three through config.Load to a non-zero +// value and assert on them directly. +func TestRedisConfig_UsernameDBTLSMapped(t *testing.T) { + t.Setenv("WH_SETTINGS_DIR", t.TempDir()) + t.Setenv("WH_CACHE_BACKEND", "redis") + t.Setenv("WH_CACHE_REDIS_ADDRS", "a:6379") + t.Setenv("WH_CACHE_REDIS_USERNAME", "u") + t.Setenv("WH_CACHE_REDIS_DB", "2") + t.Setenv("WH_CACHE_REDIS_TLS_ENABLED", "true") + t.Setenv("WH_CACHE_REDIS_TLS_SERVER_NAME", "r.internal") + loaded, err := config.Load(filepath.Join(t.TempDir(), "none.yaml")) + require.NoError(t, err) + got, err := redisConfig(loaded.Cache.Redis) + require.NoError(t, err) + assert.Equal(t, "u", got.Username) + assert.Equal(t, 2, got.DB) + require.NotNil(t, got.TLS) + assert.Equal(t, "r.internal", got.TLS.ServerName) +} + // keepalive is a config.json patch setting the stream block's keepalive pair. func keepalive(interval, buckets int) map[string]any { return map[string]any{"stream": map[string]any{"keepalive_interval": interval, "keepalive_buckets": buckets, "gap_window_minutes": 15}} From 568f2ddd20a826be68fd7bc8800e2fd2d7524df5 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:46:46 -0400 Subject: [PATCH 60/79] docs(cache): version_manager.go's own doc carries the query-key/entry-key mixup too Same overload as architecture.md's version_manager.go bullet: the type doc's "A query key folds all three versions of each dependency" means the rendered entry key QueryKey returns, not the caller's input. Reword to match. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- internal/cache/version_manager.go | 7 ++++--- 1 file changed, 4 insertions(+), 3 deletions(-) diff --git a/internal/cache/version_manager.go b/internal/cache/version_manager.go index d67a60778..d6fd7496a 100644 --- a/internal/cache/version_manager.go +++ b/internal/cache/version_manager.go @@ -16,9 +16,10 @@ import ( // place, so the index holds one entry per live tenant, table and scope // however often each is bumped, and forgetting a tenant releases all of it. // -// A query key folds all three versions of each dependency, which gives the -// lattice: a table bump orphans every scope, a scope bump that scope and the -// whole-table view, and a tenant bump everything of the tenant's. +// An entry's key (`QueryKey`) folds all three versions of each dependency, +// which gives the lattice: a table bump orphans every scope, a scope bump +// that scope and the whole-table view, and a tenant bump everything of the +// tenant's. type VersionManager struct { mu sync.RWMutex From f724033c476f95d70e03e402a3eb3dda7ff43852 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:52:50 -0400 Subject: [PATCH 61/79] fix(api): classify a WITH-led statement by INSERT INTO alone After a WITH list ClickHouse parses only SELECT, a FROM-first SELECT or INSERT INTO, yet the scanner took the first word spelled like a statement keyword for the statement. A name in the list could be spelled like any keyword, so `WITH 'd' AS desc INSERT ...` ran as a read (and a retried error wrote again), while `WITH 1 AS set SELECT set` and `WITH 1 AS x FROM system.one SELECT x` ran as writes and answered `[]`. A WITH-led statement is now a write exactly when it holds INSERT INTO outside parentheses. The cases move to internal/testutil/mutationtest so the integration suite can check each one, and every keyword as a WITH list's name, against the pinned ClickHouse's parser (EXPLAIN AST). The classifier is exported as IsMutation for it. Cases that ClickHouse rejects as syntax errors (WITH ... DELETE/ALTER/TRUNCATE among them) are marked so, and the check fails if one starts to parse. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- AGENTS.md | 3 +- CHANGELOG.md | 4 +- docs/src/content/docs/architecture.md | 4 +- internal/api/clickhouse_exec.go | 113 +++---------- internal/api/clickhouse_exec_test.go | 146 ++--------------- internal/api/pipes.go | 4 +- internal/api/query.go | 2 +- internal/testutil/mutationtest/cases.go | 175 +++++++++++++++++++++ tests/integration/ismutation_test.go | 200 ++++++++++++++++++++++++ 9 files changed, 417 insertions(+), 234 deletions(-) create mode 100644 internal/testutil/mutationtest/cases.go create mode 100644 tests/integration/ismutation_test.go diff --git a/AGENTS.md b/AGENTS.md index ea7a87cae..162f01233 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -65,7 +65,7 @@ The invariant index — what must stay true. Full narrative and rationale live i 10. **Active Sweeper** — purges NATS messages that are both ACKed (written to CH) and older than the gap window; SSE gap-fill uses `DeliverByStartTime`, no in-process ring buffer. It runs only in the process holding the `sweeper` lease (`coord.RunElected`); that lease is not fenced: an overlap cannot lose ClickHouse data, since every sweep stops at the consumer's ack floor; it can only trim SSE replay history, and only when the two holders' settings views differ (one still reading a shorter `stream.gap_window_minutes`, or missing a tenant, after a reload the other has applied) — which fencing would not prevent either. Anything that does need exclusivity must check the term's `Token`. 11. **Hasura-style access control: fail-closed (security)** — `policy.IsAdmin` (role == `admin_role`, **exact case-sensitive**, default `"admin"`) is the single admin check, shared by `Evaluate`/`ResolveRole`/`Validate`/the `/v1/ops` gate/`RoleAllowed`. Empty/absent role matches nothing (no `"*"` wildcard); `Validate` rejects empty role keys; a `nil` policy (deleted) denies **everyone incl. admin** via a role — a total lockout for token-based callers, so recovery is writing `policies.json` and reloading, never an implicit admin grant (**exception:** the operator key's `auth.IsOperator` bit passes the `/v1/ops` gate even under a `nil` policy — a deliberate break-glass that can `POST /v1/ops/settings/reload` over HTTP, see #7). Over a nested settings directory the `/v1/ops` gate reads no policy at all — those routes reach every tenant, so the operator key alone passes and an admin-role token gets `403`; `api.NewRouter` decides that from the registry's shape, not from what was wired. `default_role` is the one sanctioned roleless exception (`ResolveRole` maps empty → it pre-eval); `default_role == admin_role` is permitted but dev-only and loudly warned (`policy.DefaultRoleGrantsAdmin`). Preserve when touching `internal/policy` (policy twin of #13; see #159). Detail: architecture.md § `policy/`. 12. **Structured queries: column authz fail-closed (security)** — `POST /v1/query?table={table}`: typed AST validated against schema, permission-enforced, timestamp-bucketed for cache, `DefaultMaxRows` (10,000) cap. Every column reference — projection, aggregation args, `filters`, `group_by`, `order_by`, `time_range` — is authorized inside `query.Build` (the single chokepoint that enumerates them all), so no clause can skip the role's `allow_columns`/`deny_columns` check (#223). A `select_all` read by a *column-restricted* role expands to its allowed columns via `policy.AllowedProjection`, never a bare `SELECT *`; *unrestricted*/admin roles keep `SELECT *` (`policy.RestrictsColumns` decides). Omitting `columns` selects nothing (`ErrEmptyProjection` → `200 []`); `["*"]` is the literal column `*` (schema-gated, not a wildcard); a table-granted role with no readable columns fails closed (`ErrNoReadableColumns` → `403`). Structured and live-stream (`stream.projectIndices`) reads share the one per-column decision `policy.IsColumnAllowed`, so column visibility can't drift. Row visibility has the same one-source guarantee (#319): `Evaluate` resolves a role's row-`filter` once (`resolvePredicates`), and both surfaces consume that single resolution — the query path renders it to SQL (`predicatesToSQL`), the stream evaluates it in memory per subscriber (`ResolvedPermissions.RowVisible`, whose type-aware comparison fails closed on anything it can't prove about the ingested payload — `policy.ColumnSpec`, with `DateTime`/`DateTime64` operands compared as instants through the ingest grammar (`discovery.Column.TimeParser`) and claim constants rendered canonically and digit-exact by the one shared rule `policy.CanonicalScalar` (#457 — which also refuses a float64 at/past 2^53 rather than match a neighboring ID, and whose ok=false — an absent claim, a structured value, no canonical form — makes the predicate match no rows on BOTH surfaces: `1 = 0` in SQL, every row withheld in memory); numeric comparison runs in the column's STORAGE domain (`policy.NumericSpec`, classified by `discovery.NumericStorageOf` — Float width rounding, Decimal scale truncation, integer exactness, both operands narrowed as ClickHouse narrows stored value and bound constant, out-of-range operands refused rather than modeled; the `tests/integration` differential oracle holds in-range verdicts equal to a live ClickHouse's and the never-admit-where-SQL-hides direction for the refused out-of-range ones); an event whose insert later fails into the DLQ is the one residual payload-vs-stored asymmetry, documented in the access-control enforcement caution) — so row visibility can't drift either. Preserve when touching `internal/query` or the structured-query handler. Detail: architecture.md § `query/`. -13. **Named query pipes: fail-closed (security)** — pre-defined SQL templates (Tinybird-style) with param binding + caching — reads only: a pipe whose SQL `isMutation` classifies as a write bypasses the cache and singleflight, since a cached or coalesced write is a dropped one (#386); `GET/POST /v1/pipes/{name}` sit outside `RequireAdmin`, so per-pipe `allowed_roles` is the *only* execute-path gate, via `policy.RoleAllowed`: exact allowlist membership (no `"*"`), admin always passes, empty/absent role and empty-string entries authorize nobody, and no `allowed_roles` → admin-only. Preserve and exercise via `testutil.RunRoleMatrix` / `StandardRoleMatrix` (see #159). Detail: architecture.md § `pipes/`. +13. **Named query pipes: fail-closed (security)** — pre-defined SQL templates (Tinybird-style) with param binding + caching — reads only: a pipe whose SQL `IsMutation` classifies as a write bypasses the cache and singleflight, since a cached or coalesced write is a dropped one (#386); `GET/POST /v1/pipes/{name}` sit outside `RequireAdmin`, so per-pipe `allowed_roles` is the *only* execute-path gate, via `policy.RoleAllowed`: exact allowlist membership (no `"*"`), admin always passes, empty/absent role and empty-string entries authorize nobody, and no `allowed_roles` → admin-only. Preserve and exercise via `testutil.RunRoleMatrix` / `StandardRoleMatrix` (see #159). Detail: architecture.md § `pipes/`. 14. **TypeScript SDK** — `@wavehouse/sdk`: typed query builder, real-time SSE over `fetch`, live queries (incrementable/decomposable/poll aggregation), codegen CLI. Exactly one runtime dependency — `eventsource-parser` (SSE framing, itself dependency-free); adding a second needs the same scrutiny the first got. The canonical client (see §SDK Sync). 15. **Observability invariants** — stdout always 100% (sampling is OTLP-push-only); WARN+ERROR always export at 100% (a non-configurable floor — don't expose it); gRPC OTel exporters dial lazily so an unreachable collector never blocks startup; the OTel Prometheus exporter uses a **private** `prometheus.Registry`. The OTLP endpoint/TLS/custom-CA/mTLS/headers are delegated to the OpenTelemetry SDK's standard `OTEL_EXPORTER_OTLP_*` env vars — `InitProvider` passes **no** endpoint/header options. Known gap, intentionally not patched in WaveHouse app code: the pinned gRPC logs exporter (`otlploggrpc` v0.19/v0.20) ignores the env TLS-cert vars, so a custom/private CA and mutual TLS apply to traces/metrics but **not** the logs signal (public-CA/system-roots TLS and plaintext still work for logs) — upstream bug open-telemetry/opentelemetry-go#6661. A malformed `OTEL_EXPORTER_OTLP_HEADERS` is logged and skipped by the SDK (fail-soft), not fatal. Preserve when touching the logger/sampler/provider. Detail: architecture.md § `observability/`. 16. **Bearer-token-only CORS posture (security)** — Bearer JWT on every request, no cookies/sessions; `corsMiddleware` deliberately **never** emits `Access-Control-Allow-Credentials` (not needed, and `*` + credentials is a spec violation browsers reject). `cors.allowed_origins` (settings directory, per tenant: a tenant route is decorated from the list of the tenant it names, everything else from tenant `0`'s — `corsOrigins`) controls who can *read* responses, not cookie scope; CSRF protection is structural. Don't reintroduce cookie auth or `Allow-Credentials` without a design discussion — answers GitHub #29/#30. Code: `internal/api/router.go`. @@ -130,6 +130,7 @@ Tooling notes (the non-obvious bits `make help` won't tell you): - **JWT helpers**: Use `testutil.MakeJWT(t, claims)` and `testutil.MakeExpiredJWT(t, claims)` for auth tests. See `testutil/jwt.go`. - **Schema helpers**: Use `testutil.NewTestSchemaRegistry(t, tables)` for schema-aware tests — it builds the registry through the real discovery path (`Refresh` against a mock ClickHouse connection), so timestamp specs are precomputed like production. - **Cache backends**: every `cache.Cache` backend runs `cachetest.Run` (`internal/testutil/cachetest`), the backend-agnostic conformance suite; a behavior the contract promises goes there, not in one backend's tests. +- **Write classifier cases**: `api.IsMutation`'s cases live in `internal/testutil/mutationtest`, shared by its unit test and `tests/integration/ismutation_test.go`, which checks each against ClickHouse's own parser; add a case there, not to either test. - **Policy helpers**: Use `policy.Static(p)` for a fixed `policy.Source` in tests. - **Pipes helpers**: Use `pipes.Static(queries...)` for a fixed `pipes.Source` in tests. - **Response assertions**: Use `testutil.AssertJSONResponse(t, rec, status, expected)` and `testutil.AssertJSONContains(t, rec, status, substring)`. diff --git a/CHANGELOG.md b/CHANGELOG.md index 06279b65a..8b2836db1 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -85,8 +85,8 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed -- **A pipe that writes runs on every call instead of being answered from the cache** (`internal/api/{pipes,ch_errors}.go` (+ tests), `docs/src/content/docs/{pipes.mdx,api.md,architecture.md,configuration.mdx,ingest-pipeline.md,sdk/pipes.md,sdk/reference.md}`, `clients/ts/src/pipes.ts` (doc comment), `AGENTS.md`): fixes [#386](https://github.com/Wave-RF/WaveHouse/issues/386). `/v1/pipes/{name}` sent a write's SQL to ClickHouse through `Exec`, but still cached the `[]` it returned and coalesced identical calls in flight, so a repeat within the TTL answered `200` without executing and concurrent identical calls became one write — silently dropped writes, and with a shared cache ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)) on every instance. A pipe whose bound SQL `isMutation` classifies as a write — the same classifier that picks `Exec` — now skips the cache lookup, the fill and singleflight, and answers `X-Cache: BYPASS` with `Cache-Control: no-store`, so an HTTP cache in front of a `GET` cannot drop the write either. Classification stays automatic rather than a declared pipe property, so an operator cannot forget to mark one, and costs no ClickHouse round trip. A failed write answers with the status and `code` a failed read gets (see the ClickHouse-errors entry below), but always `retryable: false` and with no `Retry-After`, `503 clickhouse.unavailable` included: the statement may have run, so the SDK does not retry it. A write refused before it is sent, the tenant on no pool, keeps its `503` with `Retry-After: 30`. Read pipes are unchanged. Not in this fix: a write pipe still does not invalidate cached reads of the table it writes ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). -- **The write classifier skips whitespace, comments and quoted text the way ClickHouse's lexer does** (`internal/api/clickhouse_exec.go` (+ tests)): `isMutation` picks `Exec` for a write, and since [#386](https://github.com/Wave-RF/WaveHouse/issues/386) keeps a write pipe out of the cache. It missed a write behind a backslash-escaped quote (`'it\'s'`, and the same inside `"…"` and `` `…` ``), a heredoc (`$$ ( $$`, `$tag$ … $tag$`), a curly-quoted literal or identifier (`‘(’`, `“c(d”`), a `//` line comment, a nested block comment (`/* a /* b */ SELECT */ INSERT …`), or leading whitespace other than space, tab, CR and LF: `\v`, `\f`, a no-break space, a byte-order mark, and the other Unicode spaces ClickHouse skips. A missed write went through `Query`, which ran it and then failed the call with a `5xx` the TypeScript SDK retries, so one call could write three times. The same gaps, and a word led by `_` (`_delete`) whose tail was read as a verb, could make a read look like a write, which runs through `Exec` and answers `[]`. +- **A pipe that writes runs on every call instead of being answered from the cache** (`internal/api/{pipes,ch_errors}.go` (+ tests), `docs/src/content/docs/{pipes.mdx,api.md,architecture.md,configuration.mdx,ingest-pipeline.md,sdk/pipes.md,sdk/reference.md}`, `clients/ts/src/pipes.ts` (doc comment), `AGENTS.md`): fixes [#386](https://github.com/Wave-RF/WaveHouse/issues/386). `/v1/pipes/{name}` sent a write's SQL to ClickHouse through `Exec`, but still cached the `[]` it returned and coalesced identical calls in flight, so a repeat within the TTL answered `200` without executing and concurrent identical calls became one write — silently dropped writes, and with a shared cache ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)) on every instance. A pipe whose bound SQL `IsMutation` classifies as a write — the same classifier that picks `Exec` — now skips the cache lookup, the fill and singleflight, and answers `X-Cache: BYPASS` with `Cache-Control: no-store`, so an HTTP cache in front of a `GET` cannot drop the write either. Classification stays automatic rather than a declared pipe property, so an operator cannot forget to mark one, and costs no ClickHouse round trip. A failed write answers with the status and `code` a failed read gets (see the ClickHouse-errors entry below), but always `retryable: false` and with no `Retry-After`, `503 clickhouse.unavailable` included: the statement may have run, so the SDK does not retry it. A write refused before it is sent, the tenant on no pool, keeps its `503` with `Retry-After: 30`. Read pipes are unchanged. Not in this fix: a write pipe still does not invalidate cached reads of the table it writes ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). +- **The write classifier skips whitespace, comments and quoted text the way ClickHouse's lexer does** (`internal/api/clickhouse_exec.go` (+ tests), `internal/testutil/mutationtest` (new), `tests/integration/ismutation_test.go` (new), `AGENTS.md`): `IsMutation` picks `Exec` for a write, and since [#386](https://github.com/Wave-RF/WaveHouse/issues/386) keeps a write pipe out of the cache. It missed a write behind a backslash-escaped quote (`'it\'s'`, and the same inside `"…"` and `` `…` ``), a heredoc (`$$ ( $$`, `$tag$ … $tag$`), a curly-quoted literal or identifier (`‘(’`, `“c(d”`), a `//` line comment, a nested block comment (`/* a /* b */ SELECT */ INSERT …`), or leading whitespace other than space, tab, CR and LF: `\v`, `\f`, a no-break space, a byte-order mark, and the other Unicode spaces ClickHouse skips. A missed write went through `Query`, which ran it and then failed the call with a `5xx` the TypeScript SDK retries, so one call could write three times. The same gaps, and a word led by `_` (`_delete`) whose tail was read as a verb, could make a read look like a write, which runs through `Exec` and answers `[]`. After a `WITH` list, which ClickHouse follows only with `SELECT`, a FROM-first `SELECT` or `INSERT INTO`, a name spelled like a keyword was taken for the statement: `WITH 'd' AS desc INSERT …` and `WITH 1 AS select INSERT …` ran as reads, and `WITH 1 AS set SELECT set` and `WITH 1 AS x FROM system.one SELECT x` as writes. A `WITH`-led statement is now a write exactly when it holds `INSERT INTO` outside parentheses. The classifier, exported as `IsMutation` for it, is now checked against the pinned ClickHouse's own parser (`EXPLAIN AST`) in the integration suite: every test case, and every ClickHouse keyword as a `WITH` list's name ahead of each statement a `WITH` list can lead. - **The pipes page no longer says a parameter can never break out of its literal** (`docs/src/content/docs/pipes.mdx`, `internal/pipes/pipes.go`): that holds only for a placeholder written bare. A string value brings its own quotes, so inside a quoted placeholder they close the template's: the body `{"id": " OR 1=1 OR id = "}` turns `WHERE id = '{{id}}'` into `WHERE id = '' OR 1=1 OR id = ''`, which matches every row. The page now says to write each placeholder bare, never inside quotes. - **A failed ClickHouse query answers by what went wrong, not a flat `500`/`502`** (`internal/api/ch_errors.go` (new, + tests), `internal/api/{errors,query,structured_query,pipes,schema,ch_settings}.go`, `internal/chconn/errclass.go` (`HTTPStatus` exported), `clients/ts/src/errors.ts` (+ tests), `tests/integration/query_errors_test.go` (new), `tests/integration/query_limits_test.go`, `internal/app/app_test.go`, `tests/e2e/sdk/{admin,query}.test.ts`, `AGENTS.md`, `docs/src/content/docs/{api,architecture}.md`, `docs/src/content/docs/{access-control,configuration}.mdx`, `docs/src/content/docs/sdk/{reference.md,index.mdx}`): fixes [#403](https://github.com/Wave-RF/WaveHouse/issues/403) and [#271](https://github.com/Wave-RF/WaveHouse/issues/271), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). ClickHouse answers a syntax error, a missing grant and an overloaded server alike with HTTP `500`, so `/v1/ops/query` turned a bad statement into a `502` and `/v1/query` and pipes into a `500` the SDK retried. All three now class the failure with `chconn.Classify` through one helper, `writeCHError`: a statement ClickHouse refused is `400 clickhouse.rejected`; a query over a rows/bytes limit, the role's own memory cap, or its time cap where that is no longer than `query_timeout` is `400 clickhouse.limit_exceeded`; `ACCESS_DENIED` is `403 clickhouse.access_denied`; credentials, user or database refused, or a redirect or `4xx` with no exception code from whatever fronts ClickHouse, is `502 clickhouse.misconfigured`; ClickHouse down, unreachable or overloaded is `503 clickhouse.unavailable` with `Retry-After: 5`; a failure with no verdict stays `500` (`502` on the proxy) as `clickhouse.unknown`. The error envelope gains `code` and `retryable` next to `error` on these responses — additive. A role with `max_execution_time` now queries with no context deadline and a cancel two seconds past the cap instead: clickhouse-go overwrote the cap's `max_execution_time` with deadline+5s for any deadline over 1s, so an overrun came back as a bare deadline, indistinguishable from waiting for a pooled connection; ClickHouse now enforces the cap itself and reports `TIMEOUT_EXCEEDED`. `POST /v1/ops/schema/refresh` against an unreachable ClickHouse is a `503` with `Retry-After` instead of a `500`. **SDK:** `WaveHouseError.code` and `retryable` now take the server's `code`/`retryable` when the body has them (`HTTP_` and "5xx retries" otherwise), so a rejected query is `clickhouse.rejected` rather than `HTTP_500`, and is not retried. - **An unavailable ClickHouse is retried with backoff instead of dead-lettering every row** (`internal/chconn/errclass.go` (new, + tests), `internal/ingest/{worker,backoff}.go` (`backoff.go` new, + tests), `internal/mq/{mq,embedded}.go`, `internal/testutil/mocks.go`, `tests/integration/ingest_outage_test.go` (new), `AGENTS.md`, `README.md`, `docs/src/content/docs/{ingest-pipeline,architecture,api,deployment,why-wavehouse}.md`, `docs/src/content/docs/{settings-directory,index,access-control}.mdx`): workstream A of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). A failed batch insert used to go through row-by-row isolation whatever the failure, so a ClickHouse that was down, overloaded or read-only failed every row twice and parked the whole batch on the DLQ. `chconn.Classify` now classes the failure first — `Rejected` (any ClickHouse exception code outside the availability and credential lists: the server read the row and refused it), `Unavailable` (connection refused/reset, timeouts, `TOO_MANY_SIMULTANEOUS_QUERIES`, `SERVER_OVERLOADED`, `MEMORY_LIMIT_EXCEEDED`, `TOO_MANY_PARTS`, `READONLY`, `TABLE_IS_READ_ONLY`, `KEEPER_EXCEPTION`, …), `Denied` (`AUTHENTICATION_FAILED`, `ACCESS_DENIED`, …) or `Unknown` (no code, no recognizable transport failure). Only `Rejected` is isolated and dead-lettered as before, and a multi-row batch refused with `TOO_MANY_PARTS` or `MEMORY_LIMIT_EXCEEDED` is split row by row first (`chconn.Splittable`), because a batch spanning too many partitions or too much memory can fail where each of its rows inserts; every other class hands the batch back to the queue with a delayed nak (`mq.Message.NakWithDelay`, new) under a jittered 1 s → 30 s backoff shared by every table on the same ClickHouse pool (a failure of one table — read-only, too many parts or mutations, a grant missing on it, `chconn.TableScoped` — backs off that table alone), which turns rows away without a request while it runs and probes once per window, and ClickHouse going away mid-isolation stops isolation and retries the rows it had not settled. Counted by the new `wavehouse_ingest_retries_total{table, reason}`; logged at `WARN` when an outage starts and at most every 30 s during it. A long outage now shows as a growing ingest stream and, at `mq.max_bytes_gb`, ingest `503`s — not as a full DLQ; a lasting failure of one table holds back its tenant's other tables once its waiting rows reach `maxAckPending`. Retried rows come back out of arrival order, which matters only to a `ReplacingMergeTree` without a version column or a `CollapsingMergeTree`. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index e243dfc90..ab4c00e92 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -80,7 +80,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request - **router.go** — Route definitions. Public: `/livez`, `/readyz`, and the content-free `/v1/health` SDK ping (plus the permanent `/healthz` alias and the deprecated `/health`, `/ready` aliases). Policy-gated: `/v1/ingest?table={table}`, `/v1/query?table={table}` (structured), `/v1/pipes/{name}` (named pipes), `/v1/stream`. Admin-only (`RequireAdmin` — role == `policy.admin_role`, or a request bearing the operator key's operator bit, which passes even under a nil policy; over a nested settings directory `NewRouter` mounts the gate with no policy at all, whatever `Dependencies.PolicySource` was wired, so the operator key alone passes): `/v1/ops/schema/*`, `/v1/ops/dlq/stats`, `GET /v1/ops/pipes[/{name}]`, `/v1/ops/settings/reload`, `/v1/ops/query` (raw SQL — same gate as the rest of `/v1/ops/*`). `NewOpsRouter` is the router of a process without the `api` role: the probes and their aliases, `/version` and the same-port metrics path (the part it shares with `NewRouter`, `newProbeRouter`), and `POST /v1/ops/settings/reload` behind `RequireAdmin(nil)`, so only the operator key passes; every other route is a 404, under `/v1/ops` only once that gate has passed. - **auth middleware** — the JWT/JWKS authentication middleware is its own package, [`auth/`](#auth--authentication); the router runs it on every `/v1/*` route. - **tenant.go** — `TenantMW` resolves the request's tenant ahead of the auth middleware on every `/v1` route outside `/v1/ops/*`: the [`X-Tenant-ID`](/deployment#multi-tenant-deployments) header (absent means `tenant.Default`), validated by `tenant.Parse` (`400`), looked up in the `settings.Registry` (`resolveStore`: `404` for an id it does not hold, a bare `503` for a tenant whose folder was rejected — the findings stay out of a body answered before authentication), and the resolved `*settings.Store` stored in the request context (`WithStore` / `StoreFromContext` — here rather than in `tenant/`, because `settings` names `tenant.ID`). A handler reads the store once and passes it down as an argument — the per-tenant getters it holds take it as a parameter (`(*settings.Store).Policy`, `.DedupeFor`, `.DefaultMaxRows`, … in production) — and nothing below a handler reads the context; a tenant route reached without a resolved store answers `500` rather than fall back to a tenant. The probes, `/version`, the metrics path, and `/v1/ops/*` are tenant-exempt; an ops route that addresses one tenant — the admin pipe reads, the schema routes, the raw-SQL proxy, the settings reload and the DLQ stats — names it in `?tenant=` (`opsTenant`, and `opsStore` over it for the routes that need the tenant's store; the DLQ stats need none, since the MQ holds the queue), parsed strictly so that a query `url.ParseQuery` would half-read is a `400` rather than a read of the default tenant, which is what it means when absent on the reads (on the reload, absent is the whole directory). -- **pipes.go** — Named query pipe handlers: admin listing (`GET /v1/ops/pipes[/{name}]`, read per request from its `pipes.Source`) and execution with parameter binding. A read is cached and coalesced; a write — bound SQL that `isMutation` (`clickhouse_exec.go`) classifies as one — bypasses both and runs every call. `pipes.json` is the only way to define or change a pipe. +- **pipes.go** — Named query pipe handlers: admin listing (`GET /v1/ops/pipes[/{name}]`, read per request from its `pipes.Source`) and execution with parameter binding. A read is cached and coalesced; a write — bound SQL that `IsMutation` (`clickhouse_exec.go`) classifies as one — bypasses both and runs every call. `pipes.json` is the only way to define or change a pipe. - **structured_query.go** — Handler for `POST /v1/query?table={table}`: validates query AST, enforces permissions, builds and executes SQL. - **ch_errors.go** — `writeCHError`, the one mapping from a failed ClickHouse query to a response, shared by `/v1/query`, pipes and `/v1/ops/query` so they cannot drift apart: `chconn.Classify` decides the class, and the class the status, `code` and `retryable` ([ClickHouse errors on the query paths](/api#clickhouse-errors-on-the-query-paths)). A write pipe answers through `writeCHWriteError`, the same mapping with `retryable` always `false` and no `Retry-After`, since the write may have run. - **ingest.go** — Accepts `POST /v1/ingest?table={table}` in three body shapes: one flat JSON object, a JSON array of them, or NDJSON. The **required** `Content-Type` chooses the format *family* — `application/json` versus the four NDJSON spellings — and within the JSON family the body's first non-whitespace byte picks array versus single object; the bytes never choose the family. Anything that is not exactly one readable media type is a `415`, decided before the body is read: the header is parsed per RFC 9110 §8.3, and because `Content-Type` is a singleton field, repeated header lines must all resolve to the same format and a value carrying a comma is refused unless the value as a whole parses as one media type — a comma inside a *quoted* parameter value is data, so `application/json; a=", application/x-ndjson; b="` is accepted. It then reads the whole (`MaxBytesReader`-capped) body into a pooled buffer and runs the per-format record readers over those bytes, so the `413` lands before any record is processed and peak memory per request is O(body) rather than O(record). Then it validates each record against the discovered schema, optional dedup, and publishes each row through `mq.Publisher` on `mq.Topic{Tenant, Table, Scope}` (the request's tenant, read off its resolved store — `store.Tenant()` — and raw names; the subject it becomes is `internal/mq`'s; a full queue comes back as `mq.ErrQueueFull`, which is the `503` + `Retry-After: 30`, and a broker that cannot be reached or does not answer in time as `mq.ErrUnavailable`, the `503` + `Retry-After: 5`). When dedup is on, a row missing the configured `id_field` can't be deduped: it is logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total` (labeled by `table`), then published un-deduped — or rejected when `dedupe.require_id` is set ([#219](https://github.com/Wave-RF/WaveHouse/issues/219)). @@ -331,7 +331,7 @@ Client POST /v1/ops/query (browser, CDN, corp proxy) caches the result. ``` -The proxy-pattern wins are: zero classification logic on the WaveHouse side (no isMutation heuristic to maintain), and any ClickHouse statement type — including verbs added in future versions and inline FORMAT overrides — works without WaveHouse code changes. Multi-statement input (`SELECT 1; TRUNCATE t`) is supported when the upstream ClickHouse has multi-query enabled, which is the default on recent versions; older or restrictively-configured servers will return a clear error from ClickHouse itself for the second statement. The proxy buffers the response in memory with a 64 MiB cap (502 with `clickhouse response exceeded N bytes` on overflow, to keep a runaway `SELECT *` from pinning RAM on the API server), and passes ClickHouse's `Content-Type` through when an inline `FORMAT` directive overrides the default JSON envelope. The structured query endpoint and pipes still go through `clickhouse-go`'s native driver (Query/Exec) for performance and to keep the cached row-array shape consistent. +The proxy-pattern wins are: zero classification logic on the WaveHouse side (no `IsMutation` heuristic to maintain), and any ClickHouse statement type — including verbs added in future versions and inline FORMAT overrides — works without WaveHouse code changes. Multi-statement input (`SELECT 1; TRUNCATE t`) is supported when the upstream ClickHouse has multi-query enabled, which is the default on recent versions; older or restrictively-configured servers will return a clear error from ClickHouse itself for the second statement. The proxy buffers the response in memory with a 64 MiB cap (502 with `clickhouse response exceeded N bytes` on overflow, to keep a runaway `SELECT *` from pinning RAM on the API server), and passes ClickHouse's `Content-Type` through when an inline `FORMAT` directive overrides the default JSON envelope. The structured query endpoint and pipes still go through `clickhouse-go`'s native driver (Query/Exec) for performance and to keep the cached row-array shape consistent. ### Streaming Path diff --git a/internal/api/clickhouse_exec.go b/internal/api/clickhouse_exec.go index e34de64e7..83202628d 100644 --- a/internal/api/clickhouse_exec.go +++ b/internal/api/clickhouse_exec.go @@ -43,7 +43,7 @@ func timeoutOf(timeout func(*settings.Store) time.Duration, store *settings.Stor // The raw-SQL endpoint (/v1/ops/query) proxies straight to ClickHouse // over HTTP and never calls this; see internal/api/query.go. func executeCHQuery(ctx context.Context, conn driver.Conn, sql string, params []any) ([]map[string]any, error) { - if isMutation(sql) { + if IsMutation(sql) { if err := conn.Exec(ctx, sql, params...); err != nil { return nil, fmt.Errorf("clickhouse exec: %w", err) } @@ -114,17 +114,16 @@ var mutationVerbs = map[string]struct{}{ "SYSTEM": {}, } -// isMutation reports whether sql's leading statement is a non-SELECT — i.e. +// IsMutation reports whether sql's leading statement is a non-SELECT — i.e. // one that returns no result set and must go through Exec, not Query. // Leading whitespace and comments are skipped as ClickHouse's lexer skips // them, then the first bareword is matched whole, case-insensitively, against -// mutationVerbs. A leading WITH clause (CTE) routes through a paren-aware scan -// because ClickHouse accepts `WITH cte AS (...) INSERT INTO t SELECT * FROM -// cte` as equivalent to `INSERT INTO t WITH cte AS (...) SELECT * FROM cte` -// (see https://clickhouse.com/docs/sql-reference/statements/insert-into). -// A write classified as a read goes through Query, which runs it and then -// fails the call, so a client that retries the error writes again. -func isMutation(sql string) bool { +// mutationVerbs. After a WITH list ClickHouse parses only SELECT, a FROM-first +// SELECT or INSERT INTO, so a WITH-led statement is a write exactly when it +// holds INSERT INTO at the top level (hasTopLevelInsertInto). A write +// classified as a read goes through Query, which runs it and then fails the +// call, so a client that retries the error writes again. +func IsMutation(sql string) bool { s := stripLeadingSQLComments(sql) end := skipWord(s, 0) if end == 0 { @@ -135,52 +134,18 @@ func isMutation(sql string) bool { _, ok := mutationVerbs[first] return ok } - return containsMutationVerbAtTopLevel(s[end:]) + return hasTopLevelInsertInto(s[end:]) } -// nonMutationVerbs is the read/metadata-statement counterpart to -// mutationVerbs. Together they cover every ClickHouse statement-introducing -// keyword that can legally follow a CTE list. The CTE-aware scanner in -// containsMutationVerbAtTopLevel needs the union to identify *which* token -// is the statement keyword — without it, ordinary identifiers in the CTE -// list (table names, database names like the ClickHouse-built-in `system`) -// can collide with mutation-verb names and false-positive the classifier. -var nonMutationVerbs = map[string]struct{}{ - "SELECT": {}, - "SHOW": {}, - "DESCRIBE": {}, - "DESC": {}, - "EXPLAIN": {}, - "EXISTS": {}, - "CHECK": {}, -} - -// containsMutationVerbAtTopLevel scans s for the statement-introducing -// keyword at paren-depth 0, stepping over string literals and quoted -// identifiers (skipQuoted, skipCurlyQuoted), heredocs (skipHeredoc), -// parenthesized CTE subqueries, and comments (skipComment). The CTE list -// contains ordinary identifiers (CTE names, table/database names) that must -// not be matched as mutation verbs — `system` would otherwise pattern-match -// `SYSTEM` and route a `WITH … SELECT * FROM system.tables` read through -// `Exec` (silent empty-array result instead of the actual rows). Two-part fix: -// -// 1. Skip identifiers whose next non-whitespace, non-comment token is -// `AS` (case-insensitive) or `(` — those are CTE definition names -// (with optional column list before AS). This catches the harder -// class where the CTE alias is itself a mutation-verb name -// (`WITH set AS (…) SELECT …`, `WITH alter AS (…) …`, etc.). -// 2. Among the remaining identifiers, stop on the FIRST that's a -// known statement keyword (mutation OR read-class), and decide -// based on mutationVerbs membership. -// -// Tokens that aren't CTE names and aren't statement keywords (RECURSIVE, -// MATERIALIZED, scalar CTE aliases, etc.) are skipped silently. Returns -// false if no statement keyword is found — the SQL is syntactically -// incomplete or unrecognised; safer to treat as non-mutation than to -// silently route an unknown verb through Exec (an Exec'd SELECT returns -// `[]` with no error; a Query'd unrecognised statement surfaces a clear -// error). -func containsMutationVerbAtTopLevel(s string) bool { +// hasTopLevelInsertInto reports whether s holds INSERT INTO outside +// parentheses, stepping over string literals and quoted identifiers +// (skipQuoted, skipCurlyQuoted), heredocs (skipHeredoc) and comments +// (skipComment). No other word is taken for the statement: a WITH list's +// names and aliases may be spelled like any keyword (`WITH 1 AS select`, +// `WITH desc AS (…)`, `WITH set -> 1 AS f`, `WITH t.from AS y`), but only the +// INSERT statement puts INTO after an insert. The exception, a read's +// `… insert INTO OUTFILE 'f'`, is refused by the server either way. +func hasTopLevelInsertInto(s string) bool { depth := 0 i := 0 for i < len(s) { @@ -214,30 +179,12 @@ func containsMutationVerbAtTopLevel(s string) bool { } case isWordByte(c): // A word led by a digit or `_` is read whole, so its tail is - // never taken for a keyword (`_delete`). + // never taken for a keyword (`_insert`). start := i i = skipWord(s, i) - if depth == 0 { - kw := strings.ToUpper(s[start:i]) - // Check non-mutation statement keywords (SELECT, SHOW, - // DESCRIBE, …) FIRST — these can legitimately be followed - // by `(` (e.g. `SELECT (1) FROM …`, `SELECT (a, b) FROM …` - // for tuple syntax), so we must not let the CTE-name - // lookahead below misclassify them as CTE aliases. - if _, ok := nonMutationVerbs[kw]; ok { - return false - } - // CTE name suppression: an identifier that ISN'T a - // non-mutation statement keyword and is followed by `AS` - // or `(` is a CTE definition name (with optional column - // list before AS). Skip without checking mutationVerbs - // — protects against CTE aliases that share a spelling - // with a mutation verb (`WITH set AS (...)`, - // `WITH alter AS (...)`, etc.). - if isCTENameLookahead(s, i) { - continue - } - if _, ok := mutationVerbs[kw]; ok { + if depth == 0 && strings.EqualFold(s[start:i], "INSERT") { + next := skipSpaceAndComments(s, i) + if strings.EqualFold(s[next:skipWord(s, next)], "INTO") { return true } } @@ -250,22 +197,6 @@ func containsMutationVerbAtTopLevel(s string) bool { return false } -// isCTENameLookahead returns true if the next non-whitespace, non-comment -// token at or after pos is `AS` (case-insensitive, word-boundary terminated) -// or `(` — signaling that whatever identifier just ended at pos is a CTE -// definition name (with optional column list before AS). Returns false on EOF -// or any other token. -func isCTENameLookahead(s string, pos int) bool { - i := skipSpaceAndComments(s, pos) - if i >= len(s) { - return false - } - if s[i] == '(' { - return true - } - return strings.EqualFold(s[i:skipWord(s, i)], "AS") -} - // skipWord returns the index just past the bareword at s[i]: ClickHouse's // barewords run over ASCII letters, digits, `_` and `$`. func skipWord(s string, i int) int { diff --git a/internal/api/clickhouse_exec_test.go b/internal/api/clickhouse_exec_test.go index 55aacdd50..c0d8d9b05 100644 --- a/internal/api/clickhouse_exec_test.go +++ b/internal/api/clickhouse_exec_test.go @@ -7,6 +7,7 @@ import ( "time" "github.com/ClickHouse/clickhouse-go/v2/lib/driver" + "github.com/Wave-RF/WaveHouse/internal/testutil/mutationtest" "github.com/google/uuid" "github.com/stretchr/testify/assert" "github.com/stretchr/testify/require" @@ -39,139 +40,14 @@ func (c *stubConn) Query(_ context.Context, _ string, _ ...any) (driver.Rows, er return &chainEmptyRows{}, nil } +// TestIsMutation runs the shared cases; the integration suite checks the +// same cases against ClickHouse's parser. func TestIsMutation(t *testing.T) { t.Parallel() - tests := []struct { - name string - sql string - want bool - }{ - {"select", "SELECT 1", false}, - {"select lower", "select 1", false}, - {"with cte", "WITH x AS (SELECT 1) SELECT * FROM x", false}, - {"show", "SHOW TABLES", false}, - {"describe", "DESCRIBE clicks", false}, - {"explain", "EXPLAIN SELECT 1", false}, - {"exists", "EXISTS TABLE clicks", false}, - - {"insert", "INSERT INTO t VALUES (1)", true}, - {"update", "UPDATE t SET a=1 WHERE b=2", true}, - {"delete", "DELETE FROM t WHERE id=1", true}, - {"truncate", "TRUNCATE TABLE t", true}, - {"truncate lower", "truncate table t", true}, - {"drop", "DROP TABLE t", true}, - {"alter", "ALTER TABLE t ADD COLUMN c String", true}, - {"create", "CREATE TABLE t (a Int)", true}, - {"rename", "RENAME TABLE a TO b", true}, - {"exchange", "EXCHANGE TABLES t1 AND t2", true}, - {"optimize", "OPTIMIZE TABLE t", true}, - {"replace", "REPLACE INTO t VALUES (1)", true}, - {"grant", "GRANT SELECT ON t TO u", true}, - {"revoke", "REVOKE SELECT ON t FROM u", true}, - {"system", "SYSTEM RELOAD CONFIG", true}, - {"attach", "ATTACH TABLE t FROM '/path'", true}, - {"detach", "DETACH TABLE t", true}, - {"kill", "KILL QUERY WHERE query_id = 'abc'", true}, - {"set", "SET max_threads = 4", true}, - {"use", "USE mydb", true}, - - {"leading whitespace", " \n\tTRUNCATE TABLE t", true}, - {"line comment then mutation", "-- drop guard\nDROP TABLE t", true}, - {"hash line comment then mutation", "# audit\nDROP TABLE t", true}, - {"block comment then mutation", "/* admin */ ALTER TABLE t ADD COLUMN c Int", true}, - {"mixed comments then select", "-- foo\n# bar\n/* baz */ SELECT 1", false}, - {"with insert", "WITH cte AS (SELECT 1) INSERT INTO t SELECT * FROM cte", true}, - {"with insert lower", "with cte as (select 1) insert into t select * from cte", true}, - {"with delete", "WITH cte AS (SELECT id FROM x) DELETE FROM t WHERE id IN (SELECT id FROM cte)", true}, - {"with update", "WITH cte AS (SELECT 1) ALTER TABLE t UPDATE a=1 WHERE id IN (SELECT id FROM cte)", true}, - {"with truncate", "WITH cte AS (SELECT 1) TRUNCATE TABLE t", true}, - {"with multi-cte insert", "WITH a AS (SELECT 1), b AS (SELECT 2) INSERT INTO t SELECT * FROM a JOIN b", true}, - {"with nested parens insert", "WITH cte AS (SELECT id FROM t WHERE id IN (1,2,3)) INSERT INTO t2 SELECT * FROM cte", true}, - {"with paren-in-string insert", "WITH cte AS (SELECT ')' AS x) INSERT INTO t2 SELECT * FROM cte", true}, - {"with materialized insert", "WITH cte AS MATERIALIZED (SELECT 1) INSERT INTO t SELECT * FROM cte", true}, - {"with recursive select", "WITH RECURSIVE x AS (SELECT 1 UNION ALL SELECT * FROM x) SELECT * FROM x", false}, - {"with nested select", "WITH x AS (SELECT 1) SELECT * FROM (SELECT * FROM x)", false}, - {"with scalar insert", "WITH '/path' AS p INSERT INTO files VALUES (p)", true}, - {"with line comment containing DELETE then select", "WITH cte AS (SELECT 1) -- old DELETE approach\nSELECT * FROM cte", false}, - {"with hash comment containing TRUNCATE then select", "WITH cte AS (SELECT 1) # was TRUNCATE\nSELECT * FROM cte", false}, - {"with block comment containing INSERT then select", "WITH cte AS (SELECT 1) /* INSERT reminder */ SELECT * FROM cte", false}, - {"with comment then real mutation", "WITH cte AS (SELECT 1) -- explanatory\nINSERT INTO t SELECT * FROM cte", true}, - {"with unclosed block comment", "WITH cte AS (SELECT 1) /* unterminated comment DELETE", false}, - {"with select from system tables (collision regression)", "WITH x AS (SELECT 1) SELECT * FROM system.tables", false}, - {"with select from system columns lower (collision regression)", "with x as (select 1) select name from system.columns", false}, - {"with select aliased as set (false positive regression)", "WITH cte AS (SELECT 1) SELECT * FROM cte AS set", false}, - {"with select from system tables then real insert", "WITH x AS (SELECT * FROM system.tables) INSERT INTO snapshot SELECT * FROM x", true}, - {"with CTE alias named set (read)", "WITH set AS (SELECT 1) SELECT * FROM set", false}, - {"with CTE alias named alter (read)", "WITH alter AS (SELECT 1) SELECT id FROM alter", false}, - {"with CTE alias named drop lowercase (read)", "with drop as (select 1) select * from drop", false}, - {"with CTE alias named update then real update", "WITH update AS (SELECT id FROM x) ALTER TABLE other UPDATE c=1 WHERE id IN (SELECT id FROM update)", true}, - {"with CTE name with column list (read)", "WITH cte (a, b) AS (SELECT 1, 2) SELECT * FROM cte", false}, - {"with multi-CTE both with verb-name aliases (read)", "WITH set AS (SELECT 1), kill AS (SELECT 2) SELECT * FROM set JOIN kill", false}, - {"with parenthesized SELECT then system table (CTE-lookahead ordering regression)", "WITH x AS (SELECT 1) SELECT (1) FROM system.tables", false}, - {"with tuple-shape SELECT then system table", "WITH x AS (SELECT 1) SELECT (a, b) FROM system.parts", false}, - - // A backslash escapes the next byte inside all three quote kinds, so an - // escaped quote does not end the literal or identifier. - {"with backslash-escaped quote in literal then insert", `WITH m AS (SELECT 'it\'s' AS s) INSERT INTO t SELECT s FROM m`, true}, - {"with backslash-escaped quote in literal then select", `WITH m AS (SELECT 'a\'b' AS s) SELECT 'x) INSERT' FROM m`, false}, - {"with backslash-escaped double quote then insert", `WITH m AS (SELECT 'x' AS "a\"(b") INSERT INTO t SELECT * FROM m`, true}, - {"with backslash-escaped double quote then select", `WITH m AS (SELECT 1 AS "a\"b") SELECT 2 AS "x) INSERT" FROM m`, false}, - {"with backslash-escaped backtick then insert", "WITH m AS (SELECT 'x' AS `a\\`(b`) INSERT INTO t SELECT * FROM m", true}, - {"with backslash-escaped backtick then select", "WITH m AS (SELECT 1 AS `a\\`b`) SELECT 2 AS `x) INSERT` FROM m", false}, - - // A heredoc ($$…$$, $tag$…$tag$) is a literal: its parens, quotes - // and words are not the statement's. - {"with heredoc holding a paren then insert", "WITH $$ ( $$ AS s INSERT INTO t SELECT s", true}, - {"with tagged heredoc holding a quote then insert", "WITH $x$ it's $x$ AS s INSERT INTO t SELECT s", true}, - {"with tagged heredoc holding a paren and another tag then insert", "WITH $x$ ( $y$ $x$ AS s INSERT INTO t SELECT s", true}, - {"with heredoc holding a verb then select", "WITH $$INSERT$$ AS s SELECT s", false}, - {"with tagged heredoc holding a paren and a verb then select", "WITH $x$ ) INSERT $x$ AS s SELECT s", false}, - {"with CTE alias set$ (read)", "WITH set$ AS (SELECT 1 AS v) SELECT * FROM set$", false}, - - // A word led by `_` is one bareword, never a keyword's tail, and a - // leading bareword is matched whole, never by its first letters. - {"leading bareword insert_log", "insert_log VALUES (1)", false}, - {"leading bareword insert2", "insert2 INTO t VALUES (1)", false}, - {"with alias _delete (read)", "WITH 1 AS _delete SELECT _delete", false}, - {"with alias _set (read)", "WITH [1,2] AS _set SELECT has(_set, 1)", false}, - - // `//` starts a line comment. - {"slash comment then insert", "// note\nINSERT INTO t VALUES (1)", true}, - {"slash comment hiding insert then select", "// INSERT\nSELECT 1", false}, - {"with slash comment holding a paren then insert", "WITH x AS (SELECT 'a' AS s) // (\nINSERT INTO t SELECT * FROM x", true}, - - // ‘…’ is a string literal and “…” a quoted identifier; nothing escapes - // inside them. - {"with curly-quoted literal holding a paren then insert", "WITH x AS (SELECT \u2018(\u2019 AS s) INSERT INTO t SELECT s FROM x", true}, - {"with curly-quoted literal holding a verb then select", "WITH x AS (SELECT \u2018) INSERT\u2019 AS s) SELECT s FROM x", false}, - {"with curly-quoted identifier holding a paren then insert", "WITH x AS (SELECT 'q' AS \u201cc(d\u201d) INSERT INTO t SELECT * FROM x", true}, - - // ClickHouse's lexer skips \v, \f and Unicode spaces as whitespace - // (TestIsMutation_ClickHouseWhitespace covers the whole set). - {"leading form feed then insert", "\fINSERT INTO t VALUES (1)", true}, - {"leading vertical tab then insert", "\vINSERT INTO t VALUES (1)", true}, - {"leading NBSP then insert", "\u00a0INSERT INTO t VALUES (1)", true}, - {"leading BOM then insert", "\ufeffINSERT INTO t VALUES (1)", true}, - {"leading NBSP then select", "\u00a0SELECT 1", false}, - {"with CTE alias named set before form feed AS (read)", "WITH set\fAS (SELECT 1) SELECT * FROM set", false}, - {"with CTE alias named set before NBSP AS (read)", "WITH set\u00a0AS (SELECT 1) SELECT * FROM set", false}, - - // ClickHouse block comments nest. - {"nested block comment hiding select then insert", "/* a /* b */ SELECT */ INSERT INTO t VALUES (1)", true}, - {"nested block comment hiding insert then select", "/* a /* b */ INSERT */ SELECT 1", false}, - {"unclosed nested block comment", "/* a /* b */ INSERT INTO t VALUES (1)", false}, - {"with nested block comment hiding select then insert", "WITH x AS (SELECT 1) /* a /* b */ SELECT */ INSERT INTO t SELECT * FROM x", true}, - {"with nested block comment hiding insert then select", "WITH x AS (SELECT 1) /* a /* b */ INSERT */ SELECT * FROM x", false}, - {"with nested block comment before CTE AS (read)", "WITH set /* a /* b */ c */ AS (SELECT 1) SELECT * FROM set", false}, - - {"empty", "", false}, - {"comment only", "-- just a comment", false}, - {"unclosed block comment", "/* never closed", false}, - } - for _, tt := range tests { - t.Run(tt.name, func(t *testing.T) { + for _, tc := range mutationtest.Cases { + t.Run(tc.Name, func(t *testing.T) { t.Parallel() - assert.Equal(t, tt.want, isMutation(tt.sql)) + assert.Equal(t, tc.Mutation, IsMutation(tc.SQL)) }) } } @@ -187,13 +63,13 @@ func TestIsMutation_ClickHouseWhitespace(t *testing.T) { } for _, r := range spaces { ws := string(r) - assert.True(t, isMutation(ws+"INSERT INTO t VALUES (1)"), "U+%04X before INSERT", r) - assert.False(t, isMutation(ws+"SELECT 1"), "U+%04X before SELECT", r) - assert.True(t, isMutation("WITH x AS (SELECT 1)"+ws+"INSERT INTO t SELECT * FROM x"), "U+%04X before a WITH's INSERT", r) - assert.False(t, isMutation("WITH set"+ws+"AS (SELECT 1) SELECT * FROM set"), "U+%04X between a CTE name and AS", r) + assert.True(t, IsMutation(ws+"INSERT INTO t VALUES (1)"), "U+%04X before INSERT", r) + assert.False(t, IsMutation(ws+"SELECT 1"), "U+%04X before SELECT", r) + assert.True(t, IsMutation("WITH x AS (SELECT 1)"+ws+"INSERT INTO t SELECT * FROM x"), "U+%04X before a WITH's INSERT", r) + assert.True(t, IsMutation("WITH 1 AS x INSERT"+ws+"INTO t SELECT x"), "U+%04X between a WITH's INSERT and INTO", r) } // Not whitespace to ClickHouse (it rejects the statement), so not skipped. - assert.False(t, isMutation("\u1680INSERT INTO t VALUES (1)")) + assert.False(t, IsMutation("\u1680INSERT INTO t VALUES (1)")) } func TestExecuteCHQuery_MutationRoutesToExec(t *testing.T) { diff --git a/internal/api/pipes.go b/internal/api/pipes.go index 9e28ea468..3a850d3d5 100644 --- a/internal/api/pipes.go +++ b/internal/api/pipes.go @@ -152,7 +152,7 @@ func (h *PipesHandler) Execute(w http.ResponseWriter, r *http.Request) { return } - if isMutation(sql) { + if IsMutation(sql) { h.executeWrite(w, r, store, sql, params) return } @@ -211,7 +211,7 @@ func (h *PipesHandler) Execute(w http.ResponseWriter, r *http.Request) { // executeWrite runs a pipe that writes, on every call: a cached or coalesced // response would answer a repeat without executing it, silently dropping the -// write (#386) — on every instance once the cache is shared. isMutation is the +// write (#386) — on every instance once the cache is shared. IsMutation is the // classifier executeCHQuery routes Exec by, so what bypasses here is exactly // what runs as a write. no-store keeps an HTTP cache in front of a GET from // answering a repeat the same way. diff --git a/internal/api/query.go b/internal/api/query.go index 0959bf670..fa782b0c8 100644 --- a/internal/api/query.go +++ b/internal/api/query.go @@ -37,7 +37,7 @@ import ( // the upstream ClickHouse has multi-query enabled, which is the // default in recent versions; older or restrictively-configured // servers may reject the second statement with a clear error. -// - There is no isMutation heuristic to maintain — no leading-verb table, +// - There is no IsMutation heuristic to maintain — no leading-verb table, // no comment stripper, no CTE-aware paren scanner, no class of bug // where a future ClickHouse verb routes the wrong way. // - ClickHouse's own error messages reach the admin verbatim, which is diff --git a/internal/testutil/mutationtest/cases.go b/internal/testutil/mutationtest/cases.go new file mode 100644 index 000000000..a89c1df21 --- /dev/null +++ b/internal/testutil/mutationtest/cases.go @@ -0,0 +1,175 @@ +// Package mutationtest holds the statements api.IsMutation is tested on, +// shared by its unit test and by the integration test that checks each one +// against ClickHouse's own parser. Add a case here, not to either test. +package mutationtest + +// Case is a statement and how api.IsMutation classifies it. +type Case struct { + Name string + SQL string + // Mutation is true for a statement that goes through Exec. + Mutation bool + // Unparsed marks a statement ClickHouse rejects as a syntax error, so its + // parser has no answer to check Mutation against; the integration test + // fails if one starts to parse. + Unparsed bool +} + +// Cases is every statement api.IsMutation is tested on. +var Cases = []Case{ + {"select", "SELECT 1", false, false}, + {"select lower", "select 1", false, false}, + {"with cte", "WITH x AS (SELECT 1) SELECT * FROM x", false, false}, + {"show", "SHOW TABLES", false, false}, + {"describe", "DESCRIBE clicks", false, false}, + {"explain", "EXPLAIN SELECT 1", false, false}, + {"exists", "EXISTS TABLE clicks", false, false}, + + {"insert", "INSERT INTO t VALUES (1)", true, false}, + {"update", "UPDATE t SET a=1 WHERE b=2", true, false}, + {"delete", "DELETE FROM t WHERE id=1", true, false}, + {"truncate", "TRUNCATE TABLE t", true, false}, + {"truncate lower", "truncate table t", true, false}, + {"drop", "DROP TABLE t", true, false}, + {"alter", "ALTER TABLE t ADD COLUMN c String", true, false}, + {"create", "CREATE TABLE t (a Int)", true, false}, + {"rename", "RENAME TABLE a TO b", true, false}, + {"exchange", "EXCHANGE TABLES t1 AND t2", true, false}, + {"optimize", "OPTIMIZE TABLE t", true, false}, + {"replace into", "REPLACE INTO t VALUES (1)", true, true}, + {"replace table", "REPLACE TABLE t (a Int) ENGINE = Memory", true, false}, + {"grant", "GRANT SELECT ON t TO u", true, false}, + {"revoke", "REVOKE SELECT ON t FROM u", true, false}, + {"system", "SYSTEM RELOAD CONFIG", true, false}, + {"attach", "ATTACH TABLE t FROM '/path'", true, false}, + {"detach", "DETACH TABLE t", true, false}, + {"kill", "KILL QUERY WHERE query_id = 'abc'", true, false}, + {"set", "SET max_threads = 4", true, false}, + {"use", "USE mydb", true, false}, + + {"leading whitespace", " \n\tTRUNCATE TABLE t", true, false}, + {"line comment then mutation", "-- drop guard\nDROP TABLE t", true, false}, + {"hash line comment then mutation", "# audit\nDROP TABLE t", true, false}, + {"block comment then mutation", "/* admin */ ALTER TABLE t ADD COLUMN c Int", true, false}, + {"mixed comments then select", "-- foo\n# bar\n/* baz */ SELECT 1", false, false}, + {"with insert", "WITH cte AS (SELECT 1) INSERT INTO t SELECT * FROM cte", true, false}, + {"with insert lower", "with cte as (select 1) insert into t select * from cte", true, false}, + {"with multi-cte insert", "WITH a AS (SELECT 1), b AS (SELECT 2) INSERT INTO t SELECT * FROM a CROSS JOIN b", true, false}, + {"with nested parens insert", "WITH cte AS (SELECT id FROM t WHERE id IN (1,2,3)) INSERT INTO t2 SELECT * FROM cte", true, false}, + {"with paren-in-string insert", "WITH cte AS (SELECT ')' AS x) INSERT INTO t2 SELECT * FROM cte", true, false}, + {"with materialized insert", "WITH cte AS MATERIALIZED (SELECT 1) INSERT INTO t SELECT * FROM cte", true, false}, + {"with recursive select", "WITH RECURSIVE x AS (SELECT 1 UNION ALL SELECT * FROM x) SELECT * FROM x", false, false}, + {"with nested select", "WITH x AS (SELECT 1) SELECT * FROM (SELECT * FROM x)", false, false}, + {"with scalar insert", "WITH '/path' AS p INSERT INTO files VALUES (p)", true, false}, + {"with line comment containing DELETE then select", "WITH cte AS (SELECT 1) -- old DELETE approach\nSELECT * FROM cte", false, false}, + {"with hash comment containing TRUNCATE then select", "WITH cte AS (SELECT 1) # was TRUNCATE\nSELECT * FROM cte", false, false}, + {"with block comment containing INSERT then select", "WITH cte AS (SELECT 1) /* INSERT reminder */ SELECT * FROM cte", false, false}, + {"with comment then real mutation", "WITH cte AS (SELECT 1) -- explanatory\nINSERT INTO t SELECT * FROM cte", true, false}, + {"with unclosed block comment", "WITH cte AS (SELECT 1) /* unterminated comment DELETE", false, true}, + {"with select from system tables (collision regression)", "WITH x AS (SELECT 1) SELECT * FROM system.tables", false, false}, + {"with select from system columns lower (collision regression)", "with x as (select 1) select name from system.columns", false, false}, + {"with select aliased as set (false positive regression)", "WITH cte AS (SELECT 1) SELECT * FROM cte AS set", false, false}, + {"with select from system tables then real insert", "WITH x AS (SELECT * FROM system.tables) INSERT INTO snapshot SELECT * FROM x", true, false}, + {"with CTE alias named set (read)", "WITH set AS (SELECT 1) SELECT * FROM set", false, false}, + {"with CTE alias named alter (read)", "WITH alter AS (SELECT 1) SELECT id FROM alter", false, false}, + {"with CTE alias named drop lowercase (read)", "with drop as (select 1) select * from drop", false, false}, + {"with CTE name with column list (read)", "WITH cte (a, b) AS (SELECT 1, 2) SELECT * FROM cte", false, false}, + {"with multi-CTE both with verb-name aliases (read)", "WITH set AS (SELECT 1), kill AS (SELECT 2) SELECT * FROM set CROSS JOIN kill", false, false}, + {"with parenthesized SELECT then system table (CTE-lookahead ordering regression)", "WITH x AS (SELECT 1) SELECT (1) FROM system.tables", false, false}, + {"with tuple-shape SELECT then system table", "WITH x AS (SELECT 1) SELECT (a, b) FROM system.parts", false, false}, + + // A backslash escapes the next byte inside all three quote kinds, so an + // escaped quote does not end the literal or identifier. + {"with backslash-escaped quote in literal then insert", `WITH m AS (SELECT 'it\'s' AS s) INSERT INTO t SELECT s FROM m`, true, false}, + {"with backslash-escaped quote in literal then select", `WITH m AS (SELECT 'a\'b' AS s) SELECT 'x) INSERT' FROM m`, false, false}, + {"with backslash-escaped double quote then insert", `WITH m AS (SELECT 'x' AS "a\"(b") INSERT INTO t SELECT * FROM m`, true, false}, + {"with backslash-escaped double quote then select", `WITH m AS (SELECT 1 AS "a\"b") SELECT 2 AS "x) INSERT" FROM m`, false, false}, + {"with backslash-escaped backtick then insert", "WITH m AS (SELECT 'x' AS `a\\`(b`) INSERT INTO t SELECT * FROM m", true, false}, + {"with backslash-escaped backtick then select", "WITH m AS (SELECT 1 AS `a\\`b`) SELECT 2 AS `x) INSERT` FROM m", false, false}, + + // A heredoc ($$…$$, $tag$…$tag$) is a literal: its parens, quotes + // and words are not the statement's. + {"with heredoc holding a paren then insert", "WITH $$ ( $$ AS s INSERT INTO t SELECT s", true, false}, + {"with tagged heredoc holding a quote then insert", "WITH $x$ it's $x$ AS s INSERT INTO t SELECT s", true, false}, + {"with tagged heredoc holding a paren and another tag then insert", "WITH $x$ ( $y$ $x$ AS s INSERT INTO t SELECT s", true, false}, + {"with heredoc holding a verb then select", "WITH $$INSERT$$ AS s SELECT s", false, false}, + {"with tagged heredoc holding a paren and a verb then select", "WITH $x$ ) INSERT $x$ AS s SELECT s", false, false}, + {"with CTE alias set$ (read)", "WITH set$ AS (SELECT 1 AS v) SELECT * FROM set$", false, false}, + + // A word led by `_` is one bareword, never a keyword's tail, and a + // leading bareword is matched whole, never by its first letters. + {"leading bareword insert_log", "insert_log VALUES (1)", false, true}, + {"leading bareword insert2", "insert2 INTO t VALUES (1)", false, true}, + {"with alias _delete (read)", "WITH 1 AS _delete SELECT _delete", false, false}, + {"with alias _set (read)", "WITH [1,2] AS _set SELECT has(_set, 1)", false, false}, + + // `//` starts a line comment. + {"slash comment then insert", "// note\nINSERT INTO t VALUES (1)", true, false}, + {"slash comment hiding insert then select", "// INSERT\nSELECT 1", false, false}, + {"with slash comment holding a paren then insert", "WITH x AS (SELECT 'a' AS s) // (\nINSERT INTO t SELECT * FROM x", true, false}, + + // ‘…’ is a string literal and “…” a quoted identifier; nothing escapes + // inside them. + {"with curly-quoted literal holding a paren then insert", "WITH x AS (SELECT \u2018(\u2019 AS s) INSERT INTO t SELECT s FROM x", true, false}, + {"with curly-quoted literal holding a verb then select", "WITH x AS (SELECT \u2018) INSERT\u2019 AS s) SELECT s FROM x", false, false}, + {"with curly-quoted identifier holding a paren then insert", "WITH x AS (SELECT 'q' AS \u201cc(d\u201d) INSERT INTO t SELECT * FROM x", true, false}, + + // ClickHouse's lexer skips \v, \f and Unicode spaces as whitespace + // (TestIsMutation_ClickHouseWhitespace covers the whole set). + {"leading form feed then insert", "\fINSERT INTO t VALUES (1)", true, false}, + {"leading vertical tab then insert", "\vINSERT INTO t VALUES (1)", true, false}, + {"leading NBSP then insert", "\u00a0INSERT INTO t VALUES (1)", true, false}, + {"leading BOM then insert", "\ufeffINSERT INTO t VALUES (1)", true, false}, + {"leading NBSP then select", "\u00a0SELECT 1", false, false}, + {"with CTE alias named set before form feed AS (read)", "WITH set\fAS (SELECT 1) SELECT * FROM set", false, false}, + {"with CTE alias named set before NBSP AS (read)", "WITH set\u00a0AS (SELECT 1) SELECT * FROM set", false, false}, + + // ClickHouse block comments nest. + {"nested block comment hiding select then insert", "/* a /* b */ SELECT */ INSERT INTO t VALUES (1)", true, false}, + {"nested block comment hiding insert then select", "/* a /* b */ INSERT */ SELECT 1", false, false}, + {"unclosed nested block comment", "/* a /* b */ INSERT INTO t VALUES (1)", false, true}, + {"with nested block comment hiding select then insert", "WITH x AS (SELECT 1) /* a /* b */ SELECT */ INSERT INTO t SELECT * FROM x", true, false}, + {"with nested block comment hiding insert then select", "WITH x AS (SELECT 1) /* a /* b */ INSERT */ SELECT * FROM x", false, false}, + {"with nested block comment before CTE AS (read)", "WITH set /* a /* b */ c */ AS (SELECT 1) SELECT * FROM set", false, false}, + + // After a WITH list ClickHouse parses only SELECT, a FROM-first SELECT and + // INSERT INTO; any other statement is a syntax error, so nothing runs. + {"with delete", "WITH cte AS (SELECT id FROM x) DELETE FROM t WHERE id IN (SELECT id FROM cte)", false, true}, + {"with alter update", "WITH cte AS (SELECT 1) ALTER TABLE t UPDATE a=1 WHERE id IN (SELECT id FROM cte)", false, true}, + {"with truncate", "WITH cte AS (SELECT 1) TRUNCATE TABLE t", false, true}, + {"with CTE alias named update then alter update", "WITH update AS (SELECT id FROM x) ALTER TABLE other UPDATE c=1 WHERE id IN (SELECT id FROM update)", false, true}, + {"with from-first select", "WITH 1 AS x FROM system.one SELECT x", false, false}, + {"with from-first select from a subquery", "WITH 1 AS x FROM (SELECT 1) SELECT x", false, false}, + + // A WITH list's names may be spelled like any keyword: a CTE name, an + // alias, a function, a lambda parameter, an operand, a qualified name's + // part, an array element or a bare element. + {"with alias desc then insert", "WITH 'd' AS desc INSERT INTO t SELECT length(desc)", true, false}, + {"with CTE named desc then insert", "WITH desc AS (SELECT 1 AS x) INSERT INTO t SELECT * FROM desc", true, false}, + {"with CTE named check then insert", "WITH check AS (SELECT 1 AS x) INSERT INTO t SELECT * FROM check", true, false}, + {"with CTE named explain then insert", "WITH explain AS (SELECT 1 AS x) INSERT INTO t SELECT * FROM explain", true, false}, + {"with alias show then insert", "WITH 1 AS show INSERT INTO t SELECT show", true, false}, + {"with alias describe then insert", "WITH 1 AS describe INSERT INTO t SELECT describe", true, false}, + {"with EXISTS expression then insert", "WITH EXISTS(SELECT 1 FROM t WHERE x = 7) AS seen INSERT INTO t2 SELECT seen", true, false}, + {"with alias select then insert", "WITH 1 AS select INSERT INTO t SELECT 1", true, false}, + {"with operand select then insert", "WITH 1 + select AS y INSERT INTO t SELECT y", true, false}, + {"with qualified select then insert", "WITH t.select AS y INSERT INTO t SELECT y", true, false}, + {"with array of select then insert", "WITH [select] AS a INSERT INTO t SELECT a", true, false}, + {"with bare element from then insert", "WITH from INSERT INTO t SELECT 1", true, false}, + {"with comment between INSERT and INTO", "WITH 1 AS x INSERT /* c */ INTO t SELECT x", true, false}, + {"with alias set (read)", "WITH 1 AS set SELECT set", false, false}, + {"with alias use (read)", "WITH 1 AS use SELECT use", false, false}, + {"with alias kill (read)", "WITH 1 AS kill SELECT kill", false, false}, + {"with alias system (read)", "WITH 1 AS system SELECT system", false, false}, + {"with alias insert (read)", "WITH 1 AS insert SELECT insert", false, false}, + {"with alias alter then from-first select", "WITH 1 AS alter FROM system.one SELECT alter", false, false}, + {"with lambda parameter set (read)", "WITH set -> 1 AS f SELECT f(2)", false, false}, + {"with lambda parameter insert (read)", "WITH insert -> 1 AS f SELECT f(2)", false, false}, + {"with operand insert (read)", "WITH insert + 1 AS y SELECT y", false, false}, + {"with bare element insert (read)", "WITH insert SELECT 1", false, false}, + {"with function insert (read)", "WITH insert(1) AS y SELECT y", false, false}, + + {"empty", "", false, true}, + {"comment only", "-- just a comment", false, true}, + {"unclosed block comment", "/* never closed", false, true}, +} diff --git a/tests/integration/ismutation_test.go b/tests/integration/ismutation_test.go new file mode 100644 index 000000000..8e36cb297 --- /dev/null +++ b/tests/integration/ismutation_test.go @@ -0,0 +1,200 @@ +//go:build integration + +package tests + +import ( + "context" + "errors" + "fmt" + "sort" + "strings" + "sync" + "testing" + + "github.com/ClickHouse/clickhouse-go/v2/lib/driver" + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + + "github.com/Wave-RF/WaveHouse/internal/api" + "github.com/Wave-RF/WaveHouse/internal/chconn" + "github.com/Wave-RF/WaveHouse/internal/testutil/mutationtest" +) + +// codeSyntaxError is ClickHouse's SYNTAX_ERROR. +const codeSyntaxError = 62 + +// astMutation is api.IsMutation's answer for each statement kind EXPLAIN AST +// names at the root of the tree. +var astMutation = map[string]bool{ + "SelectWithUnionQuery": false, + "ShowTables": false, + "DescribeQuery": false, + "Explain": false, + "ExistsTableQuery": false, + "InsertQuery": true, + "UpdateQuery": true, + "DeleteQuery": true, + "TruncateQuery": true, + "DropQuery": true, + "AlterQuery": true, + "CreateQuery": true, + "Rename": true, + "OptimizeQuery": true, + "GrantQuery": true, + "SYSTEM": true, + "AttachQuery": true, + "DetachQuery": true, + "KillQueryQuery": true, + "Set": true, + "UseQuery": true, +} + +// astRoot is the root node ClickHouse's parser gives sql, or the error it +// rejects sql with. Parsing only: nothing runs, and no table need exist. +func astRoot(ctx context.Context, conn driver.Conn, sql string) (string, error) { + rows, err := conn.Query(ctx, "EXPLAIN AST "+sql) + if err != nil { + return "", err + } + defer func() { _ = rows.Close() }() + if !rows.Next() { + if err := rows.Err(); err != nil { + return "", err + } + return "", errors.New("EXPLAIN AST returned no rows") + } + var line string + if err := rows.Scan(&line); err != nil { + return "", err + } + fields := strings.Fields(line) + if len(fields) == 0 { + return "", fmt.Errorf("EXPLAIN AST returned %q", line) + } + return fields[0], nil +} + +func isSyntaxError(err error) bool { + code, ok := chconn.ExceptionCode(err) + return ok && code == codeSyntaxError +} + +// TestIsMutation_AgreesWithClickHouseParser checks every shared IsMutation +// case against the parser of the pinned ClickHouse: a case that parses is +// classified as its statement kind, and one marked unparsed still fails to. +func TestIsMutation_AgreesWithClickHouseParser(t *testing.T) { + e := env(t) + ctx := context.Background() + for _, tc := range mutationtest.Cases { + t.Run(tc.Name, func(t *testing.T) { + root, err := astRoot(ctx, e.chConn, tc.SQL) + if tc.Unparsed { + require.Error(t, err, "ClickHouse now parses this case as %s: set Mutation to match and drop Unparsed", root) + assert.True(t, isSyntaxError(err), "want a syntax error, got %v", err) + return + } + require.NoError(t, err) + want, known := astMutation[root] + require.True(t, known, "no IsMutation answer for statement kind %s: add it to astMutation", root) + assert.Equal(t, want, tc.Mutation, "ClickHouse parses this as %s", root) + }) + } +} + +// TestIsMutation_KeywordNamesInWithList spells a WITH list's names — a CTE, +// an alias, a function, a lambda parameter, an operand, a qualified name's +// part, an array element, a bare element — as every ClickHouse keyword, ahead +// of each statement a WITH list can lead, and checks IsMutation against the +// parser on each combination that parses. +func TestIsMutation_KeywordNamesInWithList(t *testing.T) { + e := env(t) + ctx := context.Background() + + rows, err := e.chConn.Query(ctx, "SELECT keyword FROM system.keywords") + require.NoError(t, err) + seen := map[string]bool{} + for rows.Next() { + var k string + require.NoError(t, rows.Scan(&k)) + for _, w := range strings.Fields(k) { + seen[strings.ToLower(w)] = true + } + } + require.NoError(t, rows.Err()) + _ = rows.Close() + words := make([]string, 0, len(seen)) + for w := range seen { + words = append(words, w) + } + sort.Strings(words) + require.NotEmpty(t, words) + + shapes := []string{ + "WITH %s AS (SELECT 1 AS x) %s", + "WITH 1 AS a, %s AS (SELECT 1) %s", + "WITH 1 AS %s %s", + "WITH (SELECT 1) AS a, 2 AS %s %s", + "WITH %s AS y %s", + "WITH %s(1) AS y %s", + "WITH %s -> 1 AS f %s", + "WITH 1 + %s AS y %s", + "WITH %s + 1 AS y %s", + "WITH t.%s AS y %s", + "WITH [%s] AS a %s", + "WITH %s %s", + } + statements := []string{"SELECT 1", "FROM system.one SELECT 1", "INSERT INTO t SELECT 1"} + queries := make(chan string) + go func() { + defer close(queries) + for _, w := range words { + for _, shape := range shapes { + for _, stmt := range statements { + queries <- fmt.Sprintf(shape, w, stmt) + } + } + } + }() + + var ( + mu sync.Mutex + wg sync.WaitGroup + parsed int + failures []string + ) + for range 8 { + wg.Add(1) + go func() { + defer wg.Done() + for sql := range queries { + root, err := astRoot(ctx, e.chConn, sql) + var failure string + switch want, known := astMutation[root]; { + case err != nil && !isSyntaxError(err): + failure = fmt.Sprintf("%q: %v", sql, err) + case err != nil: + case !known: + failure = fmt.Sprintf("%q: no IsMutation answer for statement kind %s", sql, root) + case api.IsMutation(sql) != want: + failure = fmt.Sprintf("%q: ClickHouse parses it as %s, IsMutation says %v", sql, root, !want) + } + mu.Lock() + if err == nil { + parsed++ + } + if failure != "" { + failures = append(failures, failure) + } + mu.Unlock() + } + }() + } + wg.Wait() + + sort.Strings(failures) + for _, f := range failures { + t.Error(f) + } + // Most combinations parse; far fewer means the check stopped checking. + assert.Greater(t, parsed, len(words)*len(shapes)*len(statements)/2) +} From e8f65ac6fe9adcf04b24160bb8946dbf91d7e9ed Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:53:16 -0400 Subject: [PATCH 62/79] docs(pipes): what a write pipe skips, bounds and costs - architecture.md: an operator-authored write pipe is the other write path besides the admin raw-SQL proxy. - pipes.mdx: no write invalidates a read pipe's cached result (#343), and a write led by a verb the classifier does not know runs as a cached read (#666). - settings-directory.mdx: query_timeout bounds every call on the query paths, write pipes and /v1/ops/query included; the HTTP proxy sends no max_execution_time. Same for the settings and wire comments. - api.md: only a read pipe is cached and coalesced. - ch_errors.go: a write pipe's unavailable or unknown failure is never retryable. - CHANGELOG: the quoted-placeholder correction moves under Security, with what to check in existing templates (#662). Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- CHANGELOG.md | 5 +++-- docs/src/content/docs/api.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/pipes.mdx | 4 ++-- docs/src/content/docs/settings-directory.mdx | 2 +- internal/api/ch_errors.go | 6 ++++-- internal/app/wire.go | 4 ++-- internal/settings/settings.go | 3 ++- 8 files changed, 16 insertions(+), 12 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 8b2836db1..0ab4465af 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -85,9 +85,8 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed -- **A pipe that writes runs on every call instead of being answered from the cache** (`internal/api/{pipes,ch_errors}.go` (+ tests), `docs/src/content/docs/{pipes.mdx,api.md,architecture.md,configuration.mdx,ingest-pipeline.md,sdk/pipes.md,sdk/reference.md}`, `clients/ts/src/pipes.ts` (doc comment), `AGENTS.md`): fixes [#386](https://github.com/Wave-RF/WaveHouse/issues/386). `/v1/pipes/{name}` sent a write's SQL to ClickHouse through `Exec`, but still cached the `[]` it returned and coalesced identical calls in flight, so a repeat within the TTL answered `200` without executing and concurrent identical calls became one write — silently dropped writes, and with a shared cache ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)) on every instance. A pipe whose bound SQL `IsMutation` classifies as a write — the same classifier that picks `Exec` — now skips the cache lookup, the fill and singleflight, and answers `X-Cache: BYPASS` with `Cache-Control: no-store`, so an HTTP cache in front of a `GET` cannot drop the write either. Classification stays automatic rather than a declared pipe property, so an operator cannot forget to mark one, and costs no ClickHouse round trip. A failed write answers with the status and `code` a failed read gets (see the ClickHouse-errors entry below), but always `retryable: false` and with no `Retry-After`, `503 clickhouse.unavailable` included: the statement may have run, so the SDK does not retry it. A write refused before it is sent, the tenant on no pool, keeps its `503` with `Retry-After: 30`. Read pipes are unchanged. Not in this fix: a write pipe still does not invalidate cached reads of the table it writes ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). +- **A pipe that writes runs on every call instead of being answered from the cache** (`internal/api/{pipes,ch_errors}.go` (+ tests), `docs/src/content/docs/{pipes.mdx,api.md,architecture.md,configuration.mdx,settings-directory.mdx,ingest-pipeline.md,sdk/pipes.md,sdk/reference.md}`, `clients/ts/src/pipes.ts` (doc comment), `internal/{settings/settings,app/wire}.go` (comments), `AGENTS.md`): fixes [#386](https://github.com/Wave-RF/WaveHouse/issues/386). `/v1/pipes/{name}` sent a write's SQL to ClickHouse through `Exec`, but still cached the `[]` it returned and coalesced identical calls in flight, so a repeat within the TTL answered `200` without executing and concurrent identical calls became one write — silently dropped writes, and with a shared cache ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)) on every instance. A pipe whose bound SQL `IsMutation` classifies as a write — the same classifier that picks `Exec` — now skips the cache lookup, the fill and singleflight, and answers `X-Cache: BYPASS` with `Cache-Control: no-store`, so an HTTP cache in front of a `GET` cannot drop the write either. Classification stays automatic rather than a declared pipe property, so an operator cannot forget to mark one, and costs no ClickHouse round trip. A failed write answers with the status and `code` a failed read gets (see the ClickHouse-errors entry below), but always `retryable: false` and with no `Retry-After`, `503 clickhouse.unavailable` included: the statement may have run, so the SDK does not retry it. A write refused before it is sent, the tenant on no pool, keeps its `503` with `Retry-After: 30`. Read pipes are unchanged. Not in this fix: a write pipe still does not invalidate cached reads of the table it writes ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). - **The write classifier skips whitespace, comments and quoted text the way ClickHouse's lexer does** (`internal/api/clickhouse_exec.go` (+ tests), `internal/testutil/mutationtest` (new), `tests/integration/ismutation_test.go` (new), `AGENTS.md`): `IsMutation` picks `Exec` for a write, and since [#386](https://github.com/Wave-RF/WaveHouse/issues/386) keeps a write pipe out of the cache. It missed a write behind a backslash-escaped quote (`'it\'s'`, and the same inside `"…"` and `` `…` ``), a heredoc (`$$ ( $$`, `$tag$ … $tag$`), a curly-quoted literal or identifier (`‘(’`, `“c(d”`), a `//` line comment, a nested block comment (`/* a /* b */ SELECT */ INSERT …`), or leading whitespace other than space, tab, CR and LF: `\v`, `\f`, a no-break space, a byte-order mark, and the other Unicode spaces ClickHouse skips. A missed write went through `Query`, which ran it and then failed the call with a `5xx` the TypeScript SDK retries, so one call could write three times. The same gaps, and a word led by `_` (`_delete`) whose tail was read as a verb, could make a read look like a write, which runs through `Exec` and answers `[]`. After a `WITH` list, which ClickHouse follows only with `SELECT`, a FROM-first `SELECT` or `INSERT INTO`, a name spelled like a keyword was taken for the statement: `WITH 'd' AS desc INSERT …` and `WITH 1 AS select INSERT …` ran as reads, and `WITH 1 AS set SELECT set` and `WITH 1 AS x FROM system.one SELECT x` as writes. A `WITH`-led statement is now a write exactly when it holds `INSERT INTO` outside parentheses. The classifier, exported as `IsMutation` for it, is now checked against the pinned ClickHouse's own parser (`EXPLAIN AST`) in the integration suite: every test case, and every ClickHouse keyword as a `WITH` list's name ahead of each statement a `WITH` list can lead. -- **The pipes page no longer says a parameter can never break out of its literal** (`docs/src/content/docs/pipes.mdx`, `internal/pipes/pipes.go`): that holds only for a placeholder written bare. A string value brings its own quotes, so inside a quoted placeholder they close the template's: the body `{"id": " OR 1=1 OR id = "}` turns `WHERE id = '{{id}}'` into `WHERE id = '' OR 1=1 OR id = ''`, which matches every row. The page now says to write each placeholder bare, never inside quotes. - **A failed ClickHouse query answers by what went wrong, not a flat `500`/`502`** (`internal/api/ch_errors.go` (new, + tests), `internal/api/{errors,query,structured_query,pipes,schema,ch_settings}.go`, `internal/chconn/errclass.go` (`HTTPStatus` exported), `clients/ts/src/errors.ts` (+ tests), `tests/integration/query_errors_test.go` (new), `tests/integration/query_limits_test.go`, `internal/app/app_test.go`, `tests/e2e/sdk/{admin,query}.test.ts`, `AGENTS.md`, `docs/src/content/docs/{api,architecture}.md`, `docs/src/content/docs/{access-control,configuration}.mdx`, `docs/src/content/docs/sdk/{reference.md,index.mdx}`): fixes [#403](https://github.com/Wave-RF/WaveHouse/issues/403) and [#271](https://github.com/Wave-RF/WaveHouse/issues/271), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). ClickHouse answers a syntax error, a missing grant and an overloaded server alike with HTTP `500`, so `/v1/ops/query` turned a bad statement into a `502` and `/v1/query` and pipes into a `500` the SDK retried. All three now class the failure with `chconn.Classify` through one helper, `writeCHError`: a statement ClickHouse refused is `400 clickhouse.rejected`; a query over a rows/bytes limit, the role's own memory cap, or its time cap where that is no longer than `query_timeout` is `400 clickhouse.limit_exceeded`; `ACCESS_DENIED` is `403 clickhouse.access_denied`; credentials, user or database refused, or a redirect or `4xx` with no exception code from whatever fronts ClickHouse, is `502 clickhouse.misconfigured`; ClickHouse down, unreachable or overloaded is `503 clickhouse.unavailable` with `Retry-After: 5`; a failure with no verdict stays `500` (`502` on the proxy) as `clickhouse.unknown`. The error envelope gains `code` and `retryable` next to `error` on these responses — additive. A role with `max_execution_time` now queries with no context deadline and a cancel two seconds past the cap instead: clickhouse-go overwrote the cap's `max_execution_time` with deadline+5s for any deadline over 1s, so an overrun came back as a bare deadline, indistinguishable from waiting for a pooled connection; ClickHouse now enforces the cap itself and reports `TIMEOUT_EXCEEDED`. `POST /v1/ops/schema/refresh` against an unreachable ClickHouse is a `503` with `Retry-After` instead of a `500`. **SDK:** `WaveHouseError.code` and `retryable` now take the server's `code`/`retryable` when the body has them (`HTTP_` and "5xx retries" otherwise), so a rejected query is `clickhouse.rejected` rather than `HTTP_500`, and is not retried. - **An unavailable ClickHouse is retried with backoff instead of dead-lettering every row** (`internal/chconn/errclass.go` (new, + tests), `internal/ingest/{worker,backoff}.go` (`backoff.go` new, + tests), `internal/mq/{mq,embedded}.go`, `internal/testutil/mocks.go`, `tests/integration/ingest_outage_test.go` (new), `AGENTS.md`, `README.md`, `docs/src/content/docs/{ingest-pipeline,architecture,api,deployment,why-wavehouse}.md`, `docs/src/content/docs/{settings-directory,index,access-control}.mdx`): workstream A of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). A failed batch insert used to go through row-by-row isolation whatever the failure, so a ClickHouse that was down, overloaded or read-only failed every row twice and parked the whole batch on the DLQ. `chconn.Classify` now classes the failure first — `Rejected` (any ClickHouse exception code outside the availability and credential lists: the server read the row and refused it), `Unavailable` (connection refused/reset, timeouts, `TOO_MANY_SIMULTANEOUS_QUERIES`, `SERVER_OVERLOADED`, `MEMORY_LIMIT_EXCEEDED`, `TOO_MANY_PARTS`, `READONLY`, `TABLE_IS_READ_ONLY`, `KEEPER_EXCEPTION`, …), `Denied` (`AUTHENTICATION_FAILED`, `ACCESS_DENIED`, …) or `Unknown` (no code, no recognizable transport failure). Only `Rejected` is isolated and dead-lettered as before, and a multi-row batch refused with `TOO_MANY_PARTS` or `MEMORY_LIMIT_EXCEEDED` is split row by row first (`chconn.Splittable`), because a batch spanning too many partitions or too much memory can fail where each of its rows inserts; every other class hands the batch back to the queue with a delayed nak (`mq.Message.NakWithDelay`, new) under a jittered 1 s → 30 s backoff shared by every table on the same ClickHouse pool (a failure of one table — read-only, too many parts or mutations, a grant missing on it, `chconn.TableScoped` — backs off that table alone), which turns rows away without a request while it runs and probes once per window, and ClickHouse going away mid-isolation stops isolation and retries the rows it had not settled. Counted by the new `wavehouse_ingest_retries_total{table, reason}`; logged at `WARN` when an outage starts and at most every 30 s during it. A long outage now shows as a growing ingest stream and, at `mq.max_bytes_gb`, ingest `503`s — not as a full DLQ; a lasting failure of one table holds back its tenant's other tables once its waiting rows reach `maxAckPending`. Retried rows come back out of arrival order, which matters only to a `ReplacingMergeTree` without a version column or a `CollapsingMergeTree`. - **Schema discovery's retry loop jitters its backoff** (`internal/discovery/discovery.go` (+ tests), `internal/app/wire.go`, `internal/api/errors.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`): `RetryRefresh` slept exactly `2s * 2^n` capped at 60s, so instances retrying against one recovering ClickHouse fired in lockstep, every 60s on the same second. Each sleep is now drawn uniformly from below the backoff (full jitter), spreading the retries over the whole window and halving the mean wait — so a failing tenant's retries, their log lines and `wavehouse_schema_refresh_failures_total` come about twice as often ([#141](https://github.com/Wave-RF/WaveHouse/issues/141)). @@ -109,6 +108,8 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Security +- **The pipes page no longer says a parameter can never break out of its literal** (`docs/src/content/docs/pipes.mdx`, `internal/pipes/pipes.go`): that holds only for a placeholder written bare. A string value brings its own quotes, so inside a quoted placeholder they close the template's: the body `{"id": " OR 1=1 OR id = "}` turns `WHERE id = '{{id}}'` into `WHERE id = '' OR 1=1 OR id = ''`, which matches every row. The page now says to write each placeholder bare, never inside quotes. Check existing `pipes.json` templates for quoted placeholders (`'{{x}}'`) and write them bare ([#662](https://github.com/Wave-RF/WaveHouse/issues/662)). + - **An empty HMAC secret no longer verifies tokens signed with an empty key** (`internal/auth/auth.go` (+ tests), `SECURITY.md`): with `auth.jwt_secret` unset and no `auth.jwks_url` — the documented public-access posture, "no token can validate" — the key function handed `golang-jwt` an empty HMAC key, and the library verifies a token signed with one, so anyone could mint `{"role": "admin"}` and reach the whole data plane and `/v1/ops/*`. The verifier now refuses every token when it has neither a secret nor a JWKS URL, pinned by a test that signs with the empty key. Found by review on [#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 9 ([#604](https://github.com/Wave-RF/WaveHouse/pull/604)), which carries the same fix. - **Policy validation now rejects the fail-open rule shapes strict decoding can't see** (`internal/policy/policy.go`, `docs/src/content/docs/access-control.mdx`; closes [#460](https://github.com/Wave-RF/WaveHouse/issues/460)): four new `validateRolePerms` rejections close the fail-open shapes strict decoding can't see because the document is syntactically innocent. A `filter` entry with no operator (`"tenant_id": {}`) resolved to zero predicates — no `WHERE` clause, row security silently off, the same shape a misspelled `"eq"` for `"_eq"` used to decode to before strict decoding closed that route; it is now rejected, as is its check-path twin (an operator-less `check` entry, skipped by `Evaluate`'s resolve switch — accepted but constraining nothing) and `filter:` under an `insert:` grant (resolved and then ignored by the ingest path — the same accept-but-ignore family as [#224](https://github.com/Wave-RF/WaveHouse/issues/224), and the pointed asymmetry #460 called out against the loud `check` `_neq`/`_gt`/`_lt` rejection) along with its mirror, `check:` under a `select:` grant — the likelier authoring slip and the fail-open direction: the author believes reads are row-scoped while `Evaluate` resolves the entry and nothing on the select or stream paths reads it. Because [#508](https://github.com/Wave-RF/WaveHouse/pull/508) funneled every adoption through the one `policy.Validate` path, the four checks land on boot, the directory watch, `SIGHUP`, `POST /v1/ops/settings/reload`, and `wavehouse validate` at once. #460's migration caveat (a stored policy hard-failing at boot) has evaporated with the settings directory being new and unreleased; no shipped seed, compose, or fixture policy carries any of the rejected shapes. diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index 2076e7340..8b82643e9 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -587,7 +587,7 @@ The inbound request body is capped at 1 MiB; a body over the cap is rejected wit ### `GET/POST /v1/pipes/{name}` — Execute Named Pipe -Executes a pre-defined named query (pipe) with parameter binding. Parameters can be supplied via query string and/or JSON body. Results are cached in the shared L1 (Ristretto) with singleflight coalescing — same machinery as the structured query endpoint, keyed by [tenant](/deployment#multi-tenant-deployments) like it, and again, unlike `/v1/ops/query`. +Executes a pre-defined named query (pipe) with parameter binding. Parameters can be supplied via query string and/or JSON body. A read's results are cached in the query cache with singleflight coalescing — same machinery as the structured query endpoint, keyed by [tenant](/deployment#multi-tenant-deployments) like it, and again, unlike `/v1/ops/query`; a [pipe that writes](/pipes#pipes-that-write) is neither cached nor coalesced (see Response). **Query Parameters:** Any key matching a pipe parameter name. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index ab4c00e92..26e2aac24 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -148,7 +148,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ ### `ingest/` — Ingest Pipeline, DLQ & Sweeping -- **worker.go** — `StartIngestWorker` launches an ingest pipeline: a durable `buffer-consumer` consumer of the ingest queue (created through `mq.ConsumerManager`) reads events, batches them per tenant table — the tenant read off each message's `mq.Topic` — and performs bulk INSERTs to ClickHouse. The pipeline is **insert-only**. The wire format `EventMessage` carries `{table_name, scope, received_timestamp, format, columns, row}` — the row positionally as one `JSONCompactEachRow` line, with `columns` naming its positions (the table's insertable columns — a computed one cannot be named in an `INSERT`); the worker batches per (tenant, table, column list) and writes `INSERT INTO … (cols) FORMAT JSONCompactEachRow`. It accepts any table name (events are addressed by `mq.Topic{Tenant, Table, Scope}` with raw names; `internal/mq` encodes them into subject tokens), then bulk-INSERTs. The embedded NATS server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only publishers that can reach the ingest queue are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (`policy.admin_role`) — see the Query Path section below; the `/v1/ops/*` `RequireAdmin` middleware enforces the check at the API layer, so a no/invalid-token request (resolved to `default_role`, not admin in a production config) never reaches the proxy. A batch whose tenant has no ClickHouse connection (no longer served, or no pool could be opened for it, such as by the connection ceiling) is never tried: no row of it could pass, so `parkBatch` takes it to the DLQ switch whole, logging once per batch rather than twice per row. Otherwise a bulk-insert failure is first classed by `chconn.Classify`: a ClickHouse that cannot take the insert (unavailable, denied, or no verdict at all) sends the batch back to the MQ for a delayed redelivery (`retryLater` → `mq.Message.NakWithDelay`), under a backoff shared by every table on the same pool (a failure of one table — read-only, too many parts — backs off that table alone), and never to the DLQ — the same when it stops answering mid-isolation. Only when ClickHouse rejects the batch, or refuses a multi-row batch for its size (`chconn.Splittable`: too many partitions for one INSERT, the memory limit), is it re-inserted row by row: rows that succeed are acked, and only the rows ClickHouse rejects again are routed to the DLQ (`sendToDLQ` → `mq.DeadLetterer.DeadLetter`), which parks the as-published `EventMessage` envelope under the topic it arrived on (`dlq.{tenant}.{table}` subjects inside `internal/mq`) with the failure context in `X-DLQ-*` headers when the tenant's `dlq.enabled` is on for the table — see [Ingest Pipeline](/ingest-pipeline) for the worker internals. +- **worker.go** — `StartIngestWorker` launches an ingest pipeline: a durable `buffer-consumer` consumer of the ingest queue (created through `mq.ConsumerManager`) reads events, batches them per tenant table — the tenant read off each message's `mq.Topic` — and performs bulk INSERTs to ClickHouse. The pipeline is **insert-only**. The wire format `EventMessage` carries `{table_name, scope, received_timestamp, format, columns, row}` — the row positionally as one `JSONCompactEachRow` line, with `columns` naming its positions (the table's insertable columns — a computed one cannot be named in an `INSERT`); the worker batches per (tenant, table, column list) and writes `INSERT INTO … (cols) FORMAT JSONCompactEachRow`. It accepts any table name (events are addressed by `mq.Topic{Tenant, Table, Scope}` with raw names; `internal/mq` encodes them into subject tokens), then bulk-INSERTs. The embedded NATS server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only publishers that can reach the ingest queue are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (`policy.admin_role`) — see the Query Path section below; the `/v1/ops/*` `RequireAdmin` middleware enforces the check at the API layer, so a no/invalid-token request (resolved to `default_role`, not admin in a production config) never reaches the proxy; the only other write path is an operator-authored [pipe that writes](/pipes#pipes-that-write). A batch whose tenant has no ClickHouse connection (no longer served, or no pool could be opened for it, such as by the connection ceiling) is never tried: no row of it could pass, so `parkBatch` takes it to the DLQ switch whole, logging once per batch rather than twice per row. Otherwise a bulk-insert failure is first classed by `chconn.Classify`: a ClickHouse that cannot take the insert (unavailable, denied, or no verdict at all) sends the batch back to the MQ for a delayed redelivery (`retryLater` → `mq.Message.NakWithDelay`), under a backoff shared by every table on the same pool (a failure of one table — read-only, too many parts — backs off that table alone), and never to the DLQ — the same when it stops answering mid-isolation. Only when ClickHouse rejects the batch, or refuses a multi-row batch for its size (`chconn.Splittable`: too many partitions for one INSERT, the memory limit), is it re-inserted row by row: rows that succeed are acked, and only the rows ClickHouse rejects again are routed to the DLQ (`sendToDLQ` → `mq.DeadLetterer.DeadLetter`), which parks the as-published `EventMessage` envelope under the topic it arrived on (`dlq.{tenant}.{table}` subjects inside `internal/mq`) with the failure context in `X-DLQ-*` headers when the tenant's `dlq.enabled` is on for the table — see [Ingest Pipeline](/ingest-pipeline) for the worker internals. - **backoff.go** — The retry backoff behind `retryLater`: a small circuit breaker per ClickHouse pool (the target's URL, user and database), and one per pool and table for a failure of one table (`chconn.TableScoped`). A failure opens it for 1 s, doubling to a 30 s cap, each window jittered down to half; while it is open, flushes and arriving rows are handed back without a request, and once it elapses one flush probes. Any answer that is not an outage closes it. - **types.go** — `EventMessage` struct (TableName, Scope — reserved, always empty today, ReceivedTimestamp, Format, Columns, Row; `Format` is `FormatJSONCompactEachRow` and `Row` is one positional line whose slots `Columns` names) and `BufferConsumerName` constant, shared across API handlers and the ingest pipeline. - **compact.go** — `EncodeCompactRow`, the positional row encoder every published row goes through, rendering one record over the table's **insertable** columns in declaration order. Serialization only: it validates nothing and judges no value. diff --git a/docs/src/content/docs/pipes.mdx b/docs/src/content/docs/pipes.mdx index 9a916730e..34b8060dc 100644 --- a/docs/src/content/docs/pipes.mdx +++ b/docs/src/content/docs/pipes.mdx @@ -180,13 +180,13 @@ The response is a JSON array of rows. Results flow through the shared in-process ### Pipes that write -A pipe's SQL may be a write: a statement led by a write verb WaveHouse recognizes — `INSERT`, `UPDATE`, `DELETE`, `ALTER` (so `ALTER … DELETE`), `CREATE`, `DROP`, `TRUNCATE`, `RENAME`, `EXCHANGE`, `REPLACE`, `OPTIMIZE`, `ATTACH`, `DETACH`, `GRANT`, `REVOKE`, `KILL`, `SET`, `USE` or `SYSTEM` — directly or after a `WITH` list, as in `WITH … INSERT`. Such a pipe runs on **every** call: it never reads or fills the cache and is never coalesced with an identical call in flight, so ten identical calls are ten writes. The response is `[]` with `X-Cache: BYPASS` and `Cache-Control: no-store`, so an HTTP cache in front of a `GET` does not answer a repeat either. WaveHouse classifies the statement from its leading keyword (after any `WITH` list), with the same classifier that sends it to ClickHouse as a write, so no pipe property marks it; a statement led by any other keyword runs as a read. +A pipe's SQL may be a write: a statement led by a write verb WaveHouse recognizes — `INSERT`, `UPDATE`, `DELETE`, `ALTER` (so `ALTER … DELETE`), `CREATE`, `DROP`, `TRUNCATE`, `RENAME`, `EXCHANGE`, `REPLACE`, `OPTIMIZE`, `ATTACH`, `DETACH`, `GRANT`, `REVOKE`, `KILL`, `SET`, `USE` or `SYSTEM` — directly or after a `WITH` list, as in `WITH … INSERT`. Such a pipe runs on **every** call: it never reads or fills the cache and is never coalesced with an identical call in flight, so ten identical calls are ten writes. The response is `[]` with `X-Cache: BYPASS` and `Cache-Control: no-store`, so an HTTP cache in front of a `GET` does not answer a repeat either. WaveHouse classifies the statement from its leading keyword (after any `WITH` list), with the same classifier that sends it to ClickHouse as a write, so no pipe property marks it; a statement led by any other keyword runs as a read — cached and coalesced like any read, so a repeat within the TTL does not run: don't put a write led by another verb (`BACKUP`, `RESTORE`, `UNDROP`) in a pipe ([#666](https://github.com/Wave-RF/WaveHouse/issues/666)). A failed write is not retried automatically, because it may have run. It answers with the status and `code` a failed read would ([ClickHouse errors on the query paths](/api#clickhouse-errors-on-the-query-paths)), but always with `retryable: false` and no `Retry-After`, `503 clickhouse.unavailable` included: once the statement is on its way to ClickHouse, WaveHouse cannot tell whether it ran. The [SDK](/sdk/pipes) does not retry such an answer, so check whether the write landed before you send it again. A call refused before anything is sent — the tenant on no ClickHouse pool, `503` with `Retry-After: 30` — cannot have run, and the SDK retries it. The SDK also retries when WaveHouse's own answer never reaches it — a dropped connection, or a `502`/`503`/`504` from a proxy in front of WaveHouse that gave up waiting — so a write can still run twice that way; give a client that runs write pipes [`options.maxRetries`](/sdk#clientconfigdb) `0` if that matters. `allowed_roles` is a write pipe's only gate: the [policy engine](/access-control)'s insert rules do not apply to it, so any role you list — including a [`default_role`](/access-control#default_role--public-unauthenticated-access) that anonymous callers resolve to — can run the write. The operator fixes the statement and its predicate when authoring the pipe; callers supply only literal values, provided every placeholder is written bare ([how a value becomes SQL](#how-a-value-becomes-sql)). -Two things a write pipe does not do yet: it does not invalidate cached reads of the table it writes — a structured query or read pipe over that table can serve pre-write rows until its TTL ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)) — and its rows do not reach [`/v1/stream`](/api#get-v1stream--server-sent-events-stream) subscribers, which only the [ingest pipeline](/ingest-pipeline) feeds ([#362](https://github.com/Wave-RF/WaveHouse/issues/362)). For writes that should be seen at once, use [`POST /v1/ingest`](/api#post-v1ingesttabletable--ingest-data). +Two things a write pipe does not do yet: it does not invalidate cached reads of the table it writes — a structured query over that table can serve pre-write rows until its TTL ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)); a read pipe's cached result expires only with its TTL whatever writes the table, ingest included ([#343](https://github.com/Wave-RF/WaveHouse/issues/343)) — and its rows do not reach [`/v1/stream`](/api#get-v1stream--server-sent-events-stream) subscribers, which only the [ingest pipeline](/ingest-pipeline) feeds ([#362](https://github.com/Wave-RF/WaveHouse/issues/362)). For writes that stream subscribers and structured queries should see at once, use [`POST /v1/ingest`](/api#post-v1ingesttabletable--ingest-data). ## End-to-end example diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index 900e6a2c2..67992b6fb 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -108,7 +108,7 @@ The tenant tunables. Every key is required (a missing one is a validation error) | `clickhouse.http_scheme` | `http` | `http` or `https` for that HTTP hop — one of the two *outbound* TLS switches, with `tls.enabled` for the native hop; unrelated to your clients' TLS. | | `clickhouse.database` | `default` | Database tables are discovered from. | | `clickhouse.username` | `default` | Connection user; the password is boot config (`WH_CH_PASSWORD`). | -| `clickhouse.query_timeout` | `30` | Seconds (`>= 1`) a read may take. On `/v1/query` under a role's `max_execution_time`, the smaller of the two is sent to ClickHouse as `max_execution_time`; otherwise it bounds the client deadline, from which the driver derives a server-side `max_execution_time`. | +| `clickhouse.query_timeout` | `30` | Seconds (`>= 1`) a ClickHouse call may take on the query paths — structured queries, pipes (a write pipe included) and `/v1/ops/query`. On `/v1/query` under a role's `max_execution_time`, the smaller of the two is sent to ClickHouse as `max_execution_time`; otherwise, on the native paths, it bounds the client deadline, from which the driver derives a server-side `max_execution_time`; on `/v1/ops/query` it is the HTTP request's deadline. | | `clickhouse.tls.enabled` | `false` | Switches the native-protocol hop (`addr`) to TLS. The HTTP hop's switch stays `http_scheme`; the rest of the `tls` block applies to whichever hop uses TLS. See [ClickHouse](#clickhouse). | | `clickhouse.tls.ca_file` | `""` | PEM bundle the server certificate is verified against; empty uses the system roots. A path, read when the connection is built and re-read when the `tls` block changes — validation does not open it. | | `clickhouse.tls.cert_file` | `""` | Client certificate for mutual TLS, PEM; set together with `key_file` or not at all. | diff --git a/internal/api/ch_errors.go b/internal/api/ch_errors.go index bf5e3eb41..adb61aaf3 100644 --- a/internal/api/ch_errors.go +++ b/internal/api/ch_errors.go @@ -28,11 +28,13 @@ const ( // caller's. 502. codeCHMisconfigured = "clickhouse.misconfigured" // codeCHUnavailable: ClickHouse, or the way to it, could not take the - // query now. 503 with Retry-After. + // query now. 503 with Retry-After — without it for a write pipe, which + // may have run and is never retryable. codeCHUnavailable = "clickhouse.unavailable" // codeCHResponseTooLarge: the raw-SQL proxy's response cap. 502. codeCHResponseTooLarge = "clickhouse.response_too_large" - // codeCHUnknown: a failure with no verdict. 5xx, retryable. + // codeCHUnknown: a failure with no verdict. 5xx, retryable unless a + // write pipe's. codeCHUnknown = "clickhouse.unknown" ) diff --git a/internal/app/wire.go b/internal/app/wire.go index efcb44fc9..c6f7c7f97 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -367,8 +367,8 @@ func (a *App) registryFor(s *settings.Store) *discovery.SchemaRegistry { return a.discoveries.For(s.Tenant()) } -// queryTimeout is the tenant's read deadline, a per-call setting rather -// than a property of the pool it shares. +// queryTimeout is the tenant's deadline for a call on the query paths, a +// per-call setting rather than a property of the pool it shares. func queryTimeout(s *settings.Store) time.Duration { return s.ClickHouse().QueryTimeout } // wireDiscovery builds one schema registry per served tenant, each with a diff --git a/internal/settings/settings.go b/internal/settings/settings.go index c1e439238..9ad25b2e7 100644 --- a/internal/settings/settings.go +++ b/internal/settings/settings.go @@ -98,7 +98,8 @@ type ClickHouseConfig struct { HTTPScheme *string `json:"http_scheme"` Database *string `json:"database"` Username *string `json:"username"` - // QueryTimeout is the read deadline in seconds (>= 1). + // QueryTimeout is the deadline in seconds (>= 1) of a call on the query + // paths: structured queries, pipes (writes included) and the raw-SQL proxy. QueryTimeout *int `json:"query_timeout"` // TLS is the TLS wiring of both hops: `enabled` switches the native // protocol to TLS, `http_scheme` stays the HTTP hop's switch, and the From bf564035e5fca6ba55e6fcb9a4acd2e2d289b325 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:53:38 -0400 Subject: [PATCH 63/79] fix(cache): trip on rejected credentials; give the probe a dial budget - A password rotated under a running process refused every new connection's handshake with WRONGPASS, which counted as the server being up: the breaker stayed closed and nothing was logged. WRONGPASS and NOAUTH now open it at once, logged at ERROR. NOPERM still names one key or command and does not. - rueidis redials under the calling operation's context, so a probe bounded by Timeout could never complete a reconnect slower than it. The probe now runs under DialTimeout plus Timeout. It reconnects only the connection it lands on; the client's others still reconnect under the op timeout, and the docs say so (refs #664). - Restoring an RDB/AOF snapshot is a rollback, not a lost token: the docs no longer say a restart can only cause misses, and a test pins the rollback. - The stored-size cap is now tested: a unit count of both size limits, and an incompressible value between them on every server. - Docs: invalidations_total's ok and deferred overlap; the breaker.go and pending.go bullets describe their own files and name the redis.go functions around them; the cluster test covers CROSSSLOT and the topology read, not routing; the timing assertion is stated as written. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 8 +- internal/cache/breaker.go | 8 +- internal/cache/breaker_test.go | 5 +- internal/cache/metrics.go | 2 +- internal/cache/redis.go | 34 +++- internal/cache/redis_integration_test.go | 199 +++++++++++++++++++---- internal/cache/redis_test.go | 102 ++++++++---- 8 files changed, 279 insertions(+), 81 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 563dd2e76..fde8df608 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A Redis-compatible shared cache backend, built but not yet selectable** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/cache.go`, `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`, `CONTRIBUTING.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart can only cause misses — `maxmemory-policy allkeys-lru` is safe. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), open a circuit breaker that skips the server until a probe write succeeds (connections are replaced every minute, so after a failover behind a stable address the process reaches the new primary, and delivers the bumps it owes, within about a minute, provided a fresh connection's dial and handshake fit in the per-operation timeout), and deferred invalidations are retried until they land — the first at once, and at the probe's cadence while the breaker is open — collapsing to one tenant-wide bump per tenant past 100,000 keys; until one lands, the process that owes it bypasses the lookups it would orphan. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. Nothing selects it yet: `cache.backend` accepts only `local` until the wiring (E4) adds `redis` and the `cache.redis.*` settings, so every deployment still runs `LocalCache`. Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, compression, server-stops-answering, refused-writes (a demoted primary, a full `noeviction` server), failover-behind-a-stable-address and owed-invalidation cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. +- **A Redis-compatible shared cache backend, built but not yet selectable** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/cache.go`, `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`, `CONTRIBUTING.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart without persistence can only cause misses — `maxmemory-policy allkeys-lru` is safe. Restoring an RDB or AOF snapshot, or a backup, is a rollback instead: the old tokens return with their values, so invalidations made since are undone until those entries' TTL. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like) or the credentials (`WRONGPASS` or `NOAUTH` after a password rotation, logged at `ERROR`), open a circuit breaker that skips the server until a probe write succeeds (connections are replaced every minute, so after a failover behind a stable address the process reaches the new primary, and delivers the bumps it owes, within about a minute; the probe's budget covers a reconnect slower than the per-operation timeout, but the client's other connections reconnect within that timeout, so the timeout should still exceed a reconnect), and deferred invalidations are retried until they land — the first at once, and at the probe's cadence while the breaker is open — collapsing to one tenant-wide bump per tenant past 100,000 keys; until one lands, the process that owes it bypasses the lookups it would orphan. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. Nothing selects it yet: `cache.backend` accepts only `local` until the wiring (E4) adds `redis` and the `cache.redis.*` settings, so every deployment still runs `LocalCache`. Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, snapshot-rollback, compression, stored-size, server-stops-answering, refused-writes (a demoted primary, a full `noeviction` server), rotated-credentials, slow-reconnect, failover-behind-a-stable-address and owed-invalidation cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. - **Leases for work that must run in one process at a time, and the sweeper runs under one** (`internal/coord/` (new: `coord.go`, `local.go`, `elect.go`, `coordtest/`, + tests), `internal/app/{app,wire}.go` (+ tests), `internal/config/backends.go`, `internal/ingest/sweeper.go`, `.github/labeler.yml`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development,ingest-pipeline}.md`, `docs/src/content/docs/configuration.mdx`, `config.yaml`, `AGENTS.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `coord.Coordinator` hands out named leases (`TryAcquire` → a `Term` with a strictly increasing fencing `Token`, a `Done` channel and `Resign`; `ErrHeld` while another holder's term is live), and `coord.RunElected` runs a loop only while its process holds the lease, resigning when the loop returns and campaigning again every 2s. `coord.Local` is the in-process implementation, and `coordtest.Conformance` is the suite every implementation runs — the NATS KV backend that lets several replicas share one queue comes next. The sweeper now runs through `RunElected` under the `sweeper` lease; with the in-process coordinator the one process always holds it, so nothing changes for a single-process deployment beyond one `coord: elected` log line at startup. `coord.backend` now selects the coordinator (`local`, the only value), so a `config.Config` built without `config.Load` must name it as well as the other three layers' backends. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 51239a095..519316940 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -115,11 +115,11 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on — one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)) — each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope by their raw names, which the cache escapes where it builds a key. `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The snapshot is taken before any input a bump invalidates is chosen, the tenant's connection included: a reload that moves the tenant to another address or database runs `Pools.Reconcile` and then `InvalidateTenant` (a repoint that keeps both, such as a username or `tls` change, reads the same tables and bumps nothing), so a request that took the old pool files the old database's rows under a version that bump orphans, whether its `Set` lands before the bump or after. The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. - **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: a namespace key is `..
.
.`, its fields joined with `keyenc.Join` so a dot in a table or scope name is escaped rather than read as a separator, and an entry's key, `|.||…` with the caller's query key escaped whole (its `:` become `%3A`), folds in the tenant's version and each dependency's namespace key and namespace version, so bumping a table (a scopeless write) or one scope — scope is reserved and empty today, so every write is the whole-table bump — orphans every dependent entry without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results, and the flat directory's single tenant simply carries the `0` prefix. `BumpTenant` (behind `InvalidateTenant`) advances the tenant version that leads every namespace key of one tenant, orphaning its every namespace and every cached query in one step (the tenant version is folded into every query key, so a pipe result with no dependencies is orphaned too) — a table no bump ever keyed included, which an enumeration of the index would miss — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. -- **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET` and `MGET`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot; the table and scope are escaped with `keyenc` after the fixed prefix, so a `:` in a name is never read as the separator. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, the hash over their escaped forms, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token (a string of another length, a list, a hash) is replaced, so a token lost to eviction, expiry or a restart is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. While this process owes a bump on any of a lookup's tokens, that lookup is a bypass that files nothing: the bump would orphan whatever it found (other processes, which cannot know, serve those entries until it lands). `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing, each attempt bounded by `DialTimeout` (1 s), and a cluster client's topology read after the handshake by the larger of it and `Timeout`, which is also how long a connection waits on a silent server before it is redialed. Nothing selects this backend yet: `cache.backend` accepts only `local` until [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s E4 adds `redis`, the `cache.redis.*` settings and the wiring. +- **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET` and `MGET`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot; the table and scope are escaped with `keyenc` after the fixed prefix, so a `:` in a name is never read as the separator. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, the hash over their escaped forms, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token (a string of another length, a list, a hash) is replaced, so a token lost to eviction, expiry, `FLUSHALL` or a restart without persistence is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. Restoring an RDB or AOF snapshot, or a backup, is not a loss but a rollback: the old tokens come back with the values filed under them, so what was invalidated since is served again until its TTL, as after a failover to a replica that missed the bumps. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. While this process owes a bump on any of a lookup's tokens, that lookup is a bypass that files nothing: the bump would orphan whatever it found (other processes, which cannot know, serve those entries until it lands). `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing, each attempt bounded by `DialTimeout` (1 s), and a cluster client's topology read after the handshake by the larger of it and `Timeout`, which is also how long a connection waits on a silent server before it is redialed. Every connection is replaced after a minute (`clientOption`; rueidis retries what was in flight), because one that outlives a failover behind a stable address stays on the demoted node, which answers but refuses writes; the replacement re-resolves the address, so a bypassed process reaches the new primary, and delivers the bumps it owes, within about that long. rueidis dials a replacement lazily, under the context of the operation that lands on it, so a reconnect (dial and handshake, TLS included) slower than `Timeout` fails that operation. The breaker's probe runs under `DialTimeout` plus `Timeout`, so such a reconnect still closes an open breaker; but rueidis spreads commands over several connections (up to four to one server, by `GOMAXPROCS`, and one per cluster node) and the probe reconnects only the one it lands on, so size `Timeout` above a reconnect, or operations that land on the others keep failing. Nothing selects this backend yet: `cache.backend` accepts only `local` until [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s E4 adds `redis`, the `cache.redis.*` settings and the wiring. - **redis_codec.go** — the key schema (`tokenKeys`, `bumpKeys`, `valueKey`, which take a `Namespace`'s raw names and escape them) and the value frame: format, flags, expiry, the token list, the payload, zstd-compressed from `CompressMinBytes` when that is smaller. A stored value is capped at `MaxValueBytes` and a decoded one at eight times that, which refuses a zip bomb planted in a shared server; an unknown format, as a newer process writes during a rolling upgrade, is a miss. -- **breaker.go** — the circuit breaker: `BreakerThreshold` (5) consecutive failures open it, and while open every operation skips the server at once; after `BreakerOpenFor` (5 s) one background write (`SET :probe`) decides whether it closes, and only that probe's success closes it. Every connection is replaced after a minute (rueidis retries what was in flight), because one that outlives a failover behind a stable address stays on the demoted node, which answers but refuses writes; the replacement re-resolves the address, so a bypassed process reaches the new primary, and delivers the bumps it owes, within about that long. rueidis dials a replacement lazily, under the `Timeout` of the operation that needs it rather than `DialTimeout`, so this holds only where a fresh connection's dial and handshake (TLS included) fit in `Timeout`; size `Timeout` above that, or every reconnect, at the lifetime or after a drop, fails and the next one starts over. A transport failure or a timeout counts against the server; an error reply saying it takes no writes right now — `READONLY` (a demoted primary), `MASTERDOWN`, `OOM` (full under `maxmemory-policy noeviction`, Redis's default, so give the cache an evicting policy), `NOREPLICAS`, `MISCONF`, `LOADING`, `BUSY`, `CLUSTERDOWN` — opens it at once, whatever the threshold, since no bump can land. Any other reply, an error reply about one key (`WRONGTYPE`, `NOPERM`, `TRYAGAIN`) or one the backend cannot use included, counts as a success, and a caller that gave up first counts as nothing. -- **pending.go** — invalidations the server did not take, kept per token key (repeats coalesce) and retried by a background loop until they land: the first one owed at once, then with backoff from 100 ms to 10 s, and while the breaker is open at each probe, which the loop starts when due, so a process that makes no lookups (ingest only) recovers as soon as one that does. `Invalidate` sends its bumps 1,000 to a round trip, as the retry does. Landing late is still correct: a fresh token orphans the pre-write entries and any fill made meanwhile. Past `PendingMax` (100,000) keys the set collapses to one tenant bump per affected tenant. `Close` makes one last attempt at the bumps still owed, past the breaker (an open one is why they are owed); what that attempt cannot deliver is lost, and the entries they would orphan are served until their TTL — the same failure as a worker stopping between an insert and its invalidation. -- **metrics.go** — the shared backend's instruments (meter `wavehouse-cache`, `backend="redis"`, no tenant attribute): `wavehouse_cache_lookups_total` by `result` (`hit`, `miss`, `stale` — filed under since-bumped tokens, `bypass` — server skipped, or held by a bump this process owes, `error`), `wavehouse_cache_op_duration_seconds` by `op` (`lookup`, `set`, `invalidate`), `wavehouse_cache_breaker_open` (1 while bypassed, including before the first connection), `wavehouse_cache_invalidations_total` by `result` (`ok`, `deferred`), `wavehouse_cache_invalidations_pending` (bumps owed: entries they would orphan may be served stale meanwhile), `wavehouse_cache_value_bytes` (stored size), `wavehouse_cache_oversize_total` and `wavehouse_cache_set_failures_total` by `reason` (`oom`, `timeout`, `other`). +- **breaker.go** — the circuit breaker's state machine: `BreakerThreshold` (5) consecutive failures open it, `trip` opens it at once, and while it is open every operation skips the server; once `BreakerOpenFor` (5 s) has passed, `allow` hands one caller the probe, and only the probe's success closes it. What feeds it is `redis.go`'s. `record` counts a transport failure or a timeout against the server, and trips the breaker on an error reply that means no bump can land: one saying the server takes no writes right now, as `refusesWork` lists them — `READONLY` (a demoted primary), `MASTERDOWN`, `OOM` (full under `maxmemory-policy noeviction`, Redis's default, so give the cache an evicting policy), `NOREPLICAS`, `MISCONF`, `LOADING`, `BUSY`, `CLUSTERDOWN` — or one refusing the credentials, as `rejectsCredentials` lists them — `WRONGPASS`, `NOAUTH`, what a new connection's handshake meets after a password rotation — which it logs at `ERROR`. Any other reply, an error reply about one key or command (`WRONGTYPE`, `NOPERM`, `TRYAGAIN`) or one the backend cannot use included, counts as a success, and a caller that gave up first counts as nothing. The probe (`probe`) is a write, `SET :probe`, so a server that answers but refuses writes stays bypassed. +- **pending.go** — the invalidations owed (`pendingBumps`): kept per token key, repeats coalescing, and past `PendingMax` (100,000) keys collapsed to one tenant bump per affected tenant; `owesAny` tells `Lookup` which lookups to hold. `redis.go` delivers them: `Invalidate` sends its bumps 1,000 to a round trip and defers those the server does not take, and `drainLoop` retries them through `drain` until they land — the first one owed at once, then with backoff from 100 ms to 10 s, and while the breaker is open at each probe, which the loop starts when due, so a process that makes no lookups (ingest only) recovers as soon as one that does. Landing late is still correct: a fresh token orphans the pre-write entries and any fill made meanwhile. `Close` makes one last attempt at the bumps still owed, past the breaker (an open one is why they are owed); what that attempt cannot deliver is lost, and the entries they would orphan are served until their TTL — the same failure as a worker stopping between an insert and its invalidation. +- **metrics.go** — the shared backend's instruments (meter `wavehouse-cache`, `backend="redis"`, no tenant attribute): `wavehouse_cache_lookups_total` by `result` (`hit`, `miss`, `stale` — filed under since-bumped tokens, `bypass` — server skipped, or held by a bump this process owes, `error`), `wavehouse_cache_op_duration_seconds` by `op` (`lookup`, `set`, `invalidate`), `wavehouse_cache_breaker_open` (1 while bypassed, including before the first connection), `wavehouse_cache_invalidations_total` by `result` (`ok` counts every bump that lands, retried ones included, and `deferred` each time one is put off, a repeat of one already owed included, so the two overlap), `wavehouse_cache_invalidations_pending` (bumps owed: entries they would orphan may be served stale meanwhile), `wavehouse_cache_value_bytes` (stored size), `wavehouse_cache_oversize_total` and `wavehouse_cache_set_failures_total` by `reason` (`oom`, `timeout`, `other`). ### `config/` — Configuration diff --git a/internal/cache/breaker.go b/internal/cache/breaker.go index a33c94dae..6bcb89d4c 100644 --- a/internal/cache/breaker.go +++ b/internal/cache/breaker.go @@ -65,11 +65,15 @@ func (b *breaker) failure() { } // trip opens the breaker at once, for a reply that says the server cannot -// do the work: one is as conclusive as any number. -func (b *breaker) trip() { +// do the work: one is as conclusive as any number. It reports whether the +// breaker was closed or probing, so a caller logs once per opening and once +// per refused probe rather than once per operation in flight. +func (b *breaker) trip() bool { b.mu.Lock() defer b.mu.Unlock() + opened := !b.open || b.probing b.open, b.openedAt, b.probing = true, b.now(), false + return opened } func (b *breaker) isOpen() bool { diff --git a/internal/cache/breaker_test.go b/internal/cache/breaker_test.go index 3b2109d96..ace2f13f1 100644 --- a/internal/cache/breaker_test.go +++ b/internal/cache/breaker_test.go @@ -68,8 +68,9 @@ func TestBreaker_TripAndProbeSchedule(t *testing.T) { _, open := b.untilProbe() assert.False(t, open) - b.trip() + assert.True(t, b.trip(), "it opened") assert.True(t, b.isOpen(), "no threshold for a refusal") + assert.False(t, b.trip(), "already open") b.success() assert.True(t, b.isOpen(), "a success that is not the probe's leaves it open") d, open := b.untilProbe() @@ -85,7 +86,7 @@ func TestBreaker_TripAndProbeSchedule(t *testing.T) { require.True(t, probe) d, _ = b.untilProbe() assert.Equal(t, 5*time.Second, d, "while the probe runs, wait out a whole period") - b.trip() // the probe was refused too + assert.True(t, b.trip(), "the probe was refused too") d, _ = b.untilProbe() assert.Equal(t, 5*time.Second, d) diff --git a/internal/cache/metrics.go b/internal/cache/metrics.go index 34539f13b..c1dc12dc0 100644 --- a/internal/cache/metrics.go +++ b/internal/cache/metrics.go @@ -45,7 +45,7 @@ func newMetrics(backend string, breakerOpen func() bool, pending func() int) (*m metric.WithDescription("Shared-cache round-trip time by op: lookup, set, invalidate"), metric.WithUnit("s"), metric.WithExplicitBucketBoundaries(.0001, .00025, .0005, .001, .0025, .005, .01, .025, .05, .1, .25)) m.invalidation, errs[2] = meter.Int64Counter("wavehouse_cache_invalidations_total", - metric.WithDescription("Version-token bumps by result: ok, or deferred to the pending retry set")) + metric.WithDescription("Version-token bumps by result: ok counts every bump that lands, retried ones included; deferred counts each time one is put off, a repeat of one already owed included. They overlap: wavehouse_cache_invalidations_pending is what is still owed")) m.valueBytes, errs[3] = meter.Int64Histogram("wavehouse_cache_value_bytes", metric.WithDescription("Size of each value written to the shared cache, after compression"), metric.WithUnit("By"), metric.WithExplicitBucketBoundaries(256, 1<<10, 4<<10, 16<<10, 64<<10, 256<<10, 1<<20, 4<<20)) diff --git a/internal/cache/redis.go b/internal/cache/redis.go index 9d6733d4e..b722e446a 100644 --- a/internal/cache/redis.go +++ b/internal/cache/redis.go @@ -187,9 +187,11 @@ func (c RedisConfig) clientOption() rueidis.ClientOption { // Versions are random tokens, one per tenant, per table and per scope, // under the tenant's hash tag; a bump sets a fresh one. A value carries the // tokens it was computed under and is a hit only while they are all still -// current, so a lost token (eviction, expiry, a restart) can only cause -// misses. A lookup is one round trip. The server failing or timing out is -// a miss, a skipped fill and a deferred invalidation — never a failed query. +// current, so a lost token (eviction, expiry, a restart without persistence) +// can only cause misses. Restoring a snapshot is not a loss but a rollback: +// the old tokens return with their values. A lookup is one round trip. The +// server failing or timing out is a miss, a skipped fill and a deferred +// invalidation — never a failed query. type RedisCache struct { cfg RedisConfig opt rueidis.ClientOption @@ -310,9 +312,12 @@ func (r *RedisCache) conn() rueidis.Client { } // probe decides whether an open breaker closes. It writes: a server that -// answers but refuses writes (refusesWork) would take no bump either. +// answers but refuses writes (refusesWork) would take no bump either. rueidis +// redials under the calling operation's context, so the probe's budget is +// DialTimeout for a reconnect plus Timeout for the write: a reconnect slower +// than Timeout fails the operations waiting on it, but not the probe. func (r *RedisCache) probe(c rueidis.Client) { - ctx, cancel := context.WithTimeout(r.ctx, r.cfg.Timeout) + ctx, cancel := context.WithTimeout(r.ctx, r.cfg.DialTimeout+r.cfg.Timeout) defer cancel() r.record(r.ctx, c.Do(ctx, c.B().Set().Key(r.cfg.KeyPrefix+":probe").Value("1").Ex(time.Minute).Build()).Error()) if r.breaker.isOpen() { @@ -337,9 +342,15 @@ func (r *RedisCache) record(parent context.Context, err error) { return } if re, ok := rueidis.IsRedisErr(err); ok { - if refusesWork(re.Error()) { + switch msg := re.Error(); { + case rejectsCredentials(msg): + if r.breaker.trip() { + slog.ErrorContext(parent, "cache: redis rejected the credentials; bypassing the cache until they work", + "addrs", r.cfg.Addrs, "error", msg) + } + case refusesWork(msg): r.breaker.trip() - } else { + default: r.breaker.success() } return @@ -371,6 +382,15 @@ func refusesWork(msg string) bool { return false } +// rejectsCredentials reports whether an error reply refuses this process's +// credentials, as a connection's handshake after a password rotation does: +// every operation meets it, so it refuses all work. NOPERM is not one: it +// names a key or a command, which the rest of the work may not touch. +func rejectsCredentials(msg string) bool { + code, _, _ := strings.Cut(msg, " ") + return code == "WRONGPASS" || code == "NOAUTH" +} + // Lookup reads the tokens deps fold and the entry for sha in one pipelined // round trip. A token that does not exist yet is created, never read as a // value, so its first use is a miss. diff --git a/internal/cache/redis_integration_test.go b/internal/cache/redis_integration_test.go index 3169e9230..d40a2fae7 100644 --- a/internal/cache/redis_integration_test.go +++ b/internal/cache/redis_integration_test.go @@ -5,6 +5,7 @@ package cache_test import ( "bytes" "context" + "crypto/rand" "crypto/tls" "fmt" "io" @@ -12,6 +13,7 @@ import ( "slices" "strconv" "strings" + "sync" "sync/atomic" "testing" "time" @@ -92,11 +94,12 @@ func startDragonfly(t *testing.T) *server { } // startCluster runs a one-node Redis Cluster owning every slot: enough for -// the server to enforce cluster semantics — CROSSSLOT on a multi-key -// command, MOVED routing through the client — which is what the key schema -// must survive. The node announces 127.0.0.1 on a host port bound to the -// same number, so the address the client learns from CLUSTER SLOTS is -// dialable from the test. +// the server to enforce CROSSSLOT on a multi-key command, which is what the +// key schema must survive, and for the cluster client to read the topology. +// It never routes across nodes: a node owning every slot never answers +// MOVED. The node announces 127.0.0.1 on a host port bound to the same +// number, so the address the client learns from CLUSTER SLOTS is dialable +// from the test. func startCluster(t *testing.T) *server { t.Helper() port := freePort(t) @@ -233,6 +236,10 @@ func TestRedis_Conformance(t *testing.T) { t.Parallel() testCompression(t, s) }) + t.Run("an incompressible value over the stored limit is not stored", func(t *testing.T) { + t.Parallel() + testIncompressible(t, s) + }) }) } } @@ -316,6 +323,24 @@ func testCompression(t *testing.T, s *server) { assert.Less(t, stored, int64(len(rows)/10)) } +// The stored-size limit, past the raw one: random bytes do not compress, so +// a value between the two is refused only once it is encoded. +func testIncompressible(t *testing.T, s *server) { + ctx := context.Background() + c := open(t, s, uniquePrefix()) + rows := make([]byte, maxValue+1<<10) + _, _ = rand.Read(rows) + require.Less(t, len(rows), maxValue*cache.DecodedFactor, "within the raw limit") + deps := []cache.Namespace{{Tenant: "acme", Table: "events"}} + _, snap, err := c.Lookup(ctx, "acme", "big", deps) + require.NoError(t, err) + require.NoError(t, c.Set(ctx, snap, rows, time.Minute), "declined, not failed") + e, _, err := c.Lookup(ctx, "acme", "big", deps) + require.NoError(t, err) + assert.Nil(t, e.Value) + assert.Zero(t, valueCount(raw(t, s))(c), "nothing stored") +} + // A token that is lost — evicted, expired, flushed, a restart without // persistence — is recreated fresh, so a value stored under its predecessor // can only miss. A counter recreated at its initial value would serve the @@ -424,8 +449,9 @@ func dockerClient(t *testing.T) *testcontainers.DockerClient { return d } -// A server that stops answering costs a request at most about the op -// timeout, then nothing: the breaker opens and the cache is bypassed. +// A server that stops answering costs a lookup a bounded wait — asserted +// under ten times the op timeout, as the suite runs in parallel under -race — +// then nothing: the breaker opens and the cache is bypassed. // Invalidations made meanwhile are kept and land once it answers again, and // a process that boots while it is down starts bypassed and connects later. func TestRedis_ServerStopsAnswering(t *testing.T) { @@ -460,7 +486,7 @@ func TestRedis_ServerStopsAnswering(t *testing.T) { e, snap, err := a.Lookup(ctx, "acme", "q", deps) require.Error(t, err, "lookup %d", i) assert.Nil(t, e.Value) - assert.Less(t, time.Since(start), 10*timeout, "a lookup costs at most about the timeout") + assert.Less(t, time.Since(start), 10*timeout, "a lookup returns within ten times the timeout") require.NoError(t, a.Set(ctx, snap, []byte("rows"), time.Minute), "the failed lookup's snapshot files nothing") } require.True(t, cache.Bypassed(a), "three timeouts open the breaker") @@ -560,35 +586,63 @@ func (c unanswered) Write(b []byte) (int, error) { return c.Conn.Write(b) } -// forward proxies each connection it accepts to the address target holds -// at that moment, so a switch moves new connections only, as a stable DNS -// name or a proxy does after a failover. It returns the address to dial. -func forward(t *testing.T, target *atomic.Pointer[string]) string { +// proxy forwards each connection it accepts to the address target holds at +// that moment, so a switch moves new connections only, as a stable DNS name +// or a proxy does after a failover. It holds a new connection for delay +// before forwarding it, as a slow dial and handshake would. +type proxy struct { + addr string // to dial + target atomic.Pointer[string] + delay atomic.Int64 // a time.Duration + + mu sync.Mutex + conns []net.Conn +} + +func newProxy(t *testing.T, target string) *proxy { t.Helper() var lc net.ListenConfig ln, err := lc.Listen(context.Background(), "tcp", "127.0.0.1:0") require.NoError(t, err) t.Cleanup(func() { _ = ln.Close() }) + p := &proxy{addr: ln.Addr().String()} + p.target.Store(&target) go func() { for { c, err := ln.Accept() if err != nil { return } - go func() { - defer func() { _ = c.Close() }() - var d net.Dialer - u, err := d.DialContext(context.Background(), "tcp", *target.Load()) - if err != nil { - return - } - defer func() { _ = u.Close() }() - go func() { _, _ = io.Copy(u, c); _ = u.Close() }() - _, _ = io.Copy(c, u) - }() + p.mu.Lock() + p.conns = append(p.conns, c) + p.mu.Unlock() + go p.serve(c) } }() - return ln.Addr().String() + return p +} + +func (p *proxy) serve(c net.Conn) { + defer func() { _ = c.Close() }() + time.Sleep(time.Duration(p.delay.Load())) + var d net.Dialer + u, err := d.DialContext(context.Background(), "tcp", *p.target.Load()) + if err != nil { + return + } + defer func() { _ = u.Close() }() + go func() { _, _ = io.Copy(u, c); _ = u.Close() }() + _, _ = io.Copy(c, u) +} + +// drop closes every connection p carries, so the client must reconnect. +func (p *proxy) drop() { + p.mu.Lock() + defer p.mu.Unlock() + for _, c := range p.conns { + _ = c.Close() + } + p.conns = nil } // A failover behind a stable address: the connections the process holds @@ -616,9 +670,8 @@ func TestRedis_FailoverBehindAStableAddress(t *testing.T) { return err == nil && strings.Contains(info, "master_link_status:up") }, 30*time.Second, 50*time.Millisecond, "the replica syncs") - var target atomic.Pointer[string] - target.Store(&primary.addr) - stable := &server{addr: forward(t, &target), mode: cache.RedisStandalone} + fwd := newProxy(t, primary.addr) + stable := &server{addr: fwd.addr, mode: cache.RedisStandalone} prefix := uniquePrefix() a := open(t, stable, prefix, func(c *cache.RedisConfig) { c.BreakerThreshold, c.BreakerOpenFor = 1000, 200*time.Millisecond @@ -637,7 +690,7 @@ func TestRedis_FailoverBehindAStableAddress(t *testing.T) { _, err = a.Invalidate(ctx, deps) require.ErrorContains(t, err, "READONLY") require.True(t, cache.Bypassed(a)) - target.Store(&replica.addr) + fwd.target.Store(&replica.addr) deadline := time.Now().Add(15 * time.Second) for cache.Pending(a) > 0 { @@ -814,3 +867,93 @@ func TestRedis_RefusedWrites(t *testing.T) { }) } } + +// Restoring a snapshot — RDB, AOF or a backup — is a rollback, not a lost +// token: the old tokens come back with the values filed under them, so what +// was invalidated since is served again, as after a failover to a replica +// that missed the bumps. The documented exception to "lost tokens miss". +func TestRedis_RestoredSnapshotIsARollback(t *testing.T) { + t.Parallel() + ctx := context.Background() + s := startStandalone(t, redisImage, "redis-server", "--save", "", "--appendonly", "no", "--enable-debug-command", "yes") + r := raw(t, s) + prefix := uniquePrefix() + c := open(t, s, prefix) + deps := []cache.Namespace{{Tenant: "acme", Table: "events"}} + _, snap, err := c.Lookup(ctx, "acme", "q", deps) + require.NoError(t, err) + require.NoError(t, c.Set(ctx, snap, []byte("pre-write rows"), time.Minute)) + command(t, r, "SAVE") + + _, err = c.Invalidate(ctx, deps) + require.NoError(t, err) + e, _, err := c.Lookup(ctx, "acme", "q", deps) + require.NoError(t, err) + require.Nil(t, e.Value) + + command(t, r, "DEBUG", "RELOAD", "NOSAVE") + e, _, err = open(t, s, prefix).Lookup(ctx, "acme", "q", deps) + require.NoError(t, err) + assert.Equal(t, "pre-write rows", string(e.Value), "the restore brought back the pre-write token") +} + +// A reconnect slower than the op timeout (a dial and TLS handshake across +// zones) fails the operations waiting on it, since rueidis dials under their +// context; they open the breaker, and the probe, whose budget covers a dial, +// completes the reconnect and closes it. Only the connection the probe +// lands on: rueidis spreads commands over several, and the rest still +// reconnect under the op timeout. +func TestRedis_ProbeFitsASlowReconnect(t *testing.T) { + t.Parallel() + ctx := context.Background() + s := startRedis(t) + p := newProxy(t, s.addr) + const timeout = 100 * time.Millisecond + a := open(t, &server{addr: p.addr, mode: cache.RedisStandalone}, uniquePrefix(), func(c *cache.RedisConfig) { + c.Timeout, c.DialTimeout = timeout, 2*time.Second + c.BreakerThreshold, c.BreakerOpenFor = 1, 200*time.Millisecond + }) + deps := []cache.Namespace{{Tenant: "acme", Table: "events"}} + _, _, err := a.Lookup(ctx, "acme", "q", deps) + require.NoError(t, err) + + p.delay.Store(int64(5 * timeout)) + p.drop() + _, _, err = a.Lookup(ctx, "acme", "q", deps) + require.Error(t, err, "the reconnect does not fit in the lookup's timeout") + require.True(t, cache.Bypassed(a)) + require.Eventually(t, func() bool { return !cache.Bypassed(a) }, 10*time.Second, 10*time.Millisecond, + "the probe reconnects within the dial timeout") +} + +// A password rotated under a running process refuses its next connection's +// handshake, and so every operation: one such reply opens the breaker, +// whatever the threshold, and the probe closes it once the credentials work. +func TestRedis_RejectedCredentialsOpenTheBreaker(t *testing.T) { + t.Parallel() + ctx := context.Background() + s := startRedis(t) + r := raw(t, s) + command(t, r, "ACL", "SETUSER", "rotating", "on", ">old", "+@all", "~*") + a := open(t, s, uniquePrefix(), func(c *cache.RedisConfig) { + c.Username, c.Password = "rotating", "old" + c.BreakerThreshold, c.BreakerOpenFor = 1000, 200*time.Millisecond + }) + deps := []cache.Namespace{{Tenant: "acme", Table: "events"}} + _, _, err := a.Lookup(ctx, "acme", "q", deps) + require.NoError(t, err) + + command(t, r, "ACL", "SETUSER", "rotating", "resetpass", ">new") + command(t, r, "CLIENT", "KILL", "USER", "rotating") // as the connection lifetime would + require.Eventually(t, func() bool { + _, _, err = a.Lookup(ctx, "acme", "q", deps) + return err != nil && strings.Contains(err.Error(), "WRONGPASS") + }, 5*time.Second, 10*time.Millisecond, "the reconnect is refused") + assert.True(t, cache.Bypassed(a), "one refused handshake opens the breaker") + time.Sleep(time.Second) // several refused probes + assert.True(t, cache.Bypassed(a)) + + command(t, r, "ACL", "SETUSER", "rotating", ">old") + require.Eventually(t, func() bool { return !cache.Bypassed(a) }, 10*time.Second, 10*time.Millisecond, + "the probe closes it once the credentials work") +} diff --git a/internal/cache/redis_test.go b/internal/cache/redis_test.go index a9ccc3711..deb4f2c56 100644 --- a/internal/cache/redis_test.go +++ b/internal/cache/redis_test.go @@ -119,8 +119,11 @@ func TestRedis_UnreachableIsBypassed(t *testing.T) { require.NoError(t, r.Close(), "idempotent") } +// Each size limit is counted: the raw one before compression, the stored one +// after it, which a server-less Set cannot otherwise tell from a bypass. func TestRedis_SetDeclinesWithoutTouchingTheServer(t *testing.T) { - t.Parallel() + // No t.Parallel(): swaps the global meter provider. + reader := meterReader(t) r, err := NewRedis(RedisConfig{Addrs: []string{closedAddr(t)}, DialTimeout: 100 * time.Millisecond, MaxValueBytes: 64}) require.NoError(t, err) t.Cleanup(func() { _ = r.Close() }) @@ -139,6 +142,7 @@ func TestRedis_SetDeclinesWithoutTouchingTheServer(t *testing.T) { } { require.NoError(t, r.Set(ctx, tt.snap, tt.value, tt.ttl), tt.name) } + assert.Equal(t, int64(2), sumOf(t, collect(t, reader), "wavehouse_cache_oversize_total", "", "")) } func TestRedis_Record(t *testing.T) { @@ -156,6 +160,15 @@ func TestRedis_Record(t *testing.T) { assert.True(t, r.breaker.isOpen()) } +func TestRejectsCredentials(t *testing.T) { + t.Parallel() + assert.True(t, rejectsCredentials("WRONGPASS invalid username-password pair or user is disabled.")) + assert.True(t, rejectsCredentials("NOAUTH Authentication required.")) + assert.False(t, rejectsCredentials("NOPERM No permissions to access a key"), "about a key, not the credentials") + assert.False(t, rejectsCredentials("READONLY You can't write against a read only replica.")) + assert.False(t, rejectsCredentials("ERR WRONGPASS")) +} + func TestRefusesWork(t *testing.T) { t.Parallel() for _, msg := range []string{ @@ -224,8 +237,10 @@ func TestReadTokens_NotAnArrayIsMalformed(t *testing.T) { require.ErrorIs(t, err, errMalformedReply) } -func TestRedisMetrics(t *testing.T) { - // No t.Parallel(): swaps the global meter provider. +// meterReader makes a manual reader the global meter provider's for the +// test's duration, which must then not run in parallel. +func meterReader(t *testing.T) *sdkmetric.ManualReader { + t.Helper() saved := otel.GetMeterProvider() reader := sdkmetric.NewManualReader() mp := sdkmetric.NewMeterProvider(sdkmetric.WithReader(reader)) @@ -234,51 +249,66 @@ func TestRedisMetrics(t *testing.T) { _ = mp.Shutdown(context.Background()) otel.SetMeterProvider(saved) }) + return reader +} - r, err := NewRedis(RedisConfig{Addrs: []string{closedAddr(t)}, DialTimeout: 100 * time.Millisecond}) - require.NoError(t, err) - t.Cleanup(func() { _ = r.Close() }) - ctx := context.Background() - _, _, _ = r.Lookup(ctx, "acme", "q", nil) - _, _ = r.Invalidate(ctx, []Namespace{{Tenant: "acme", Table: "events"}}) - r.metrics.op("set", time.Now()) - r.metrics.stored(10) - r.metrics.tooLarge() - r.metrics.setFailed("oom") - +// collect reads every instrument's data from reader, by name. +func collect(t *testing.T, reader *sdkmetric.ManualReader) map[string]metricdata.Aggregation { + t.Helper() var rm metricdata.ResourceMetrics - require.NoError(t, reader.Collect(ctx, &rm)) + require.NoError(t, reader.Collect(context.Background(), &rm)) got := map[string]metricdata.Aggregation{} for _, sm := range rm.ScopeMetrics { for _, m := range sm.Metrics { got[m.Name] = m.Data } } - sumOf := func(name, key, value string) int64 { - t.Helper() - var n int64 - switch d := got[name].(type) { - case metricdata.Sum[int64]: - for _, dp := range d.DataPoints { - if v, ok := dp.Attributes.Value(attribute.Key(key)); key == "" || ok && v.AsString() == value { - n += dp.Value - } - } - case metricdata.Gauge[int64]: - for _, dp := range d.DataPoints { + return got +} + +// sumOf adds up an int64 counter's or gauge's points whose attribute key +// is value; key "" takes every point. +func sumOf(t *testing.T, got map[string]metricdata.Aggregation, name, key, value string) int64 { + t.Helper() + var n int64 + switch d := got[name].(type) { + case metricdata.Sum[int64]: + for _, dp := range d.DataPoints { + if v, ok := dp.Attributes.Value(attribute.Key(key)); key == "" || ok && v.AsString() == value { n += dp.Value } - default: - t.Fatalf("%s: %T", name, got[name]) } - return n + case metricdata.Gauge[int64]: + for _, dp := range d.DataPoints { + n += dp.Value + } + default: + t.Fatalf("%s: %T", name, got[name]) } - assert.Equal(t, int64(1), sumOf("wavehouse_cache_lookups_total", "result", resultBypass)) - assert.Equal(t, int64(1), sumOf("wavehouse_cache_invalidations_total", "result", "deferred")) - assert.Equal(t, int64(1), sumOf("wavehouse_cache_invalidations_pending", "", "")) - assert.Equal(t, int64(1), sumOf("wavehouse_cache_breaker_open", "", "")) - assert.Equal(t, int64(1), sumOf("wavehouse_cache_oversize_total", "", "")) - assert.Equal(t, int64(1), sumOf("wavehouse_cache_set_failures_total", "reason", "oom")) + return n +} + +func TestRedisMetrics(t *testing.T) { + // No t.Parallel(): swaps the global meter provider. + reader := meterReader(t) + r, err := NewRedis(RedisConfig{Addrs: []string{closedAddr(t)}, DialTimeout: 100 * time.Millisecond}) + require.NoError(t, err) + t.Cleanup(func() { _ = r.Close() }) + ctx := context.Background() + _, _, _ = r.Lookup(ctx, "acme", "q", nil) + _, _ = r.Invalidate(ctx, []Namespace{{Tenant: "acme", Table: "events"}}) + r.metrics.op("set", time.Now()) + r.metrics.stored(10) + r.metrics.tooLarge() + r.metrics.setFailed("oom") + + got := collect(t, reader) + assert.Equal(t, int64(1), sumOf(t, got, "wavehouse_cache_lookups_total", "result", resultBypass)) + assert.Equal(t, int64(1), sumOf(t, got, "wavehouse_cache_invalidations_total", "result", "deferred")) + assert.Equal(t, int64(1), sumOf(t, got, "wavehouse_cache_invalidations_pending", "", "")) + assert.Equal(t, int64(1), sumOf(t, got, "wavehouse_cache_breaker_open", "", "")) + assert.Equal(t, int64(1), sumOf(t, got, "wavehouse_cache_oversize_total", "", "")) + assert.Equal(t, int64(1), sumOf(t, got, "wavehouse_cache_set_failures_total", "reason", "oom")) assert.Contains(t, got, "wavehouse_cache_op_duration_seconds") assert.Contains(t, got, "wavehouse_cache_value_bytes") } From a10ba14cb98b03ddb08becbc2d7e96f710da9031 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 09:01:32 -0400 Subject: [PATCH 64/79] docs(agents): list mutationtest in the file structure Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- AGENTS.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/AGENTS.md b/AGENTS.md index 9666db1bf..cff241d62 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -448,7 +448,7 @@ internal/query/ → Structured query AST + SQL builder internal/settings/ → Settings directory (validate, adopted snapshot + reload, watcher, embedded seed) internal/stream/ → SSE fan-out (event Hub: project once per role, Subscriber outbound queue, Bucket fan-out, keepalive Heartbeater wheel) internal/tenant/ → Tenant id (type, grammar, reserved default, request header name) -internal/testutil/ → Shared test helpers (mocks, JWT + schema helpers; logtest/ captures or silences the default logger; cachetest/ is the conformance suite every cache.Cache backend runs) +internal/testutil/ → Shared test helpers (mocks, JWT + schema helpers; logtest/ captures or silences the default logger; cachetest/ is the conformance suite every cache.Cache backend runs; mutationtest/ holds the shared write-classifier cases) tests/ → Integration & E2E tests tests/integration/ → Go integration tests (//go:build integration; ClickHouse testcontainer) tests/e2e/ → E2E test stack (scripts/orchestrator boots a ClickHouse testcontainer + the wavehouse-cov binary) From c42b6db2603dc8d517441bba981dc6685833c135 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 09:04:44 -0400 Subject: [PATCH 65/79] docs(cache): fix the one query-key/entry-key spot the terminology pass missed MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit version_manager.go's tenants field comment still said "the first query key built for it" — the same overload the type doc three lines above and the other doc fixes in this round eliminated (query key = the caller's input, entry key = what QueryKey renders). Both confirmation reviewers caught this independently. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- internal/cache/version_manager.go | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/internal/cache/version_manager.go b/internal/cache/version_manager.go index d6fd7496a..6b1de3b5f 100644 --- a/internal/cache/version_manager.go +++ b/internal/cache/version_manager.go @@ -23,8 +23,8 @@ import ( type VersionManager struct { mu sync.RWMutex - // tenants holds each tenant's index from the first query key built for - // it until the tenant is bumped or pruned. + // tenants holds each tenant's index from the first entry key (`QueryKey`) + // built for it until the tenant is bumped or pruned. tenants map[tenant.ID]*tenantVersions // lastGen is the last generation handed to a tenant; see tenantVersions.gen. From ce0da484740928e07cd721831ed5afd54159c401 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 09:32:52 -0400 Subject: [PATCH 66/79] docs(cache): match configuration.mdx and deployment.md to #626's fixes Verified against the merged internal/cache/{redis,breaker}.go: - rejectsCredentials (WRONGPASS, NOAUTH) now trips the breaker at once at runtime too, logged at ERROR once per opening and once per refused probe (breaker.trip()); NOPERM deliberately does not (it names one key/command, not every operation, so an ACL denial on one tenant's traffic shouldn't bypass the cache for all of them). configuration.mdx's circuit-breaker paragraph didn't mention this at all; added it, distinct from the dial-time isAuthError logging (a separate function, used only in dial()/dialLoop(), where WRONGPASS, NOAUTH and NOPERM are all ERROR since any of them blocks the connection outright before a client exists to send a scoped command on). - deployment.md's failover bullet gets the probe's DialTimeout+Timeout budget and the caveat that every other connection still redials under Timeout alone, matching architecture.md's redis.go bullet. - deployment.md's metrics list presented invalidations_total's `ok`/ `deferred` as if they partition the count; they overlap (a retried landing counts `ok` again), per metrics.go's own doc comment. Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- docs/src/content/docs/configuration.mdx | 2 +- docs/src/content/docs/deployment.md | 4 ++-- 2 files changed, 3 insertions(+), 3 deletions(-) diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index de970487c..c8878afee 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -168,7 +168,7 @@ The `redis` backend's settings, read only when `cache.backend` is `redis`. It is | `cache.redis.compress_min_bytes` | `WH_CACHE_REDIS_COMPRESS_MIN_BYTES` | `1024` | Results at least this large are zstd-compressed when that makes them smaller. `0` never compresses. | | `cache.redis.version_ttl` | `WH_CACHE_REDIS_VERSION_TTL` | `168h` | How long a table's or tenant's version token outlives its last write, so the tokens of dropped tables and removed tenants eventually expire. At least `2s`. An expired token only causes misses. | -**When the server is unreachable or misbehaves, the cache is bypassed; queries are not.** A failure or a timeout makes the lookup a miss and the fill a no-op. Five in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), open a circuit breaker that skips the server entirely until a probe write, every 5 s, succeeds. Queries then go straight to ClickHouse, still coalesced per instance by `singleflight`. An invalidation the server did not take is kept and retried until it lands, and until then the instance that owes it bypasses the lookups it would orphan. `/readyz` does not depend on the cache. +**When the server is unreachable or misbehaves, the cache is bypassed; queries are not.** A failure or a timeout makes the lookup a miss and the fill a no-op. Five in a row, one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), or one refusing the credentials (`WRONGPASS`, `NOAUTH` — what an already-open connection meets once the password is rotated; logged at `ERROR`, once per opening and once per refused probe) open a circuit breaker that skips the server entirely until a probe write, every 5 s, succeeds. `NOPERM` does not open it: it names one key or command an ACL user cannot use, not every operation, so it should not bypass the cache for every tenant. Queries then go straight to ClickHouse, still coalesced per instance by `singleflight`. An invalidation the server did not take is kept and retried until it lands, and until then the instance that owes it bypasses the lookups it would orphan. `/readyz` does not depend on the cache. **Boot does not wait for the server.** A malformed block (an address without a port, `mode: cluster` with `db` other than `0`, an unreadable or unparsable TLS file) refuses boot. A server that cannot be reached, or that refuses the credentials, does not: the process boots with the cache bypassed and keeps reconnecting, with backoff up to 30 s. A rejected credential (`WRONGPASS`, `NOAUTH`, or `NOPERM` for an ACL user missing a connection command) is logged at `ERROR` on every attempt; any other failure at `WARN`. This is deliberate: a rotated Redis password must not crash-loop every instance at once. Watch `wavehouse_cache_breaker_open`, which reads `1` while the cache is bypassed, including before the first connection. diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index fcbcc2523..b3fdcf801 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -432,7 +432,7 @@ The query-result cache is the layer that can be shared today. With the default ` **What another instance can see.** Ingest is already asynchronous: `/v1/ingest` answers before the batch is inserted. Once the inserting instance's worker has written the batch to ClickHouse, it replaces the table's version token in Redis, and from then on a lookup on any instance misses and reads the new rows. The cache adds no delay of its own beyond that single write. The exceptions: - **The server is unreachable from the inserting instance.** The invalidation is kept and retried until it lands (`wavehouse_cache_invalidations_pending` counts what is owed). Meanwhile other instances that can still reach the server keep serving the older results, for as long as the outage lasts and at most until each entry's TTL. An instance that stops while invalidations are still owed loses them, with the same bound. The same thing happens today when a process stops between an insert and its invalidation. -- **A failover to a replica that had not yet received the latest token writes** can bring back entries filed under the older tokens, bounded by the replication lag at the moment of failover and those entries' TTL. WaveHouse never reads from replicas. Behind a stable address (a managed primary endpoint), an instance still connected to the demoted node has its writes refused, which bypasses its cache; connections are replaced every minute, so it reaches the new primary and delivers the invalidations it owes within about that long. +- **A failover to a replica that had not yet received the latest token writes** can bring back entries filed under the older tokens, bounded by the replication lag at the moment of failover and those entries' TTL. WaveHouse never reads from replicas. Behind a stable address (a managed primary endpoint), an instance still connected to the demoted node has its writes refused, which bypasses its cache; connections are replaced every minute, so it reaches the new primary and delivers the invalidations it owes within about that long. The breaker's own probe gets a longer reconnect budget (`dial_timeout` plus `timeout`) so a slow reconnect still closes it, but every other connection redials under `timeout` alone: size it above how long a reconnect actually takes, or operations that land on one of those keep failing after the probe has already succeeded. - **The server is full and `maxmemory-policy` is `noeviction`.** It refuses the token writes. The inserting instance keeps its invalidations and retries them, bypassing its cache meanwhile, but every other instance serves the results from before the insert until one lands, up to their TTL. - **A pipe that writes** (an `INSERT` in `pipes.json`) has its result cached like a read, so a repeated identical call is answered from the cache and the write does not run again ([#386](https://github.com/Wave-RF/WaveHouse/issues/386)). With a shared cache that holds on every instance, until the entry's TTL. - **Admin writes through `POST /v1/ops/query`** do not invalidate the cache ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). With a shared cache, the stale results they leave are served by every instance, not only one. @@ -443,7 +443,7 @@ The query-result cache is the layer that can be shared today. With the default ` **Coalescing stays per instance.** `singleflight` collapses identical concurrent queries within each instance, so a cold hot query costs at most one ClickHouse query per instance, not one per request. -**Metrics** (meter `wavehouse-cache`, every series labeled `backend="redis"`, no tenant label): `wavehouse_cache_lookups_total{result}` (`hit`, `miss`, `stale`, `bypass`, `error`), `wavehouse_cache_op_duration_seconds{op}` (`lookup`, `set`, `invalidate`), `wavehouse_cache_breaker_open` (1 while the cache is bypassed), `wavehouse_cache_invalidations_total{result}` (`ok`, `deferred`), `wavehouse_cache_invalidations_pending`, `wavehouse_cache_value_bytes`, `wavehouse_cache_oversize_total` and `wavehouse_cache_set_failures_total{reason}` (`oom`, `timeout`, `other`). Two signals are worth alerting on: `wavehouse_cache_breaker_open` at 1, or `wavehouse_cache_invalidations_pending` above 0, for more than a few minutes. +**Metrics** (meter `wavehouse-cache`, every series labeled `backend="redis"`, no tenant label): `wavehouse_cache_lookups_total{result}` (`hit`, `miss`, `stale`, `bypass`, `error`), `wavehouse_cache_op_duration_seconds{op}` (`lookup`, `set`, `invalidate`), `wavehouse_cache_breaker_open` (1 while the cache is bypassed), `wavehouse_cache_invalidations_total{result}` (`ok` counts every bump that lands, retried ones included, and `deferred` each time one is put off, a repeat of one already owed included — the two overlap, not a split), `wavehouse_cache_invalidations_pending`, `wavehouse_cache_value_bytes`, `wavehouse_cache_oversize_total` and `wavehouse_cache_set_failures_total{reason}` (`oom`, `timeout`, `other`). Two signals are worth alerting on: `wavehouse_cache_breaker_open` at 1, or `wavehouse_cache_invalidations_pending` above 0, for more than a few minutes. For local development, `docker compose -f deployments/compose/dependencies.yaml --profile redis up -d` starts a Redis on `localhost:6379` with no persistence. From 0eea1cbc69558819aeec2ff9c9eed95dd222eb16 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 12:52:53 -0400 Subject: [PATCH 67/79] fix(cache): close the breaker only on a probe write within Timeout The probe's reconnect allowance also covered its write, so a server answering slower than Timeout, but within the allowance, passed every probe: the breaker closed, the next operations timed out on the server, and it reopened, every cycle. A probe write slower than Timeout is now repeated under Timeout alone, and the repeat decides. rueidis bounds a reconnect's dial, TLS included, by DialTimeout and then its handshake by DialTimeout again, so the allowance is now twice DialTimeout. The slow-reconnect test now reconnects over TLS, with the dial and the handshake each taking most of DialTimeout, and a new test keeps a server that answers slower than Timeout bypassed. Docs: a restart after a crash that reloads the last save is a rollback, like restoring a snapshot; what the deferred invalidation count counts; the client's other connections must reconnect within Timeout. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 4 +- internal/cache/metrics.go | 2 +- internal/cache/redis.go | 38 ++++--- internal/cache/redis_integration_test.go | 122 +++++++++++++++++++---- 5 files changed, 134 insertions(+), 34 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 366b49152..a75b24028 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A Redis-compatible shared cache backend, built but not yet selectable** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/cache.go`, `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`, `CONTRIBUTING.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart without persistence can only cause misses — `maxmemory-policy allkeys-lru` is safe. Restoring an RDB or AOF snapshot, or a backup, is a rollback instead: the old tokens return with their values, so invalidations made since are undone until those entries' TTL. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like) or the credentials (`WRONGPASS` or `NOAUTH` after a password rotation, logged at `ERROR`), open a circuit breaker that skips the server until a probe write succeeds (connections are replaced every minute, so after a failover behind a stable address the process reaches the new primary, and delivers the bumps it owes, within about a minute; the probe's budget covers a reconnect slower than the per-operation timeout, but the client's other connections reconnect within that timeout, so the timeout should still exceed a reconnect), and deferred invalidations are retried until they land — the first at once, and at the probe's cadence while the breaker is open — collapsing to one tenant-wide bump per tenant past 100,000 keys; until one lands, the process that owes it bypasses the lookups it would orphan. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. Nothing selects it yet: `cache.backend` accepts only `local` until the wiring (E4) adds `redis` and the `cache.redis.*` settings, so every deployment still runs `LocalCache`. Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, snapshot-rollback, compression, stored-size, server-stops-answering, refused-writes (a demoted primary, a full `noeviction` server), rotated-credentials, slow-reconnect, failover-behind-a-stable-address and owed-invalidation cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. +- **A Redis-compatible shared cache backend, built but not yet selectable** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/cache.go`, `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`, `CONTRIBUTING.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart without persistence can only cause misses — `maxmemory-policy allkeys-lru` is safe. Restoring an RDB or AOF snapshot, or a backup, is a rollback instead (a restart after a crash that reloads the server's last save included, which stock Redis and Valkey make by default): the old tokens return with their values, so invalidations made since are undone until those entries' TTL. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like) or the credentials (`WRONGPASS` or `NOAUTH` after a password rotation, logged at `ERROR`), open a circuit breaker that skips the server until a probe write succeeds within the per-operation timeout (connections are replaced every minute, so after a failover behind a stable address the process reaches the new primary, and delivers the bumps it owes, within about a minute; the probe also allows up to twice the dial timeout for a reconnect, but the client's other connections must reconnect within the per-operation timeout, so that timeout should still exceed a reconnect), and deferred invalidations are retried until they land — the first at once, and at the probe's cadence while the breaker is open — collapsing to one tenant-wide bump per tenant past 100,000 keys; until one lands, the process that owes it bypasses the lookups it would orphan. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. Nothing selects it yet: `cache.backend` accepts only `local` until the wiring (E4) adds `redis` and the `cache.redis.*` settings, so every deployment still runs `LocalCache`. Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, snapshot-rollback, compression, stored-size, server-stops-answering, refused-writes (a demoted primary, a full `noeviction` server), rotated-credentials, slow-reconnect (over TLS), slow-server, failover-behind-a-stable-address and owed-invalidation cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. - **Leases for work that must run in one process at a time, and the sweeper runs under one** (`internal/coord/` (new: `coord.go`, `local.go`, `elect.go`, `coordtest/`, + tests), `internal/app/{app,wire}.go` (+ tests), `internal/config/backends.go`, `internal/ingest/sweeper.go`, `.github/labeler.yml`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development,ingest-pipeline}.md`, `docs/src/content/docs/configuration.mdx`, `config.yaml`, `AGENTS.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `coord.Coordinator` hands out named leases (`TryAcquire` → a `Term` with a strictly increasing fencing `Token`, a `Done` channel and `Resign`; `ErrHeld` while another holder's term is live), and `coord.RunElected` runs a loop only while its process holds the lease, resigning when the loop returns and campaigning again every 2s. `coord.Local` is the in-process implementation, and `coordtest.Conformance` is the suite every implementation runs — the NATS KV backend that lets several replicas share one queue comes next. The sweeper now runs through `RunElected` under the `sweeper` lease; with the in-process coordinator the one process always holds it, so nothing changes for a single-process deployment beyond one `coord: elected` log line at startup. `coord.backend` now selects the coordinator (`local`, the only value), so a `config.Config` built without `config.Load` must name it as well as the other three layers' backends. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 96618a477..2347ec727 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -115,11 +115,11 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **cache.go** — `Cache` interface: `Lookup`, `Set`, `Invalidate`, `InvalidateTenant`, `Close`, plus `QueryTimeToTTL`, which sets a result's TTL from how long its query took (10 s floor, 1 h ceiling). Every entry is one tenant's: `Lookup` takes the tenant, the caller's query key — `:query:`, built by the two cached handlers in `api/` (`queryCacheKey`, with the tenant read off the request's store — `settings.Store.Tenant`), which use it as their [singleflight](https://pkg.go.dev/golang.org/x/sync/singleflight) key too — and the `Namespace`s the result depends on — one for a structured query, none yet for a pipe (a pipe's table dependencies are [#343](https://github.com/Wave-RF/WaveHouse/pull/343)) — each of that tenant (another tenant's is `ErrForeignDependency`) and naming a table and scope by their raw names, which the cache escapes where it builds a key. `Lookup` returns the `Entry` (a nil value is a miss) and a `Snapshot` of the versions it read; on a miss the handler runs the query and passes that snapshot to `Set`, so a result is filed under the versions read *before* its query ran, and a write that lands while it runs orphans the fill rather than re-homing pre-write rows under the post-write versions ([#382](https://github.com/Wave-RF/WaveHouse/issues/382)). The snapshot is taken before any input a bump invalidates is chosen, the tenant's connection included: a reload that moves the tenant to another address or database runs `Pools.Reconcile` and then `InvalidateTenant` (a repoint that keeps both, such as a username or `tls` change, reads the same tables and bumps nothing), so a request that took the old pool files the old database's rows under a version that bump orphans, whether its `Set` lands before the bump or after. The singleflight leader's snapshot is the one used. `Set` returns an error only when the backend failed; a value the cache declines — larger than it keeps, a non-positive TTL, a zero snapshot — is not one. What a hit, a miss and a bump mean is pinned by the conformance suite every backend runs, `internal/testutil/cachetest`. - **local.go** — `LocalCache`, the in-process L1 on [Ristretto](https://github.com/dgraph-io/ristretto): one pool shared by every tenant (a heavier tenant holds more of it), sized by the boot config's `cache.l1_max_cost`. - **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: a namespace key is `..
.
.`, its fields joined with `keyenc.Join` so a dot in a table or scope name is escaped rather than read as a separator, and an entry's key, `|.||…` with the caller's query key escaped whole (its `:` become `%3A`), folds in the tenant's version and each dependency's namespace key and namespace version, so bumping a table (a scopeless write) or one scope — scope is reserved and empty today, so every write is the whole-table bump — orphans every dependent entry without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results, and the flat directory's single tenant simply carries the `0` prefix. `BumpTenant` (behind `InvalidateTenant`) advances the tenant version that leads every namespace key of one tenant, orphaning its every namespace and every cached query in one step (the tenant version is folded into every entry's key, so a pipe result with no dependencies is orphaned too) — a table no bump ever keyed included, which an enumeration of the index would miss — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. -- **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET` and `MGET`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot; the table and scope are escaped with `keyenc` after the fixed prefix, so a `:` in a name is never read as the separator. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, the hash over their escaped forms, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token (a string of another length, a list, a hash) is replaced, so a token lost to eviction, expiry, `FLUSHALL` or a restart without persistence is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. Restoring an RDB or AOF snapshot, or a backup, is not a loss but a rollback: the old tokens come back with the values filed under them, so what was invalidated since is served again until its TTL, as after a failover to a replica that missed the bumps. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. While this process owes a bump on any of a lookup's tokens, that lookup is a bypass that files nothing: the bump would orphan whatever it found (other processes, which cannot know, serve those entries until it lands). `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing, each attempt bounded by `DialTimeout` (1 s), and a cluster client's topology read after the handshake by the larger of it and `Timeout`, which is also how long a connection waits on a silent server before it is redialed. Every connection is replaced after a minute (`clientOption`; rueidis retries what was in flight), because one that outlives a failover behind a stable address stays on the demoted node, which answers but refuses writes; the replacement re-resolves the address, so a bypassed process reaches the new primary, and delivers the bumps it owes, within about that long. rueidis dials a replacement lazily, under the context of the operation that lands on it, so a reconnect (dial and handshake, TLS included) slower than `Timeout` fails that operation. The breaker's probe runs under `DialTimeout` plus `Timeout`, so such a reconnect still closes an open breaker; but rueidis spreads commands over several connections (up to four to one server, by `GOMAXPROCS`, and one per cluster node) and the probe reconnects only the one it lands on, so size `Timeout` above a reconnect, or operations that land on the others keep failing. Nothing selects this backend yet: `cache.backend` accepts only `local` until [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s E4 adds `redis`, the `cache.redis.*` settings and the wiring. +- **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET` and `MGET`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot; the table and scope are escaped with `keyenc` after the fixed prefix, so a `:` in a name is never read as the separator. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, the hash over their escaped forms, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token (a string of another length, a list, a hash) is replaced, so a token lost to eviction, expiry, `FLUSHALL` or a restart without persistence is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. Restoring an RDB or AOF snapshot, or a backup, is not a loss but a rollback — a restart after a crash that reloads the server's last save included, which stock Redis and Valkey make by default: the old tokens come back with the values filed under them, so what was invalidated since is served again until its TTL, as after a failover to a replica that missed the bumps. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. While this process owes a bump on any of a lookup's tokens, that lookup is a bypass that files nothing: the bump would orphan whatever it found (other processes, which cannot know, serve those entries until it lands). `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing, each attempt bounded by `DialTimeout` (1 s), and a cluster client's topology read after the handshake by the larger of it and `Timeout`, which is also how long a connection waits on a silent server before it is redialed. Every connection is replaced after a minute (`clientOption`; rueidis retries what was in flight), because one that outlives a failover behind a stable address stays on the demoted node, which answers but refuses writes; the replacement re-resolves the address, so a bypassed process reaches the new primary, and delivers the bumps it owes, within about that long. rueidis dials a replacement lazily, under the context of the operation that lands on it, bounding the dial (TLS included) by `DialTimeout` and then the handshake by `DialTimeout` again, so a reconnect slower than `Timeout` fails that operation. The breaker's probe gives its write twice `DialTimeout` for a reconnect on top of `Timeout`, so such a reconnect still closes an open breaker. The allowance is for a reconnect only: a probe write slower than `Timeout` is repeated under `Timeout`, and the repeat decides, so a server answering slower than `Timeout` stays bypassed. But rueidis spreads commands over several connections (up to four to one server, by `GOMAXPROCS`, and one per cluster node) and the probe reconnects only the one it lands on, so size `Timeout` above a reconnect, or operations that land on the others keep failing. Nothing selects this backend yet: `cache.backend` accepts only `local` until [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s E4 adds `redis`, the `cache.redis.*` settings and the wiring. - **redis_codec.go** — the key schema (`tokenKeys`, `bumpKeys`, `valueKey`, which take a `Namespace`'s raw names and escape them) and the value frame: format, flags, expiry, the token list, the payload, zstd-compressed from `CompressMinBytes` when that is smaller. A stored value is capped at `MaxValueBytes` and a decoded one at eight times that, which refuses a zip bomb planted in a shared server; an unknown format, as a newer process writes during a rolling upgrade, is a miss. - **breaker.go** — the circuit breaker's state machine: `BreakerThreshold` (5) consecutive failures open it, `trip` opens it at once, and while it is open every operation skips the server; once `BreakerOpenFor` (5 s) has passed, `allow` hands one caller the probe, and only the probe's success closes it. What feeds it is `redis.go`'s. `record` counts a transport failure or a timeout against the server, and trips the breaker on an error reply that means no bump can land: one saying the server takes no writes right now, as `refusesWork` lists them — `READONLY` (a demoted primary), `MASTERDOWN`, `OOM` (full under `maxmemory-policy noeviction`, Redis's default, so give the cache an evicting policy), `NOREPLICAS`, `MISCONF`, `LOADING`, `BUSY`, `CLUSTERDOWN` — or one refusing the credentials, as `rejectsCredentials` lists them — `WRONGPASS`, `NOAUTH`, what a new connection's handshake meets after a password rotation — which it logs at `ERROR`. Any other reply, an error reply about one key or command (`WRONGTYPE`, `NOPERM`, `TRYAGAIN`) or one the backend cannot use included, counts as a success, and a caller that gave up first counts as nothing. The probe (`probe`) is a write, `SET :probe`, so a server that answers but refuses writes stays bypassed. - **pending.go** — the invalidations owed (`pendingBumps`): kept per token key, repeats coalescing, and past `PendingMax` (100,000) keys collapsed to one tenant bump per affected tenant; `owesAny` tells `Lookup` which lookups to hold. `redis.go` delivers them: `Invalidate` sends its bumps 1,000 to a round trip and defers those the server does not take, and `drainLoop` retries them through `drain` until they land — the first one owed at once, then with backoff from 100 ms to 10 s, and while the breaker is open at each probe, which the loop starts when due, so a process that makes no lookups (ingest only) recovers as soon as one that does. Landing late is still correct: a fresh token orphans the pre-write entries and any fill made meanwhile. `Close` makes one last attempt at the bumps still owed, past the breaker (an open one is why they are owed); what that attempt cannot deliver is lost, and the entries they would orphan are served until their TTL — the same failure as a worker stopping between an insert and its invalidation. -- **metrics.go** — the shared backend's instruments (meter `wavehouse-cache`, `backend="redis"`, no tenant attribute): `wavehouse_cache_lookups_total` by `result` (`hit`, `miss`, `stale` — filed under since-bumped tokens, `bypass` — server skipped, or held by a bump this process owes, `error`), `wavehouse_cache_op_duration_seconds` by `op` (`lookup`, `set`, `invalidate`), `wavehouse_cache_breaker_open` (1 while bypassed, including before the first connection), `wavehouse_cache_invalidations_total` by `result` (`ok` counts every bump that lands, retried ones included, and `deferred` each time one is put off, a repeat of one already owed included, so the two overlap), `wavehouse_cache_invalidations_pending` (bumps owed: entries they would orphan may be served stale meanwhile), `wavehouse_cache_value_bytes` (stored size), `wavehouse_cache_oversize_total` and `wavehouse_cache_set_failures_total` by `reason` (`oom`, `timeout`, `other`). +- **metrics.go** — the shared backend's instruments (meter `wavehouse-cache`, `backend="redis"`, no tenant attribute): `wavehouse_cache_lookups_total` by `result` (`hit`, `miss`, `stale` — filed under since-bumped tokens, `bypass` — server skipped, or held by a bump this process owes, `error`), `wavehouse_cache_op_duration_seconds` by `op` (`lookup`, `set`, `invalidate`), `wavehouse_cache_breaker_open` (1 while bypassed, including before the first connection), `wavehouse_cache_invalidations_total` by `result` (`ok` counts every bump that lands, retried ones included, and `deferred` each bump an invalidation could not deliver when made, a repeat of one already owed included; a failed retry is not counted again, so the two overlap), `wavehouse_cache_invalidations_pending` (bumps owed: entries they would orphan may be served stale meanwhile), `wavehouse_cache_value_bytes` (stored size), `wavehouse_cache_oversize_total` and `wavehouse_cache_set_failures_total` by `reason` (`oom`, `timeout`, `other`). - **cachetest** (`internal/testutil/cachetest`) — `Run(t, factory, Options)`, the conformance suite every `Cache` backend runs, each case on a fresh cache from the factory: what a hit, a miss and each kind of bump mean, independent of where entries and versions live. `Options` describes what a backend can do beyond the `Cache` contract, and leaving one unset skips the cases it enables: `MaxValueBytes` opens the oversize case; `NewPair` returns two instances over one shared store, for the cross-instance cases a shared backend must run; `Entries` counts a cache's entries, and without it the zero-snapshot case — no `Lookup` reads the key a zero `Snapshot` would land under, so only a count shows that one stored nothing — is skipped. ### `config/` — Configuration diff --git a/internal/cache/metrics.go b/internal/cache/metrics.go index c1dc12dc0..855843443 100644 --- a/internal/cache/metrics.go +++ b/internal/cache/metrics.go @@ -45,7 +45,7 @@ func newMetrics(backend string, breakerOpen func() bool, pending func() int) (*m metric.WithDescription("Shared-cache round-trip time by op: lookup, set, invalidate"), metric.WithUnit("s"), metric.WithExplicitBucketBoundaries(.0001, .00025, .0005, .001, .0025, .005, .01, .025, .05, .1, .25)) m.invalidation, errs[2] = meter.Int64Counter("wavehouse_cache_invalidations_total", - metric.WithDescription("Version-token bumps by result: ok counts every bump that lands, retried ones included; deferred counts each time one is put off, a repeat of one already owed included. They overlap: wavehouse_cache_invalidations_pending is what is still owed")) + metric.WithDescription("Version-token bumps by result: ok counts every bump that lands, retried ones included; deferred counts each bump an invalidation could not deliver when made, a repeat of one already owed included; a failed retry is not counted again. They overlap: wavehouse_cache_invalidations_pending is what is still owed")) m.valueBytes, errs[3] = meter.Int64Histogram("wavehouse_cache_value_bytes", metric.WithDescription("Size of each value written to the shared cache, after compression"), metric.WithUnit("By"), metric.WithExplicitBucketBoundaries(256, 1<<10, 4<<10, 16<<10, 64<<10, 256<<10, 1<<20, 4<<20)) diff --git a/internal/cache/redis.go b/internal/cache/redis.go index b722e446a..92719d80e 100644 --- a/internal/cache/redis.go +++ b/internal/cache/redis.go @@ -188,10 +188,12 @@ func (c RedisConfig) clientOption() rueidis.ClientOption { // under the tenant's hash tag; a bump sets a fresh one. A value carries the // tokens it was computed under and is a hit only while they are all still // current, so a lost token (eviction, expiry, a restart without persistence) -// can only cause misses. Restoring a snapshot is not a loss but a rollback: -// the old tokens return with their values. A lookup is one round trip. The -// server failing or timing out is a miss, a skipped fill and a deferred -// invalidation — never a failed query. +// can only cause misses. Restoring a snapshot is not a loss but a rollback — +// a restart after a crash that reloads the server's last save included, +// which stock Redis and Valkey make by default: the old tokens return with +// their values. A lookup is one round trip. The server failing or timing +// out is a miss, a skipped fill and a deferred invalidation — never a +// failed query. type RedisCache struct { cfg RedisConfig opt rueidis.ClientOption @@ -312,18 +314,30 @@ func (r *RedisCache) conn() rueidis.Client { } // probe decides whether an open breaker closes. It writes: a server that -// answers but refuses writes (refusesWork) would take no bump either. rueidis -// redials under the calling operation's context, so the probe's budget is -// DialTimeout for a reconnect plus Timeout for the write: a reconnect slower -// than Timeout fails the operations waiting on it, but not the probe. +// answers but refuses writes (refusesWork) would take no bump either. +// rueidis redials under the calling operation's context, bounding the dial +// (TLS included) by DialTimeout and then the handshake by DialTimeout again, +// so the first write has twice DialTimeout for a reconnect on top of +// Timeout: a reconnect slower than Timeout fails the operations waiting on +// it, but not the probe. The allowance is for a reconnect only: a first +// write slower than Timeout is repeated under Timeout, and the repeat +// decides, so a server answering slower than Timeout stays bypassed. func (r *RedisCache) probe(c rueidis.Client) { - ctx, cancel := context.WithTimeout(r.ctx, r.cfg.DialTimeout+r.cfg.Timeout) - defer cancel() - r.record(r.ctx, c.Do(ctx, c.B().Set().Key(r.cfg.KeyPrefix+":probe").Value("1").Ex(time.Minute).Build()).Error()) + set := func(budget time.Duration) error { + ctx, cancel := context.WithTimeout(r.ctx, budget) + defer cancel() + return c.Do(ctx, c.B().Set().Key(r.cfg.KeyPrefix+":probe").Value("1").Ex(time.Minute).Build()).Error() + } + start := time.Now() + err := set(2*r.cfg.DialTimeout + r.cfg.Timeout) + if time.Since(start) > r.cfg.Timeout { + err = set(r.cfg.Timeout) + } + r.record(r.ctx, err) if r.breaker.isOpen() { return } - slog.InfoContext(ctx, "cache: redis reachable again; cache back in use") + slog.InfoContext(r.ctx, "cache: redis reachable again; cache back in use") r.nudge() } diff --git a/internal/cache/redis_integration_test.go b/internal/cache/redis_integration_test.go index d40a2fae7..13ff66557 100644 --- a/internal/cache/redis_integration_test.go +++ b/internal/cache/redis_integration_test.go @@ -5,10 +5,14 @@ package cache_test import ( "bytes" "context" + "crypto/ecdsa" + "crypto/elliptic" "crypto/rand" "crypto/tls" + "crypto/x509" "fmt" "io" + "math/big" "net" "slices" "strconv" @@ -588,24 +592,28 @@ func (c unanswered) Write(b []byte) (int, error) { // proxy forwards each connection it accepts to the address target holds at // that moment, so a switch moves new connections only, as a stable DNS name -// or a proxy does after a failover. It holds a new connection for delay -// before forwarding it, as a slow dial and handshake would. +// or a proxy does after a failover. Given a TLS config it terminates TLS. +// It holds a new connection for delay before its TLS handshake and again +// before forwarding it, as a slow dial and a slow handshake would, and +// holds everything the server sends for replyDelay, as a slow server would. type proxy struct { - addr string // to dial - target atomic.Pointer[string] - delay atomic.Int64 // a time.Duration + addr string // to dial + serverTLS *tls.Config + target atomic.Pointer[string] + delay atomic.Int64 // a time.Duration + replyDelay atomic.Int64 // a time.Duration mu sync.Mutex conns []net.Conn } -func newProxy(t *testing.T, target string) *proxy { +func newProxy(t *testing.T, target string, serverTLS *tls.Config) *proxy { t.Helper() var lc net.ListenConfig ln, err := lc.Listen(context.Background(), "tcp", "127.0.0.1:0") require.NoError(t, err) t.Cleanup(func() { _ = ln.Close() }) - p := &proxy{addr: ln.Addr().String()} + p := &proxy{addr: ln.Addr().String(), serverTLS: serverTLS} p.target.Store(&target) go func() { for { @@ -624,6 +632,14 @@ func newProxy(t *testing.T, target string) *proxy { func (p *proxy) serve(c net.Conn) { defer func() { _ = c.Close() }() + if p.serverTLS != nil { + time.Sleep(time.Duration(p.delay.Load())) + tc := tls.Server(c, p.serverTLS) + if tc.HandshakeContext(context.Background()) != nil { + return + } + c = tc + } time.Sleep(time.Duration(p.delay.Load())) var d net.Dialer u, err := d.DialContext(context.Background(), "tcp", *p.target.Load()) @@ -632,7 +648,42 @@ func (p *proxy) serve(c net.Conn) { } defer func() { _ = u.Close() }() go func() { _, _ = io.Copy(u, c); _ = u.Close() }() - _, _ = io.Copy(c, u) + buf := make([]byte, 32<<10) + for { + n, err := u.Read(buf) + if n > 0 { + time.Sleep(time.Duration(p.replyDelay.Load())) + if _, err := c.Write(buf[:n]); err != nil { + return + } + } + if err != nil { + return + } + } +} + +// selfSigned returns a server TLS config holding a fresh certificate for +// 127.0.0.1, and a client config that trusts it. +func selfSigned(t *testing.T) (srv, cli *tls.Config) { + t.Helper() + key, err := ecdsa.GenerateKey(elliptic.P256(), rand.Reader) + require.NoError(t, err) + tmpl := &x509.Certificate{ + SerialNumber: big.NewInt(1), + NotBefore: time.Now().Add(-time.Minute), + NotAfter: time.Now().Add(time.Hour), + IPAddresses: []net.IP{net.IPv4(127, 0, 0, 1)}, + ExtKeyUsage: []x509.ExtKeyUsage{x509.ExtKeyUsageServerAuth}, + } + der, err := x509.CreateCertificate(rand.Reader, tmpl, tmpl, &key.PublicKey, key) + require.NoError(t, err) + cert, err := x509.ParseCertificate(der) + require.NoError(t, err) + roots := x509.NewCertPool() + roots.AddCert(cert) + return &tls.Config{Certificates: []tls.Certificate{{Certificate: [][]byte{der}, PrivateKey: key}}, MinVersion: tls.VersionTLS12}, + &tls.Config{RootCAs: roots, MinVersion: tls.VersionTLS12} } // drop closes every connection p carries, so the client must reconnect. @@ -670,7 +721,7 @@ func TestRedis_FailoverBehindAStableAddress(t *testing.T) { return err == nil && strings.Contains(info, "master_link_status:up") }, 30*time.Second, 50*time.Millisecond, "the replica syncs") - fwd := newProxy(t, primary.addr) + fwd := newProxy(t, primary.addr, nil) stable := &server{addr: fwd.addr, mode: cache.RedisStandalone} prefix := uniquePrefix() a := open(t, stable, prefix, func(c *cache.RedisConfig) { @@ -899,31 +950,66 @@ func TestRedis_RestoredSnapshotIsARollback(t *testing.T) { // A reconnect slower than the op timeout (a dial and TLS handshake across // zones) fails the operations waiting on it, since rueidis dials under their -// context; they open the breaker, and the probe, whose budget covers a dial, -// completes the reconnect and closes it. Only the connection the probe -// lands on: rueidis spreads commands over several, and the rest still -// reconnect under the op timeout. +// context; they open the breaker, and the probe, whose budget covers a +// reconnect, completes it and closes the breaker. rueidis bounds the dial, +// TLS included, by the dial timeout and then the handshake by it again, so +// here each takes most of it: together they outlast one dial timeout plus +// the op timeout. Only the connection the probe lands on reconnects: +// rueidis spreads commands over several, and the rest still reconnect under +// the op timeout. func TestRedis_ProbeFitsASlowReconnect(t *testing.T) { t.Parallel() ctx := context.Background() s := startRedis(t) - p := newProxy(t, s.addr) - const timeout = 100 * time.Millisecond + srvTLS, cliTLS := selfSigned(t) + p := newProxy(t, s.addr, srvTLS) + const timeout, dialTimeout = 100 * time.Millisecond, time.Second a := open(t, &server{addr: p.addr, mode: cache.RedisStandalone}, uniquePrefix(), func(c *cache.RedisConfig) { - c.Timeout, c.DialTimeout = timeout, 2*time.Second + c.Timeout, c.DialTimeout, c.TLS = timeout, dialTimeout, cliTLS c.BreakerThreshold, c.BreakerOpenFor = 1, 200*time.Millisecond }) deps := []cache.Namespace{{Tenant: "acme", Table: "events"}} _, _, err := a.Lookup(ctx, "acme", "q", deps) require.NoError(t, err) - p.delay.Store(int64(5 * timeout)) + p.delay.Store(int64(dialTimeout * 7 / 10)) p.drop() _, _, err = a.Lookup(ctx, "acme", "q", deps) require.Error(t, err, "the reconnect does not fit in the lookup's timeout") require.True(t, cache.Bypassed(a)) + require.Eventually(t, func() bool { return !cache.Bypassed(a) }, 20*time.Second, 10*time.Millisecond, + "the probe reconnects within twice the dial timeout") +} + +// A server that answers, but slower than the op timeout, fails every +// operation, so the probe's reconnect allowance must not close the breaker +// on it: a probe write slower than the op timeout is repeated under it, and +// only a repeat that lands closes the breaker. +func TestRedis_ProbeKeepsASlowServerBypassed(t *testing.T) { + t.Parallel() + ctx := context.Background() + s := startRedis(t) + p := newProxy(t, s.addr, nil) + const timeout = 100 * time.Millisecond + a := open(t, &server{addr: p.addr, mode: cache.RedisStandalone}, uniquePrefix(), func(c *cache.RedisConfig) { + c.Timeout, c.DialTimeout = timeout, time.Second + c.BreakerThreshold, c.BreakerOpenFor = 1, 200*time.Millisecond + }) + deps := []cache.Namespace{{Tenant: "acme", Table: "events"}} + _, _, err := a.Lookup(ctx, "acme", "q", deps) + require.NoError(t, err) + + p.replyDelay.Store(int64(3 * timeout)) + _, _, err = a.Lookup(ctx, "acme", "q", deps) + require.Error(t, err, "a reply slower than the timeout fails the lookup") + require.True(t, cache.Bypassed(a)) + // Several probes, each answered within the reconnect allowance. + for end := time.Now().Add(4 * time.Second); time.Now().Before(end); time.Sleep(5 * time.Millisecond) { + require.True(t, cache.Bypassed(a), "a probe answered slower than the op timeout closed the breaker") + } + p.replyDelay.Store(0) require.Eventually(t, func() bool { return !cache.Bypassed(a) }, 10*time.Second, 10*time.Millisecond, - "the probe reconnects within the dial timeout") + "a probe answered in time closes it") } // A password rotated under a running process refuses its next connection's From e2366d7361d30db00e2d244d226d6bf56d52bd05 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 12:59:30 -0400 Subject: [PATCH 68/79] fix(api): see a write behind a .-led number or an EXECUTE AS prefix ClickHouse ends a number led by `.` at its digits and exponent, so in `WITH 1 AS a, .5INSERT INTO t SELECT a` it reads `.5` then INSERT. The classifier stepped over the `.` and read `5INSERT` as one word, so the write ran through Query, which ran it and then failed the call. hasTopLevelInsertInto now reads a `.`-led number as ClickHouse's lexer does: digits with `_` between two of them, then an optional exponent. `EXECUTE AS ` runs the statement as that user, and IsMutation classified it by EXECUTE, so `EXECUTE AS u INSERT ...` also ran through Query. The prefix is now looked through, for a bare, quoted or heredoc user and an optional `@host`, and the statement after it is classified by the same rules. A bare `EXECUTE AS u` switches the session's user and returns no result set, so it goes through Exec. The EXPLAIN AST oracle maps an ExecuteAsQuery root by the statement it runs. The comment on a read's `... AS insert INTO OUTFILE` now says what happens to it: it is classified as a write and answers [] uncached. Docs: pipes.mdx states the WITH rule as it is, looks through EXECUTE AS, and gives UNDROP and MOVE their real consequence (they fail after running, which the SDK may retry); architecture.md and AGENTS.md no longer say non-insert mutations must use /v1/ops/query and then name a second path. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- AGENTS.md | 2 +- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/pipes.mdx | 2 +- internal/api/clickhouse_exec.go | 95 +++++++++++++++++++++++-- internal/api/clickhouse_exec_test.go | 18 +++-- internal/testutil/mutationtest/cases.go | 33 +++++++++ tests/integration/ismutation_test.go | 48 +++++++++---- 8 files changed, 173 insertions(+), 29 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index cff241d62..69d0d5bad 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -38,7 +38,7 @@ Twenty internal packages under `internal/` (plus `internal/testutil/` for shared - **`coord/`** — leases for work that must run in one process at a time: `Coordinator.TryAcquire(ctx, name)` → a `Term` (fencing `Token`, strictly increasing per name; `Done`/`Err`, `ErrLost` on loss; `Resign`), `ErrHeld` while another holder's — or this coordinator's own — term is live; `RunElected` runs a loop only while holding its lease, resigning when the loop returns and campaigning again every `RetryPeriod`. `Local` is the in-process implementation (first taker wins, never expires; `Peer` is a second handle over the same table for tests); every implementation runs `coordtest.Conformance`. Imports only the standard library, so a distributed backend lives beside its connection (NATS KV in `internal/mq`). `internal/app`'s `wireCoord` opens the one `coord.backend` selects and the sweeper runs through `RunElected` under the `sweeper` lease - **`dedupe/`** — `Deduplicator` interface → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) -- **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy; the only other write path is an operator-authored pipe (#386). A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) +- **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`, so non-admin callers never reach the proxy) or through an operator-authored pipe that writes, gated only by its `allowed_roles` (#386). A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) - **`keyenc/`** — the one escaping composite keys are built from: `Escape` keeps `[A-Za-z0-9_-]` (exactly the tenant-id grammar, so a tenant id is its own escaped form) and writes every other byte as `%XX`, `Unescape` is `url.PathUnescape` (lenient: either hex case, and a byte left unescaped reads as itself, so a `%2D` an earlier build wrote still reads), `Join`/`AppendJoin` escape each field and put a separator between them (they panic on no fields, and on a separator the escaping could write or one outside ASCII) and `Split` reverses them. The package that builds a key takes raw names and escapes them itself, so no caller has to and no field reaches a key unescaped: NATS subjects (`Join`/`Split` after the verbatim tenant) and the cache's keys (`cache.VersionManager`) use it; changing what it keeps orphans every stored key - **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal, `ErrUnavailable` a broker that cannot be reached — both a `503`, with `Retry-After` `30` and `5`), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`, `deadletter.go`), whose subject tokens are escaped by the shared `internal/keyenc`; `internal/app` constructs it and hands everything else a `mq.Broker`. Every implementation passes the conformance suite in `internal/mq/mqtest` (`mqtest.Run`), which states the `Broker` contract as behavior; a new backend runs it from its own test, with `mqtest.Caps` only where its semantics legitimately differ - **`observability/`** — OpenTelemetry pipeline: `InitProvider` wires trace/metric/log providers via OTLP gRPC (each signal independently gated). A top-level `Prometheus` config block drives an optional `/metrics` scrape endpoint that runs independently of OTLP push — standalone (Alloy/Mimir scrape, no collector), alongside OTLP, or off. `NewLogger` produces a slog handler that fans out to stdout AND OTLP (stdout always 100%, OTLP sample-rate-aware). `TraceHandler` injects trace_id/span_id from active spans. `tracer.go` provides W3C trace context propagation over message headers (`InjectHeaders`/`ExtractHeaders` on a plain header map; `internal/mq` injects on every publish and extracts on the `Subscribe` path, so this package never sees a NATS type). diff --git a/CHANGELOG.md b/CHANGELOG.md index fdf65b628..b2a744936 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -86,7 +86,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Fixed - **A pipe that writes runs on every call instead of being answered from the cache** (`internal/api/{pipes,ch_errors}.go` (+ tests), `docs/src/content/docs/{pipes.mdx,api.md,architecture.md,configuration.mdx,settings-directory.mdx,ingest-pipeline.md,sdk/pipes.md,sdk/reference.md}`, `clients/ts/src/pipes.ts` (doc comment), `internal/{settings/settings,app/wire}.go` (comments), `AGENTS.md`): fixes [#386](https://github.com/Wave-RF/WaveHouse/issues/386). `/v1/pipes/{name}` sent a write's SQL to ClickHouse through `Exec`, but still cached the `[]` it returned and coalesced identical calls in flight, so a repeat within the TTL answered `200` without executing and concurrent identical calls became one write — silently dropped writes, and with a shared cache ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)) on every instance. A pipe whose bound SQL `IsMutation` classifies as a write — the same classifier that picks `Exec` — now skips the cache lookup, the fill and singleflight, and answers `X-Cache: BYPASS` with `Cache-Control: no-store`, so an HTTP cache in front of a `GET` cannot drop the write either. Classification stays automatic rather than a declared pipe property, so an operator cannot forget to mark one, and costs no ClickHouse round trip. A failed write answers with the status and `code` a failed read gets (see the ClickHouse-errors entry below), but always `retryable: false` and with no `Retry-After`, `503 clickhouse.unavailable` included: the statement may have run, so the SDK does not retry it. A write refused before it is sent, the tenant on no pool, keeps its `503` with `Retry-After: 30`. Read pipes are unchanged. Not in this fix: a write pipe still does not invalidate cached reads of the table it writes ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). -- **The write classifier skips whitespace, comments and quoted text the way ClickHouse's lexer does** (`internal/api/clickhouse_exec.go` (+ tests), `internal/testutil/mutationtest` (new), `tests/integration/ismutation_test.go` (new), `AGENTS.md`): `IsMutation` picks `Exec` for a write, and since [#386](https://github.com/Wave-RF/WaveHouse/issues/386) keeps a write pipe out of the cache. It missed a write behind a backslash-escaped quote (`'it\'s'`, and the same inside `"…"` and `` `…` ``), a heredoc (`$$ ( $$`, `$tag$ … $tag$`), a curly-quoted literal or identifier (`‘(’`, `“c(d”`), a `//` line comment, a nested block comment (`/* a /* b */ SELECT */ INSERT …`), or leading whitespace other than space, tab, CR and LF: `\v`, `\f`, a no-break space, a byte-order mark, and the other Unicode spaces ClickHouse skips. A missed write went through `Query`, which ran it and then failed the call with a `5xx` the TypeScript SDK retries, so one call could write three times. The same gaps, and a word led by `_` (`_delete`) whose tail was read as a verb, could make a read look like a write, which runs through `Exec` and answers `[]`. After a `WITH` list, which ClickHouse follows only with `SELECT`, a FROM-first `SELECT` or `INSERT INTO`, a name spelled like a keyword was taken for the statement: `WITH 'd' AS desc INSERT …` and `WITH 1 AS select INSERT …` ran as reads, and `WITH 1 AS set SELECT set` and `WITH 1 AS x FROM system.one SELECT x` as writes. A `WITH`-led statement is now a write exactly when it holds `INSERT INTO` outside parentheses. The classifier, exported as `IsMutation` for it, is now checked against the pinned ClickHouse's own parser (`EXPLAIN AST`) in the integration suite: every test case, and every ClickHouse keyword as a `WITH` list's name ahead of each statement a `WITH` list can lead. +- **The write classifier skips whitespace, comments and quoted text the way ClickHouse's lexer does, classifies a `WITH`-led statement by `INSERT INTO` alone, and looks through `EXECUTE AS`** (`internal/api/clickhouse_exec.go` (+ tests), `internal/testutil/mutationtest` (new), `tests/integration/ismutation_test.go` (new), `docs/src/content/docs/pipes.mdx`, `AGENTS.md`): `IsMutation` picks `Exec` for a write, and since [#386](https://github.com/Wave-RF/WaveHouse/issues/386) keeps a write pipe out of the cache. It missed a write behind a backslash-escaped quote (`'it\'s'`, and the same inside `"…"` and `` `…` ``), a heredoc (`$$ ( $$`, `$tag$ … $tag$`), a curly-quoted literal or identifier (`‘(’`, `“c(d”`), a `//` line comment, a nested block comment (`/* a /* b */ SELECT */ INSERT …`), a number led by `.` with the verb glued to it (`WITH 1 AS a, .5INSERT INTO t …`, which ClickHouse reads as `.5` then `INSERT`), an `EXECUTE AS ` prefix (`EXECUTE AS u INSERT …`), or leading whitespace other than space, tab, CR and LF: `\v`, `\f`, a no-break space, a byte-order mark, and the other Unicode spaces ClickHouse skips. A missed write went through `Query`, which ran it and then failed the call with a `5xx` the TypeScript SDK retries, so one call could write three times. The same gaps, and a word led by `_` (`_delete`) whose tail was read as a verb, could make a read look like a write, which runs through `Exec` and answers `[]`. After a `WITH` list, which ClickHouse follows only with `SELECT`, a FROM-first `SELECT` or `INSERT INTO`, a name spelled like a keyword was taken for the statement: `WITH 'd' AS desc INSERT …` and `WITH 1 AS select INSERT …` ran as reads, and `WITH 1 AS set SELECT set` and `WITH 1 AS x FROM system.one SELECT x` as writes. A `WITH`-led statement is now a write exactly when it holds `INSERT INTO` outside parentheses. The classifier, exported as `IsMutation` for it, is now checked against the pinned ClickHouse's own parser (`EXPLAIN AST`) in the integration suite: every test case, and every ClickHouse keyword as a `WITH` list's name ahead of each statement a `WITH` list can lead. - **A failed ClickHouse query answers by what went wrong, not a flat `500`/`502`** (`internal/api/ch_errors.go` (new, + tests), `internal/api/{errors,query,structured_query,pipes,schema,ch_settings}.go`, `internal/chconn/errclass.go` (`HTTPStatus` exported), `clients/ts/src/errors.ts` (+ tests), `tests/integration/query_errors_test.go` (new), `tests/integration/query_limits_test.go`, `internal/app/app_test.go`, `tests/e2e/sdk/{admin,query}.test.ts`, `AGENTS.md`, `docs/src/content/docs/{api,architecture}.md`, `docs/src/content/docs/{access-control,configuration}.mdx`, `docs/src/content/docs/sdk/{reference.md,index.mdx}`): fixes [#403](https://github.com/Wave-RF/WaveHouse/issues/403) and [#271](https://github.com/Wave-RF/WaveHouse/issues/271), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). ClickHouse answers a syntax error, a missing grant and an overloaded server alike with HTTP `500`, so `/v1/ops/query` turned a bad statement into a `502` and `/v1/query` and pipes into a `500` the SDK retried. All three now class the failure with `chconn.Classify` through one helper, `writeCHError`: a statement ClickHouse refused is `400 clickhouse.rejected`; a query over a rows/bytes limit, the role's own memory cap, or its time cap where that is no longer than `query_timeout` is `400 clickhouse.limit_exceeded`; `ACCESS_DENIED` is `403 clickhouse.access_denied`; credentials, user or database refused, or a redirect or `4xx` with no exception code from whatever fronts ClickHouse, is `502 clickhouse.misconfigured`; ClickHouse down, unreachable or overloaded is `503 clickhouse.unavailable` with `Retry-After: 5`; a failure with no verdict stays `500` (`502` on the proxy) as `clickhouse.unknown`. The error envelope gains `code` and `retryable` next to `error` on these responses — additive. A role with `max_execution_time` now queries with no context deadline and a cancel two seconds past the cap instead: clickhouse-go overwrote the cap's `max_execution_time` with deadline+5s for any deadline over 1s, so an overrun came back as a bare deadline, indistinguishable from waiting for a pooled connection; ClickHouse now enforces the cap itself and reports `TIMEOUT_EXCEEDED`. `POST /v1/ops/schema/refresh` against an unreachable ClickHouse is a `503` with `Retry-After` instead of a `500`. **SDK:** `WaveHouseError.code` and `retryable` now take the server's `code`/`retryable` when the body has them (`HTTP_` and "5xx retries" otherwise), so a rejected query is `clickhouse.rejected` rather than `HTTP_500`, and is not retried. - **An unavailable ClickHouse is retried with backoff instead of dead-lettering every row** (`internal/chconn/errclass.go` (new, + tests), `internal/ingest/{worker,backoff}.go` (`backoff.go` new, + tests), `internal/mq/{mq,embedded}.go`, `internal/testutil/mocks.go`, `tests/integration/ingest_outage_test.go` (new), `AGENTS.md`, `README.md`, `docs/src/content/docs/{ingest-pipeline,architecture,api,deployment,why-wavehouse}.md`, `docs/src/content/docs/{settings-directory,index,access-control}.mdx`): workstream A of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). A failed batch insert used to go through row-by-row isolation whatever the failure, so a ClickHouse that was down, overloaded or read-only failed every row twice and parked the whole batch on the DLQ. `chconn.Classify` now classes the failure first — `Rejected` (any ClickHouse exception code outside the availability and credential lists: the server read the row and refused it), `Unavailable` (connection refused/reset, timeouts, `TOO_MANY_SIMULTANEOUS_QUERIES`, `SERVER_OVERLOADED`, `MEMORY_LIMIT_EXCEEDED`, `TOO_MANY_PARTS`, `READONLY`, `TABLE_IS_READ_ONLY`, `KEEPER_EXCEPTION`, …), `Denied` (`AUTHENTICATION_FAILED`, `ACCESS_DENIED`, …) or `Unknown` (no code, no recognizable transport failure). Only `Rejected` is isolated and dead-lettered as before, and a multi-row batch refused with `TOO_MANY_PARTS` or `MEMORY_LIMIT_EXCEEDED` is split row by row first (`chconn.Splittable`), because a batch spanning too many partitions or too much memory can fail where each of its rows inserts; every other class hands the batch back to the queue with a delayed nak (`mq.Message.NakWithDelay`, new) under a jittered 1 s → 30 s backoff shared by every table on the same ClickHouse pool (a failure of one table — read-only, too many parts or mutations, a grant missing on it, `chconn.TableScoped` — backs off that table alone), which turns rows away without a request while it runs and probes once per window, and ClickHouse going away mid-isolation stops isolation and retries the rows it had not settled. Counted by the new `wavehouse_ingest_retries_total{table, reason}`; logged at `WARN` when an outage starts and at most every 30 s during it. A long outage now shows as a growing ingest stream and, at `mq.max_bytes_gb`, ingest `503`s — not as a full DLQ; a lasting failure of one table holds back its tenant's other tables once its waiting rows reach `maxAckPending`. Retried rows come back out of arrival order, which matters only to a `ReplacingMergeTree` without a version column or a `CollapsingMergeTree`. - **Schema discovery's retry loop jitters its backoff** (`internal/discovery/discovery.go` (+ tests), `internal/app/wire.go`, `internal/api/errors.go`, `AGENTS.md`, `docs/src/content/docs/{architecture,api,deployment}.md`): `RetryRefresh` slept exactly `2s * 2^n` capped at 60s, so instances retrying against one recovering ClickHouse fired in lockstep, every 60s on the same second. Each sleep is now drawn uniformly from below the backoff (full jitter), spreading the retries over the whole window and halving the mean wait — so a failing tenant's retries, their log lines and `wavehouse_schema_refresh_failures_total` come about twice as often ([#141](https://github.com/Wave-RF/WaveHouse/issues/141)). diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 6b8b6a10e..976e645fa 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -149,7 +149,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ ### `ingest/` — Ingest Pipeline, DLQ & Sweeping -- **worker.go** — `StartIngestWorker` launches an ingest pipeline: a durable `buffer-consumer` consumer of the ingest queue (created through `mq.ConsumerManager`) reads events, batches them per tenant table — the tenant read off each message's `mq.Topic` — and performs bulk INSERTs to ClickHouse. The pipeline is **insert-only**. The wire format `EventMessage` carries `{table_name, scope, received_timestamp, format, columns, row}` — the row positionally as one `JSONCompactEachRow` line, with `columns` naming its positions (the table's insertable columns — a computed one cannot be named in an `INSERT`); the worker batches per (tenant, table, column list) and writes `INSERT INTO … (cols) FORMAT JSONCompactEachRow`. It accepts any table name (events are addressed by `mq.Topic{Tenant, Table, Scope}` with raw names; `internal/mq` encodes them into subject tokens), then bulk-INSERTs. The embedded NATS server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only publishers that can reach the ingest queue are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (`policy.admin_role`) — see the Query Path section below; the `/v1/ops/*` `RequireAdmin` middleware enforces the check at the API layer, so a no/invalid-token request (resolved to `default_role`, not admin in a production config) never reaches the proxy; the only other write path is an operator-authored [pipe that writes](/pipes#pipes-that-write). A batch whose tenant has no ClickHouse connection (no longer served, or no pool could be opened for it, such as by the connection ceiling) is never tried: no row of it could pass, so `parkBatch` takes it to the DLQ switch whole, logging once per batch rather than twice per row. Otherwise a bulk-insert failure is first classed by `chconn.Classify`: a ClickHouse that cannot take the insert (unavailable, denied, or no verdict at all) sends the batch back to the MQ for a delayed redelivery (`retryLater` → `mq.Message.NakWithDelay`), under a backoff shared by every table on the same pool (a failure of one table — read-only, too many parts — backs off that table alone), and never to the DLQ — the same when it stops answering mid-isolation. Only when ClickHouse rejects the batch, or refuses a multi-row batch for its size (`chconn.Splittable`: too many partitions for one INSERT, the memory limit), is it re-inserted row by row: rows that succeed are acked, and only the rows ClickHouse rejects again are routed to the DLQ (`sendToDLQ` → `mq.DeadLetterer.DeadLetter`), which parks the as-published `EventMessage` envelope under the topic it arrived on (`dlq.{tenant}.{table}` subjects inside `internal/mq`) with the failure context in `X-DLQ-*` headers when the tenant's `dlq.enabled` is on for the table — see [Ingest Pipeline](/ingest-pipeline) for the worker internals. +- **worker.go** — `StartIngestWorker` launches an ingest pipeline: a durable `buffer-consumer` consumer of the ingest queue (created through `mq.ConsumerManager`) reads events, batches them per tenant table — the tenant read off each message's `mq.Topic` — and performs bulk INSERTs to ClickHouse. The pipeline is **insert-only**. The wire format `EventMessage` carries `{table_name, scope, received_timestamp, format, columns, row}` — the row positionally as one `JSONCompactEachRow` line, with `columns` naming its positions (the table's insertable columns — a computed one cannot be named in an `INSERT`); the worker batches per (tenant, table, column list) and writes `INSERT INTO … (cols) FORMAT JSONCompactEachRow`. It accepts any table name (events are addressed by `mq.Topic{Tenant, Table, Scope}` with raw names; `internal/mq` encodes them into subject tokens), then bulk-INSERTs. The embedded NATS server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only publishers that can reach the ingest queue are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (`policy.admin_role`) — see the Query Path section below; the `/v1/ops/*` `RequireAdmin` middleware enforces the check at the API layer, so a no/invalid-token request (resolved to `default_role`, not admin in a production config) never reaches the proxy — or through an operator-authored [pipe that writes](/pipes#pipes-that-write), gated only by its `allowed_roles`. A batch whose tenant has no ClickHouse connection (no longer served, or no pool could be opened for it, such as by the connection ceiling) is never tried: no row of it could pass, so `parkBatch` takes it to the DLQ switch whole, logging once per batch rather than twice per row. Otherwise a bulk-insert failure is first classed by `chconn.Classify`: a ClickHouse that cannot take the insert (unavailable, denied, or no verdict at all) sends the batch back to the MQ for a delayed redelivery (`retryLater` → `mq.Message.NakWithDelay`), under a backoff shared by every table on the same pool (a failure of one table — read-only, too many parts — backs off that table alone), and never to the DLQ — the same when it stops answering mid-isolation. Only when ClickHouse rejects the batch, or refuses a multi-row batch for its size (`chconn.Splittable`: too many partitions for one INSERT, the memory limit), is it re-inserted row by row: rows that succeed are acked, and only the rows ClickHouse rejects again are routed to the DLQ (`sendToDLQ` → `mq.DeadLetterer.DeadLetter`), which parks the as-published `EventMessage` envelope under the topic it arrived on (`dlq.{tenant}.{table}` subjects inside `internal/mq`) with the failure context in `X-DLQ-*` headers when the tenant's `dlq.enabled` is on for the table — see [Ingest Pipeline](/ingest-pipeline) for the worker internals. - **backoff.go** — The retry backoff behind `retryLater`: a small circuit breaker per ClickHouse pool (the target's URL, user and database), and one per pool and table for a failure of one table (`chconn.TableScoped`). A failure opens it for 1 s, doubling to a 30 s cap, each window jittered down to half; while it is open, flushes and arriving rows are handed back without a request, and once it elapses one flush probes. Any answer that is not an outage closes it. - **types.go** — `EventMessage` struct (TableName, Scope — reserved, always empty today, ReceivedTimestamp, Format, Columns, Row; `Format` is `FormatJSONCompactEachRow` and `Row` is one positional line whose slots `Columns` names) and `BufferConsumerName` constant, shared across API handlers and the ingest pipeline. - **compact.go** — `EncodeCompactRow`, the positional row encoder every published row goes through, rendering one record over the table's **insertable** columns in declaration order. Serialization only: it validates nothing and judges no value. diff --git a/docs/src/content/docs/pipes.mdx b/docs/src/content/docs/pipes.mdx index 34b8060dc..a6ce6417d 100644 --- a/docs/src/content/docs/pipes.mdx +++ b/docs/src/content/docs/pipes.mdx @@ -180,7 +180,7 @@ The response is a JSON array of rows. Results flow through the shared in-process ### Pipes that write -A pipe's SQL may be a write: a statement led by a write verb WaveHouse recognizes — `INSERT`, `UPDATE`, `DELETE`, `ALTER` (so `ALTER … DELETE`), `CREATE`, `DROP`, `TRUNCATE`, `RENAME`, `EXCHANGE`, `REPLACE`, `OPTIMIZE`, `ATTACH`, `DETACH`, `GRANT`, `REVOKE`, `KILL`, `SET`, `USE` or `SYSTEM` — directly or after a `WITH` list, as in `WITH … INSERT`. Such a pipe runs on **every** call: it never reads or fills the cache and is never coalesced with an identical call in flight, so ten identical calls are ten writes. The response is `[]` with `X-Cache: BYPASS` and `Cache-Control: no-store`, so an HTTP cache in front of a `GET` does not answer a repeat either. WaveHouse classifies the statement from its leading keyword (after any `WITH` list), with the same classifier that sends it to ClickHouse as a write, so no pipe property marks it; a statement led by any other keyword runs as a read — cached and coalesced like any read, so a repeat within the TTL does not run: don't put a write led by another verb (`BACKUP`, `RESTORE`, `UNDROP`) in a pipe ([#666](https://github.com/Wave-RF/WaveHouse/issues/666)). +A pipe's SQL may be a write: a statement led by a write verb WaveHouse recognizes — `INSERT`, `UPDATE`, `DELETE`, `ALTER` (so `ALTER … DELETE`), `CREATE`, `DROP`, `TRUNCATE`, `RENAME`, `EXCHANGE`, `REPLACE`, `OPTIMIZE`, `ATTACH`, `DETACH`, `GRANT`, `REVOKE`, `KILL`, `SET`, `USE` or `SYSTEM` — directly, or `INSERT INTO` after a `WITH` list (`WITH … INSERT INTO …`, the only write ClickHouse accepts there). An `EXECUTE AS ` prefix is looked through: the statement after it is the one classified. Such a pipe runs on **every** call: it never reads or fills the cache and is never coalesced with an identical call in flight, so ten identical calls are ten writes. The response is `[]` with `X-Cache: BYPASS` and `Cache-Control: no-store`, so an HTTP cache in front of a `GET` does not answer a repeat either. WaveHouse classifies the statement by its leading keyword — a `WITH`-led one by whether it holds `INSERT INTO` outside parentheses — with the same classifier that sends it to ClickHouse as a write, so no pipe property marks it. A statement led by any other keyword runs as a read: one that returns rows (`BACKUP`, `RESTORE`) is cached and coalesced, so a repeat within the TTL does not run; one that returns none (`UNDROP`, `MOVE`) fails the call after it has run, which the SDK may retry — don't put a write led by another verb in a pipe ([#666](https://github.com/Wave-RF/WaveHouse/issues/666)). A failed write is not retried automatically, because it may have run. It answers with the status and `code` a failed read would ([ClickHouse errors on the query paths](/api#clickhouse-errors-on-the-query-paths)), but always with `retryable: false` and no `Retry-After`, `503 clickhouse.unavailable` included: once the statement is on its way to ClickHouse, WaveHouse cannot tell whether it ran. The [SDK](/sdk/pipes) does not retry such an answer, so check whether the write landed before you send it again. A call refused before anything is sent — the tenant on no ClickHouse pool, `503` with `Retry-After: 30` — cannot have run, and the SDK retries it. The SDK also retries when WaveHouse's own answer never reaches it — a dropped connection, or a `502`/`503`/`504` from a proxy in front of WaveHouse that gave up waiting — so a write can still run twice that way; give a client that runs write pipes [`options.maxRetries`](/sdk#clientconfigdb) `0` if that matters. diff --git a/internal/api/clickhouse_exec.go b/internal/api/clickhouse_exec.go index 83202628d..51d41e583 100644 --- a/internal/api/clickhouse_exec.go +++ b/internal/api/clickhouse_exec.go @@ -120,11 +120,19 @@ var mutationVerbs = map[string]struct{}{ // them, then the first bareword is matched whole, case-insensitively, against // mutationVerbs. After a WITH list ClickHouse parses only SELECT, a FROM-first // SELECT or INSERT INTO, so a WITH-led statement is a write exactly when it -// holds INSERT INTO at the top level (hasTopLevelInsertInto). A write -// classified as a read goes through Query, which runs it and then fails the -// call, so a client that retries the error writes again. +// holds INSERT INTO at the top level (hasTopLevelInsertInto). An +// `EXECUTE AS ` prefix is looked through to the statement it runs. A +// write classified as a read goes through Query, which runs it and then fails +// the call, so a client that retries the error writes again. func IsMutation(sql string) bool { s := stripLeadingSQLComments(sql) + if rest, ok := skipExecuteAs(s); ok { + // Bare, it switches the session's user and returns no result set. + if rest == "" || rest[0] == ';' { + return true + } + s = rest + } end := skipWord(s, 0) if end == 0 { return false @@ -137,6 +145,45 @@ func IsMutation(sql string) bool { return hasTopLevelInsertInto(s[end:]) } +// skipExecuteAs returns what follows an `EXECUTE AS [@]` prefix +// leading s, past whitespace and comments, and true; or s and false if no such +// prefix leads it. +func skipExecuteAs(s string) (string, bool) { + i := skipWord(s, 0) + if !strings.EqualFold(s[:i], "EXECUTE") { + return s, false + } + i = skipSpaceAndComments(s, i) + j := skipWord(s, i) + if !strings.EqualFold(s[i:j], "AS") { + return s, false + } + i = skipSpaceAndComments(s, skipName(s, skipSpaceAndComments(s, j))) + if i < len(s) && s[i] == '@' { + i = skipSpaceAndComments(s, skipName(s, skipSpaceAndComments(s, i+1))) + } + return s[i:], true +} + +// skipName returns the index just past the user or host name at s[i]: a +// bareword, a quoted identifier or string literal, or a heredoc. +func skipName(s string, i int) int { + if i >= len(s) { + return i + } + switch s[i] { + case '\'', '"', '`': + return skipQuoted(s, i) + case 0xE2: + return skipCurlyQuoted(s, i) + case '$': + if j := skipHeredoc(s, i); j > i { + return j + } + } + return skipWord(s, i) +} + // hasTopLevelInsertInto reports whether s holds INSERT INTO outside // parentheses, stepping over string literals and quoted identifiers // (skipQuoted, skipCurlyQuoted), heredocs (skipHeredoc) and comments @@ -144,7 +191,9 @@ func IsMutation(sql string) bool { // names and aliases may be spelled like any keyword (`WITH 1 AS select`, // `WITH desc AS (…)`, `WITH set -> 1 AS f`, `WITH t.from AS y`), but only the // INSERT statement puts INTO after an insert. The exception, a read's -// `… insert INTO OUTFILE 'f'`, is refused by the server either way. +// `… AS insert INTO OUTFILE 'f'`, is classified as a write and answers `[]` +// uncached: harmless, and contrived. OUTFILE cannot tell the two apart, as +// `INSERT INTO outfile …` names a table. func hasTopLevelInsertInto(s string) bool { depth := 0 i := 0 @@ -177,9 +226,16 @@ func hasTopLevelInsertInto(s string) bool { } else { i = skipWord(s, i+1) } + case c == '.' && i+1 < len(s) && isDigit(s[i+1]): + // A number led by `.` ends with its digits, so a word glued to it + // is a word of its own: `.5INSERT` is `.5` then INSERT. After a + // name ClickHouse reads the `.` as a qualifier (`t.5insert`), which + // can only make a statement it rejects, or a read's INTO OUTFILE, + // look like a write. + i = skipDotNumber(s, i) case isWordByte(c): // A word led by a digit or `_` is read whole, so its tail is - // never taken for a keyword (`_insert`). + // never taken for a keyword (`_insert`, `5insert`). start := i i = skipWord(s, i) if depth == 0 && strings.EqualFold(s[start:i], "INSERT") { @@ -207,7 +263,34 @@ func skipWord(s string, i int) int { } func isWordByte(c byte) bool { - return (c >= 'A' && c <= 'Z') || (c >= 'a' && c <= 'z') || (c >= '0' && c <= '9') || c == '_' + return (c >= 'A' && c <= 'Z') || (c >= 'a' && c <= 'z') || isDigit(c) || c == '_' +} + +func isDigit(c byte) bool { return c >= '0' && c <= '9' } + +// skipDotNumber returns the index just past the number led by the `.` at s[i], +// read as ClickHouse's lexer reads one: digits, then an optional exponent (`e` +// or `E`, an optional sign, any digits), `_` allowed between two digits. +// Unlike a number led by a digit, it ends before any letters that follow it. +func skipDotNumber(s string, i int) int { + i = skipDigits(s, i+1) + if i < len(s) && (s[i] == 'e' || s[i] == 'E') { + i++ + if i < len(s) && (s[i] == '+' || s[i] == '-') { + i++ + } + i = skipDigits(s, i) + } + return i +} + +// skipDigits returns the index just past the run of digits at s[i], in which +// each `_` stands between two digits. +func skipDigits(s string, i int) int { + for i < len(s) && (isDigit(s[i]) || s[i] == '_' && i > 0 && isDigit(s[i-1]) && i+1 < len(s) && isDigit(s[i+1])) { + i++ + } + return i } // skipHeredoc returns the index just past the heredoc opening at s[i] — diff --git a/internal/api/clickhouse_exec_test.go b/internal/api/clickhouse_exec_test.go index c0d8d9b05..1ead574f3 100644 --- a/internal/api/clickhouse_exec_test.go +++ b/internal/api/clickhouse_exec_test.go @@ -87,6 +87,7 @@ func TestExecuteCHQuery_MutationRoutesToExec(t *testing.T) { "ALTER TABLE clicks ADD COLUMN c String", "INSERT INTO clicks VALUES (1)", " -- audit log\n UPDATE clicks SET v = 1 WHERE id = 2", + "EXECUTE AS writer INSERT INTO clicks VALUES (1)", } { t.Run(sql, func(t *testing.T) { t.Parallel() @@ -102,12 +103,17 @@ func TestExecuteCHQuery_MutationRoutesToExec(t *testing.T) { func TestExecuteCHQuery_SelectRoutesToQuery(t *testing.T) { t.Parallel() - conn := &stubConn{} - rows, err := executeCHQuery(context.Background(), conn, "SELECT 1", nil) - require.NoError(t, err) - assert.Zero(t, conn.execCount, "Exec must not be used for SELECT") - assert.Equal(t, 1, conn.queryCount, "Query must be used for SELECT") - assert.Equal(t, []map[string]any{}, rows, "zero-row SELECT must marshal to [] not null") + for _, sql := range []string{"SELECT 1", "EXECUTE AS reader SELECT 1"} { + t.Run(sql, func(t *testing.T) { + t.Parallel() + conn := &stubConn{} + rows, err := executeCHQuery(context.Background(), conn, sql, nil) + require.NoError(t, err) + assert.Zero(t, conn.execCount, "Exec must not be used for SELECT") + assert.Equal(t, 1, conn.queryCount, "Query must be used for SELECT") + assert.Equal(t, []map[string]any{}, rows, "zero-row SELECT must marshal to [] not null") + }) + } } // TestExecuteCHQuery_TransformsClickHouseTypes pins transformRow's contract diff --git a/internal/testutil/mutationtest/cases.go b/internal/testutil/mutationtest/cases.go index a89c1df21..0a94be19d 100644 --- a/internal/testutil/mutationtest/cases.go +++ b/internal/testutil/mutationtest/cases.go @@ -169,6 +169,39 @@ var Cases = []Case{ {"with bare element insert (read)", "WITH insert SELECT 1", false, false}, {"with function insert (read)", "WITH insert(1) AS y SELECT y", false, false}, + // A number led by `.` ends with its digits and exponent, unlike one led + // by a digit, so a word glued to it is a word of its own. + {"with dot-led number glued to insert", "WITH 1 AS a, .5INSERT INTO t SELECT a", true, false}, + {"with signed dot-led exponent glued to insert", "WITH -.5e-3INSERT INTO t SELECT 1", true, false}, + {"with dot-led number and digit separator glued to insert", "WITH 1 AS a, .5_0INSERT INTO t SELECT a", true, false}, + {"with dot-led exponent in a sum glued to insert", "WITH 1+.5e3INSERT INTO t SELECT 1", true, false}, + {"with Unicode minus and dot-led number glued to insert", "WITH −.5INSERT INTO t SELECT 1", true, false}, + + // EXECUTE AS runs the statement after the user as that user, so that + // statement is the one classified; bare, it switches the session's user + // and returns no result set. + {"execute as insert", "EXECUTE AS default INSERT INTO t SELECT 1", true, false}, + {"execute as with insert", "EXECUTE AS default WITH 1 AS a INSERT INTO t SELECT a", true, false}, + {"execute as select", "EXECUTE AS default SELECT 1", false, false}, + {"execute as with select", "EXECUTE AS default WITH 1 AS a SELECT a", false, false}, + {"execute as drop", "EXECUTE AS u1 DROP TABLE t", true, false}, + {"execute as show", "EXECUTE AS u1 SHOW TABLES", false, false}, + {"execute as backticked user then insert", "EXECUTE AS `u 1` INSERT INTO t SELECT 1", true, false}, + {"execute as double-quoted user then select", `EXECUTE AS "u1" SELECT 1`, false, false}, + {"execute as string user glued to insert", "EXECUTE AS 'u1'INSERT INTO t SELECT 1", true, false}, + {"execute as heredoc user then insert", "EXECUTE AS $$u1$$ INSERT INTO t SELECT 1", true, false}, + {"execute as curly-quoted user then insert", "EXECUTE AS “u1” INSERT INTO t SELECT 1", true, false}, + {"execute as user at host then insert", "EXECUTE AS u1@'localhost' INSERT INTO t SELECT 1", true, false}, + {"execute as user at host then select", "EXECUTE AS u1 @ `h` SELECT 1", false, false}, + {"execute as lower with comments then insert", "execute /* c */ as u1 -- c\ninsert into t select 1", true, false}, + {"execute as user named insert then select", "EXECUTE AS insert SELECT 1", false, false}, + {"execute as user named select then insert", "EXECUTE AS select INSERT INTO t SELECT 1", true, false}, + {"execute as bare", "EXECUTE AS u1", true, false}, + {"execute as bare with semicolon", "EXECUTE AS u1;", true, false}, + {"execute as user glued to insert", "EXECUTE AS u1INSERT INTO t SELECT 1", false, true}, + {"execute as nested", "EXECUTE AS u1 EXECUTE AS u2 INSERT INTO t SELECT 1", false, true}, + {"execute without as", "EXECUTE u1 INSERT INTO t SELECT 1", false, true}, + {"empty", "", false, true}, {"comment only", "-- just a comment", false, true}, {"unclosed block comment", "/* never closed", false, true}, diff --git a/tests/integration/ismutation_test.go b/tests/integration/ismutation_test.go index 8e36cb297..1458a9337 100644 --- a/tests/integration/ismutation_test.go +++ b/tests/integration/ismutation_test.go @@ -23,9 +23,12 @@ import ( // codeSyntaxError is ClickHouse's SYNTAX_ERROR. const codeSyntaxError = 62 -// astMutation is api.IsMutation's answer for each statement kind EXPLAIN AST -// names at the root of the tree. +// astMutation is api.IsMutation's answer for each statement kind astRoot +// names. var astMutation = map[string]bool{ + // A bare EXECUTE AS: it switches the session's user and returns no + // result set. + "ExecuteAsQuery": true, "SelectWithUnionQuery": false, "ShowTables": false, "DescribeQuery": false, @@ -49,29 +52,48 @@ var astMutation = map[string]bool{ "UseQuery": true, } -// astRoot is the root node ClickHouse's parser gives sql, or the error it -// rejects sql with. Parsing only: nothing runs, and no table need exist. +// astRoot is the statement kind ClickHouse's parser gives sql — the root node +// of its tree, or for an EXECUTE AS that leads a statement, that statement's +// — or the error it rejects sql with. Parsing only: nothing runs, and no table +// need exist. func astRoot(ctx context.Context, conn driver.Conn, sql string) (string, error) { rows, err := conn.Query(ctx, "EXPLAIN AST "+sql) if err != nil { return "", err } defer func() { _ = rows.Close() }() - if !rows.Next() { - if err := rows.Err(); err != nil { + var lines []string + for rows.Next() { + var line string + if err := rows.Scan(&line); err != nil { return "", err } - return "", errors.New("EXPLAIN AST returned no rows") + lines = append(lines, line) } - var line string - if err := rows.Scan(&line); err != nil { + if err := rows.Err(); err != nil { return "", err } - fields := strings.Fields(line) - if len(fields) == 0 { - return "", fmt.Errorf("EXPLAIN AST returned %q", line) + if len(lines) == 0 { + return "", errors.New("EXPLAIN AST returned no rows") + } + root := strings.Fields(lines[0]) + if len(root) == 0 { + return "", fmt.Errorf("EXPLAIN AST returned %q", lines[0]) + } + if root[0] != "ExecuteAsQuery" { + return root[0], nil + } + // Its children, indented one space: the user, then any statement. + var children []string + for _, line := range lines[1:] { + if strings.HasPrefix(line, " ") && !strings.HasPrefix(line, " ") { + children = append(children, strings.Fields(line)[0]) + } + } + if len(children) < 2 { + return root[0], nil } - return fields[0], nil + return children[1], nil } func isSyntaxError(err error) bool { From d0af402294626e666249b7de38486960230f995e Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 13:10:06 -0400 Subject: [PATCH 69/79] docs(api): name write pipes as a mutation path, not a second route The Insert-only note said every non-insert mutation must go through POST /v1/ops/query and then named write pipes as "the one other route". It now names both paths in the same sentence, as architecture.md and AGENTS.md do. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- docs/src/content/docs/api.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index 8b82643e9..214d60538 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -235,7 +235,7 @@ The inbound request body is capped at 16 MiB; a body over the cap is rejected wi The `{table}` URL query must match a table that exists in ClickHouse. WaveHouse discovers table schemas on startup and refreshes them periodically. :::note[Insert-only] -The ingest pipeline accepts only inserts. All other mutations — `DELETE`, `UPDATE`, `TRUNCATE`, `DROP`, `ALTER`, `REPLACE`, etc. — must be issued through [`POST /v1/ops/query`](#post-v1opsquery--query-clickhouse), which is restricted to the admin role (`admin_role`, the same gate as the rest of `/v1/ops/*`). The one other route is a [pipe that writes](/pipes#pipes-that-write): an operator authors its statement in `pipes.json`, and the roles in its `allowed_roles` run it with parameter values only. +The ingest pipeline accepts only inserts. All other mutations — `DELETE`, `UPDATE`, `TRUNCATE`, `DROP`, `ALTER`, `REPLACE`, etc. — must be issued through [`POST /v1/ops/query`](#post-v1opsquery--query-clickhouse), which is restricted to the admin role (`admin_role`, the same gate as the rest of `/v1/ops/*`), or through an operator-authored [pipe that writes](/pipes#pipes-that-write): an operator authors its statement in `pipes.json`, and the roles in its `allowed_roles` run it with parameter values only. The policy engine authorizes mutations by inspecting the columns being written. That works for inserts but not for predicate-driven mutations like `DELETE … WHERE` — there's no way to prove the predicate matches only rows the caller is allowed to touch. Routing those statements through the admin-gated raw-SQL surface, or through a pipe whose predicate the operator wrote, keeps the policy contract honest. ::: From ecd21c0890b5e84e5e830547a32e1e7a6c98233f Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 13:10:19 -0400 Subject: [PATCH 70/79] docs(cache): match deployment.md to #626's probe-fix (0eea1cbc) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Verified against the merged internal/cache/redis.go: - The failover bullet said the probe gets DialTimeout+Timeout; it's now 2×DialTimeout+Timeout for the first write (rueidis bounds the dial and the handshake by DialTimeout in turn), and a write slower than Timeout is repeated under Timeout alone — only the repeat decides whether the breaker closes, so a server that merely answers slowly stays bypassed instead of flapping open and shut every cycle. - Confirmed the probe is unaffected by, and doesn't affect, the boot/ Close 1s-cap arithmetic in configuration.mdx's dial_timeout row and cache_redis.go's maxRedisTimeout comment: probe() runs in an untracked goroutine (no r.wg.Add before `go r.probe(...)`), so Close's r.wg.Wait() never waits on it. No doc change needed there. - "Run it without persistence" didn't say why: stock Redis and Valkey persist by default (periodic RDB save points), so an ordinary crash- restart reloads the last save on its own — the same rollback as restoring a snapshot by hand. Reworded to match #626's CHANGELOG/ architecture.md wording (grepped for other copies; found none). Co-Authored-By: Claude Sonnet 5 Claude-Session: https://claude.ai/code/session_017aS7rLrH1RKkUMem7X4ckd --- docs/src/content/docs/deployment.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index b3fdcf801..0ba84aa5c 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -432,12 +432,12 @@ The query-result cache is the layer that can be shared today. With the default ` **What another instance can see.** Ingest is already asynchronous: `/v1/ingest` answers before the batch is inserted. Once the inserting instance's worker has written the batch to ClickHouse, it replaces the table's version token in Redis, and from then on a lookup on any instance misses and reads the new rows. The cache adds no delay of its own beyond that single write. The exceptions: - **The server is unreachable from the inserting instance.** The invalidation is kept and retried until it lands (`wavehouse_cache_invalidations_pending` counts what is owed). Meanwhile other instances that can still reach the server keep serving the older results, for as long as the outage lasts and at most until each entry's TTL. An instance that stops while invalidations are still owed loses them, with the same bound. The same thing happens today when a process stops between an insert and its invalidation. -- **A failover to a replica that had not yet received the latest token writes** can bring back entries filed under the older tokens, bounded by the replication lag at the moment of failover and those entries' TTL. WaveHouse never reads from replicas. Behind a stable address (a managed primary endpoint), an instance still connected to the demoted node has its writes refused, which bypasses its cache; connections are replaced every minute, so it reaches the new primary and delivers the invalidations it owes within about that long. The breaker's own probe gets a longer reconnect budget (`dial_timeout` plus `timeout`) so a slow reconnect still closes it, but every other connection redials under `timeout` alone: size it above how long a reconnect actually takes, or operations that land on one of those keep failing after the probe has already succeeded. +- **A failover to a replica that had not yet received the latest token writes** can bring back entries filed under the older tokens, bounded by the replication lag at the moment of failover and those entries' TTL. WaveHouse never reads from replicas. Behind a stable address (a managed primary endpoint), an instance still connected to the demoted node has its writes refused, which bypasses its cache; connections are replaced every minute, so it reaches the new primary and delivers the invalidations it owes within about that long. The breaker's own probe write gets a longer budget for a reconnect — twice `dial_timeout` (the client bounds the dial and the handshake by it in turn) plus `timeout` for the write itself — but only closes the breaker when a write actually lands within `timeout`: one slower than that is repeated under `timeout` alone, and the repeat decides, so a server that merely answers slowly stays bypassed instead of flapping open and shut. Every other connection redials under `timeout` alone: size it above how long a reconnect actually takes, or operations that land on one of those keep failing after the probe has already succeeded. - **The server is full and `maxmemory-policy` is `noeviction`.** It refuses the token writes. The inserting instance keeps its invalidations and retries them, bypassing its cache meanwhile, but every other instance serves the results from before the insert until one lands, up to their TTL. - **A pipe that writes** (an `INSERT` in `pipes.json`) has its result cached like a read, so a repeated identical call is answered from the cache and the write does not run again ([#386](https://github.com/Wave-RF/WaveHouse/issues/386)). With a shared cache that holds on every instance, until the entry's TTL. - **Admin writes through `POST /v1/ops/query`** do not invalidate the cache ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). With a shared cache, the stale results they leave are served by every instance, not only one. -**Sizing the server.** Every key WaveHouse writes has a TTL, and a version token lost to eviction, expiry or `FLUSHALL` can only cause misses, never bring back an entry it had invalidated. So set `maxmemory` and let the server evict: `maxmemory-policy allkeys-lru` (or `allkeys-lfu`, `volatile-lru`, `volatile-lfu`). Under `noeviction`, a full server refuses the writes: each refusal (a fill's is counted by `wavehouse_cache_set_failures_total{reason="oom"}`) bypasses the cache of the instance that got it, and invalidations are kept and retried, so the pre-insert results above stay served by the others: avoid `noeviction`. A stored result is capped at `cache.redis.max_value_bytes` (1 MiB compressed). A tenant's version tokens share one hash tag, so each lookup reads them in one `MGET` in cluster mode as well. The results themselves carry no hash tag and spread across shards. **Run it without persistence** (`save ""` and `appendonly no`): a restart without persistence can only cause misses, the same as any other token loss. With persistence on, a restart is not that — it reloads whatever snapshot or AOF it last wrote, tokens and values it had already invalidated included, so a fresh instance can serve the pre-write rows filed under them as hits until their TTL (up to 1 h) expires. Treat restoring a snapshot as a rollback, not a resume. +**Sizing the server.** Every key WaveHouse writes has a TTL, and a version token lost to eviction, expiry or `FLUSHALL` can only cause misses, never bring back an entry it had invalidated. So set `maxmemory` and let the server evict: `maxmemory-policy allkeys-lru` (or `allkeys-lfu`, `volatile-lru`, `volatile-lfu`). Under `noeviction`, a full server refuses the writes: each refusal (a fill's is counted by `wavehouse_cache_set_failures_total{reason="oom"}`) bypasses the cache of the instance that got it, and invalidations are kept and retried, so the pre-insert results above stay served by the others: avoid `noeviction`. A stored result is capped at `cache.redis.max_value_bytes` (1 MiB compressed). A tenant's version tokens share one hash tag, so each lookup reads them in one `MGET` in cluster mode as well. The results themselves carry no hash tag and spread across shards. **Run it without persistence** (`save ""` and `appendonly no`): stock Redis and Valkey persist by default (periodic RDB save points), so a crash that is followed by a restart reloads the last save on its own — the same rollback as restoring a snapshot by hand, not a loss. Without persistence, a restart can only cause misses, the same as any other token loss. With it on (the default), a restart reloads whatever snapshot or AOF it last wrote, tokens and values it had already invalidated included, so a fresh instance can serve the pre-write rows filed under them as hits until their TTL (up to 1 h) expires. Treat restoring a snapshot, or a crash-restart on a server that still has its defaults, as a rollback, not a resume. **The server is inside the trust boundary.** A cached result is served after the access policy has filtered it, so whoever can write to the server can change what any caller reads. Keep it on a private network, require a password or ACL user (`WH_CACHE_REDIS_PASSWORD`), use TLS across links you do not trust, and share it only with deployments you trust as much as this one. From b3068ab6563fe952ae10ee9acfed732d90cd2470 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 14:30:20 -0400 Subject: [PATCH 71/79] test(cache): pin the failover test to one connection rueidis keeps up to four connections to a standalone server by GOMAXPROCS and dials each on first use, so an operation landing on one first used after the failover reached the new primary with no connection replaced: without ConnLifetime the test still passed unless GOMAXPROCS was 1. A test-only option gives the client one connection, dialed before the failover, so the test fails without the replacement at any GOMAXPROCS. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01FyrXjhR7iDg33paioLQHFq --- internal/cache/export_test.go | 5 +++++ internal/cache/redis.go | 4 ++++ internal/cache/redis_integration_test.go | 1 + 3 files changed, 10 insertions(+) diff --git a/internal/cache/export_test.go b/internal/cache/export_test.go index 70b8544d2..73a791178 100644 --- a/internal/cache/export_test.go +++ b/internal/cache/export_test.go @@ -24,6 +24,11 @@ func Bypassed(r *RedisCache) bool { return r.bypassed() } // replaced, for a failover test. func SetConnLifetime(c *RedisConfig, d time.Duration) { c.connLifetime = d } +// SetOnePipe gives c's client one connection per node. rueidis otherwise +// keeps up to four by GOMAXPROCS, dialing each on first use, so one first +// used after a failover reaches the new primary without being replaced. +func SetOnePipe(c *RedisConfig) { c.onePipe = true } + // KeyPrefix is the prefix every key r writes leads with. func KeyPrefix(r *RedisCache) string { return r.cfg.KeyPrefix } diff --git a/internal/cache/redis.go b/internal/cache/redis.go index 92719d80e..6e567e899 100644 --- a/internal/cache/redis.go +++ b/internal/cache/redis.go @@ -80,6 +80,7 @@ type RedisConfig struct { BreakerOpenFor time.Duration // how long it stays open before a probe connLifetime time.Duration // defaultConnLifetime; tests shorten it + onePipe bool // one connection per node, whatever GOMAXPROCS; tests only } func (c RedisConfig) withDefaults() (RedisConfig, error) { @@ -174,6 +175,9 @@ func (c RedisConfig) clientOption() rueidis.ClientOption { // bypass reaches the new primary, and delivers its owed bumps, within // about this long. opt.ConnLifetime = c.connLifetime + if c.onePipe { + opt.PipelineMultiplex = -1 + } if c.Mode == RedisSentinel { opt.Sentinel = rueidis.SentinelOption{MasterSet: c.SentinelMaster, TLSConfig: c.TLS, Dialer: opt.Dialer} } diff --git a/internal/cache/redis_integration_test.go b/internal/cache/redis_integration_test.go index 13ff66557..f19ea8059 100644 --- a/internal/cache/redis_integration_test.go +++ b/internal/cache/redis_integration_test.go @@ -727,6 +727,7 @@ func TestRedis_FailoverBehindAStableAddress(t *testing.T) { a := open(t, stable, prefix, func(c *cache.RedisConfig) { c.BreakerThreshold, c.BreakerOpenFor = 1000, 200*time.Millisecond cache.SetConnLifetime(c, time.Second) + cache.SetOnePipe(c) // every operation uses the connection dialed before the failover }) deps := []cache.Namespace{{Tenant: "acme", Table: "events"}} _, snap, err := a.Lookup(ctx, "acme", "q", deps) From 822aecb17bac70d56970532e165c1c5da36790b4 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 14:30:34 -0400 Subject: [PATCH 72/79] fix(cache): log the breaker opening once, at WARN A reply refusing writes and a run of unanswered operations opened the breaker silently. Each opening, a failed probe's included, now logs one WARN with the reply or the error; operations failing while it is open log nothing. Also rewords a comment that named an internal milestone. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01FyrXjhR7iDg33paioLQHFq --- CHANGELOG.md | 2 +- internal/cache/breaker.go | 8 +++++-- internal/cache/breaker_test.go | 7 +++--- internal/cache/redis.go | 38 +++++++++++++++++++++----------- internal/cache/redis_test.go | 40 ++++++++++++++++++++++++++++++++++ 5 files changed, 76 insertions(+), 19 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index a75b24028..e14bae9c2 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A Redis-compatible shared cache backend, built but not yet selectable** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/cache.go`, `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`, `CONTRIBUTING.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart without persistence can only cause misses — `maxmemory-policy allkeys-lru` is safe. Restoring an RDB or AOF snapshot, or a backup, is a rollback instead (a restart after a crash that reloads the server's last save included, which stock Redis and Valkey make by default): the old tokens return with their values, so invalidations made since are undone until those entries' TTL. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like) or the credentials (`WRONGPASS` or `NOAUTH` after a password rotation, logged at `ERROR`), open a circuit breaker that skips the server until a probe write succeeds within the per-operation timeout (connections are replaced every minute, so after a failover behind a stable address the process reaches the new primary, and delivers the bumps it owes, within about a minute; the probe also allows up to twice the dial timeout for a reconnect, but the client's other connections must reconnect within the per-operation timeout, so that timeout should still exceed a reconnect), and deferred invalidations are retried until they land — the first at once, and at the probe's cadence while the breaker is open — collapsing to one tenant-wide bump per tenant past 100,000 keys; until one lands, the process that owes it bypasses the lookups it would orphan. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. Nothing selects it yet: `cache.backend` accepts only `local` until the wiring (E4) adds `redis` and the `cache.redis.*` settings, so every deployment still runs `LocalCache`. Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, snapshot-rollback, compression, stored-size, server-stops-answering, refused-writes (a demoted primary, a full `noeviction` server), rotated-credentials, slow-reconnect (over TLS), slow-server, failover-behind-a-stable-address and owed-invalidation cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. +- **A Redis-compatible shared cache backend, built but not yet selectable** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/cache.go`, `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`, `CONTRIBUTING.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart without persistence can only cause misses — `maxmemory-policy allkeys-lru` is safe. Restoring an RDB or AOF snapshot, or a backup, is a rollback instead (a restart after a crash that reloads the server's last save included, which stock Redis and Valkey make by default): the old tokens return with their values, so invalidations made since are undone until those entries' TTL. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like) or the credentials (`WRONGPASS` or `NOAUTH` after a password rotation), open a circuit breaker — logged once per opening, at `ERROR` for the credentials and `WARN` otherwise — that skips the server until a probe write succeeds within the per-operation timeout (connections are replaced every minute, so after a failover behind a stable address the process reaches the new primary, and delivers the bumps it owes, within about a minute; the probe also allows up to twice the dial timeout for a reconnect, but the client's other connections must reconnect within the per-operation timeout, so that timeout should still exceed a reconnect), and deferred invalidations are retried until they land — the first at once, and at the probe's cadence while the breaker is open — collapsing to one tenant-wide bump per tenant past 100,000 keys; until one lands, the process that owes it bypasses the lookups it would orphan. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. Nothing selects it yet: `cache.backend` accepts only `local` until the wiring (E4) adds `redis` and the `cache.redis.*` settings, so every deployment still runs `LocalCache`. Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, snapshot-rollback, compression, stored-size, server-stops-answering, refused-writes (a demoted primary, a full `noeviction` server), rotated-credentials, slow-reconnect (over TLS), slow-server, failover-behind-a-stable-address and owed-invalidation cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. - **Leases for work that must run in one process at a time, and the sweeper runs under one** (`internal/coord/` (new: `coord.go`, `local.go`, `elect.go`, `coordtest/`, + tests), `internal/app/{app,wire}.go` (+ tests), `internal/config/backends.go`, `internal/ingest/sweeper.go`, `.github/labeler.yml`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development,ingest-pipeline}.md`, `docs/src/content/docs/configuration.mdx`, `config.yaml`, `AGENTS.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `coord.Coordinator` hands out named leases (`TryAcquire` → a `Term` with a strictly increasing fencing `Token`, a `Done` channel and `Resign`; `ErrHeld` while another holder's term is live), and `coord.RunElected` runs a loop only while its process holds the lease, resigning when the loop returns and campaigning again every 2s. `coord.Local` is the in-process implementation, and `coordtest.Conformance` is the suite every implementation runs — the NATS KV backend that lets several replicas share one queue comes next. The sweeper now runs through `RunElected` under the `sweeper` lease; with the in-process coordinator the one process always holds it, so nothing changes for a single-process deployment beyond one `coord: elected` log line at startup. `coord.backend` now selects the coordinator (`local`, the only value), so a `config.Config` built without `config.Load` must name it as well as the other three layers' backends. diff --git a/internal/cache/breaker.go b/internal/cache/breaker.go index 6bcb89d4c..f3e065a53 100644 --- a/internal/cache/breaker.go +++ b/internal/cache/breaker.go @@ -54,14 +54,18 @@ func (b *breaker) success() { } // failure records a call the server did not answer in time, and opens the -// breaker at the threshold — or at once, for a failed probe. -func (b *breaker) failure() { +// breaker at the threshold — or at once, for a failed probe. It reports +// whether that opened a closed or probing breaker, as trip does. +func (b *breaker) failure() bool { b.mu.Lock() defer b.mu.Unlock() b.failures++ if b.probing || b.failures >= b.threshold { + opened := !b.open || b.probing b.open, b.openedAt, b.probing = true, b.now(), false + return opened } + return false } // trip opens the breaker at once, for a reply that says the server cannot diff --git a/internal/cache/breaker_test.go b/internal/cache/breaker_test.go index ace2f13f1..273a871c0 100644 --- a/internal/cache/breaker_test.go +++ b/internal/cache/breaker_test.go @@ -26,10 +26,11 @@ func TestBreaker(t *testing.T) { b.failure() b.success() // a success resets the run b.failure() - b.failure() + assert.False(t, b.failure()) assert.False(t, b.isOpen(), "two in a row is below the threshold") - b.failure() + assert.True(t, b.failure(), "it opened") assert.True(t, b.isOpen()) + assert.False(t, b.failure(), "already open") ok, probe = allow() assert.False(t, ok, "open: skip the server") @@ -43,7 +44,7 @@ func TestBreaker(t *testing.T) { assert.False(t, ok) assert.False(t, probe, "one probe at a time") - b.failure() // the probe failed: open for another period + assert.True(t, b.failure(), "the probe failed: open for another period") assert.True(t, b.isOpen()) clock.t = clock.t.Add(4 * time.Second) _, probe = allow() diff --git a/internal/cache/redis.go b/internal/cache/redis.go index 6e567e899..1889ee727 100644 --- a/internal/cache/redis.go +++ b/internal/cache/redis.go @@ -160,7 +160,7 @@ func (c RedisConfig) clientOption() rueidis.ClientOption { TLSConfig: c.TLS, Dialer: net.Dialer{Timeout: c.DialTimeout}, ClientName: "wavehouse", - DisableCache: true, // no client-side caching until the near-cache (E5) + DisableCache: true, // no client-side caching until a local near-cache in front of Redis exists ForceSingleClient: c.Mode == RedisStandalone, } // How long a connection waits on a silent server, 10 s unset. A cluster @@ -360,17 +360,7 @@ func (r *RedisCache) record(parent context.Context, err error) { return } if re, ok := rueidis.IsRedisErr(err); ok { - switch msg := re.Error(); { - case rejectsCredentials(msg): - if r.breaker.trip() { - slog.ErrorContext(parent, "cache: redis rejected the credentials; bypassing the cache until they work", - "addrs", r.cfg.Addrs, "error", msg) - } - case refusesWork(msg): - r.breaker.trip() - default: - r.breaker.success() - } + r.recordReply(parent, re.Error()) return } if errors.Is(err, errMalformedReply) { @@ -380,7 +370,29 @@ func (r *RedisCache) record(parent context.Context, err error) { if parent.Err() != nil { return } - r.breaker.failure() + if r.breaker.failure() { + slog.WarnContext(parent, "cache: redis not answering; bypassing the cache", + "addrs", r.cfg.Addrs, "error", err) + } +} + +// recordReply is record for an error reply. The breaker opening is logged +// once, not once per operation in flight. +func (r *RedisCache) recordReply(parent context.Context, msg string) { + switch { + case rejectsCredentials(msg): + if r.breaker.trip() { + slog.ErrorContext(parent, "cache: redis rejected the credentials; bypassing the cache until they work", + "addrs", r.cfg.Addrs, "error", msg) + } + case refusesWork(msg): + if r.breaker.trip() { + slog.WarnContext(parent, "cache: redis refusing writes; bypassing the cache", + "addrs", r.cfg.Addrs, "reply", msg) + } + default: + r.breaker.success() + } } // refusesWork reports whether an error reply says the server takes no diff --git a/internal/cache/redis_test.go b/internal/cache/redis_test.go index deb4f2c56..5f9d5cec4 100644 --- a/internal/cache/redis_test.go +++ b/internal/cache/redis_test.go @@ -4,7 +4,9 @@ import ( "context" "errors" "fmt" + "log/slog" "net" + "strings" "testing" "time" @@ -17,6 +19,7 @@ import ( "go.opentelemetry.io/otel/sdk/metric/metricdata" "github.com/Wave-RF/WaveHouse/internal/tenant" + "github.com/Wave-RF/WaveHouse/internal/testutil/logtest" ) func TestRedisConfig_Validation(t *testing.T) { @@ -160,6 +163,43 @@ func TestRedis_Record(t *testing.T) { assert.True(t, r.breaker.isOpen()) } +// Each opening of the breaker logs one WARN naming why; the operations that +// fail while it is open log nothing. Not parallel: it captures the default +// logger. +func TestRedis_BreakerOpeningLogsOnce(t *testing.T) { + buf := logtest.Capture(t, slog.LevelWarn) + clock := &fakeClock{t: time.Unix(0, 0)} + r := &RedisCache{breaker: newBreaker(2, time.Second, clock.now)} + live := context.Background() + warns := func() int { return strings.Count(buf.String(), `"level":"WARN"`) } + + r.recordReply(live, "READONLY You can't write against a read only replica.") + r.recordReply(live, "READONLY You can't write against a read only replica.") + r.record(live, context.DeadlineExceeded) + r.record(live, context.DeadlineExceeded) + assert.Equal(t, 1, warns(), buf.String()) + assert.Contains(t, buf.String(), `"reply":"READONLY You can't write against a read only replica."`) + + clock.t = clock.t.Add(time.Second) + _, probe := r.breaker.allow() + require.True(t, probe) + r.record(live, nil) + require.False(t, r.breaker.isOpen()) + + r.record(live, context.DeadlineExceeded) + assert.Equal(t, 1, warns(), "under the threshold") + r.record(live, context.DeadlineExceeded) + r.record(live, context.DeadlineExceeded) + assert.Equal(t, 2, warns(), buf.String()) + assert.Contains(t, buf.String(), "not answering") + + clock.t = clock.t.Add(time.Second) + _, probe = r.breaker.allow() + require.True(t, probe) + r.record(live, context.DeadlineExceeded) + assert.Equal(t, 3, warns(), "a failed probe opens it again") +} + func TestRejectsCredentials(t *testing.T) { t.Parallel() assert.True(t, rejectsCredentials("WRONGPASS invalid username-password pair or user is disabled.")) From aad9b1e5b965bd30f2a05618650f2cec1b184aad Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 14:33:21 -0400 Subject: [PATCH 73/79] docs(cache): say every breaker opening logs, a failed probe's included Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01FyrXjhR7iDg33paioLQHFq --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index e14bae9c2..5aa2462e9 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A Redis-compatible shared cache backend, built but not yet selectable** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/cache.go`, `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`, `CONTRIBUTING.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart without persistence can only cause misses — `maxmemory-policy allkeys-lru` is safe. Restoring an RDB or AOF snapshot, or a backup, is a rollback instead (a restart after a crash that reloads the server's last save included, which stock Redis and Valkey make by default): the old tokens return with their values, so invalidations made since are undone until those entries' TTL. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like) or the credentials (`WRONGPASS` or `NOAUTH` after a password rotation), open a circuit breaker — logged once per opening, at `ERROR` for the credentials and `WARN` otherwise — that skips the server until a probe write succeeds within the per-operation timeout (connections are replaced every minute, so after a failover behind a stable address the process reaches the new primary, and delivers the bumps it owes, within about a minute; the probe also allows up to twice the dial timeout for a reconnect, but the client's other connections must reconnect within the per-operation timeout, so that timeout should still exceed a reconnect), and deferred invalidations are retried until they land — the first at once, and at the probe's cadence while the breaker is open — collapsing to one tenant-wide bump per tenant past 100,000 keys; until one lands, the process that owes it bypasses the lookups it would orphan. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. Nothing selects it yet: `cache.backend` accepts only `local` until the wiring (E4) adds `redis` and the `cache.redis.*` settings, so every deployment still runs `LocalCache`. Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, snapshot-rollback, compression, stored-size, server-stops-answering, refused-writes (a demoted primary, a full `noeviction` server), rotated-credentials, slow-reconnect (over TLS), slow-server, failover-behind-a-stable-address and owed-invalidation cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. +- **A Redis-compatible shared cache backend, built but not yet selectable** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/cache.go`, `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`, `CONTRIBUTING.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart without persistence can only cause misses — `maxmemory-policy allkeys-lru` is safe. Restoring an RDB or AOF snapshot, or a backup, is a rollback instead (a restart after a crash that reloads the server's last save included, which stock Redis and Valkey make by default): the old tokens return with their values, so invalidations made since are undone until those entries' TTL. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like) or the credentials (`WRONGPASS` or `NOAUTH` after a password rotation), open a circuit breaker — logged once per opening, at `ERROR` for the credentials and `WARN` otherwise, which a failed probe repeats every 5 s while the server stays down — that skips the server until a probe write succeeds within the per-operation timeout (connections are replaced every minute, so after a failover behind a stable address the process reaches the new primary, and delivers the bumps it owes, within about a minute; the probe also allows up to twice the dial timeout for a reconnect, but the client's other connections must reconnect within the per-operation timeout, so that timeout should still exceed a reconnect), and deferred invalidations are retried until they land — the first at once, and at the probe's cadence while the breaker is open — collapsing to one tenant-wide bump per tenant past 100,000 keys; until one lands, the process that owes it bypasses the lookups it would orphan. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. Nothing selects it yet: `cache.backend` accepts only `local` until the wiring (E4) adds `redis` and the `cache.redis.*` settings, so every deployment still runs `LocalCache`. Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, snapshot-rollback, compression, stored-size, server-stops-answering, refused-writes (a demoted primary, a full `noeviction` server), rotated-credentials, slow-reconnect (over TLS), slow-server, failover-behind-a-stable-address and owed-invalidation cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. - **Leases for work that must run in one process at a time, and the sweeper runs under one** (`internal/coord/` (new: `coord.go`, `local.go`, `elect.go`, `coordtest/`, + tests), `internal/app/{app,wire}.go` (+ tests), `internal/config/backends.go`, `internal/ingest/sweeper.go`, `.github/labeler.yml`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development,ingest-pipeline}.md`, `docs/src/content/docs/configuration.mdx`, `config.yaml`, `AGENTS.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `coord.Coordinator` hands out named leases (`TryAcquire` → a `Term` with a strictly increasing fencing `Token`, a `Done` channel and `Resign`; `ErrHeld` while another holder's term is live), and `coord.RunElected` runs a loop only while its process holds the lease, resigning when the loop returns and campaigning again every 2s. `coord.Local` is the in-process implementation, and `coordtest.Conformance` is the suite every implementation runs — the NATS KV backend that lets several replicas share one queue comes next. The sweeper now runs through `RunElected` under the `sweeper` lease; with the in-process coordinator the one process always holds it, so nothing changes for a single-process deployment beyond one `coord: elected` log line at startup. `coord.backend` now selects the coordinator (`local`, the only value), so a `config.Config` built without `config.Load` must name it as well as the other three layers' backends. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 2347ec727..9a18f8307 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -117,7 +117,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: a namespace key is `..
.
.`, its fields joined with `keyenc.Join` so a dot in a table or scope name is escaped rather than read as a separator, and an entry's key, `|.||…` with the caller's query key escaped whole (its `:` become `%3A`), folds in the tenant's version and each dependency's namespace key and namespace version, so bumping a table (a scopeless write) or one scope — scope is reserved and empty today, so every write is the whole-table bump — orphans every dependent entry without touching the pool. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results, and the flat directory's single tenant simply carries the `0` prefix. `BumpTenant` (behind `InvalidateTenant`) advances the tenant version that leads every namespace key of one tenant, orphaning its every namespace and every cached query in one step (the tenant version is folded into every entry's key, so a pipe result with no dependencies is orphaned too) — a table no bump ever keyed included, which an enumeration of the index would miss — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. - **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET` and `MGET`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot; the table and scope are escaped with `keyenc` after the fixed prefix, so a `:` in a name is never read as the separator. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, the hash over their escaped forms, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token (a string of another length, a list, a hash) is replaced, so a token lost to eviction, expiry, `FLUSHALL` or a restart without persistence is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. Restoring an RDB or AOF snapshot, or a backup, is not a loss but a rollback — a restart after a crash that reloads the server's last save included, which stock Redis and Valkey make by default: the old tokens come back with the values filed under them, so what was invalidated since is served again until its TTL, as after a failover to a replica that missed the bumps. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. While this process owes a bump on any of a lookup's tokens, that lookup is a bypass that files nothing: the bump would orphan whatever it found (other processes, which cannot know, serve those entries until it lands). `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing, each attempt bounded by `DialTimeout` (1 s), and a cluster client's topology read after the handshake by the larger of it and `Timeout`, which is also how long a connection waits on a silent server before it is redialed. Every connection is replaced after a minute (`clientOption`; rueidis retries what was in flight), because one that outlives a failover behind a stable address stays on the demoted node, which answers but refuses writes; the replacement re-resolves the address, so a bypassed process reaches the new primary, and delivers the bumps it owes, within about that long. rueidis dials a replacement lazily, under the context of the operation that lands on it, bounding the dial (TLS included) by `DialTimeout` and then the handshake by `DialTimeout` again, so a reconnect slower than `Timeout` fails that operation. The breaker's probe gives its write twice `DialTimeout` for a reconnect on top of `Timeout`, so such a reconnect still closes an open breaker. The allowance is for a reconnect only: a probe write slower than `Timeout` is repeated under `Timeout`, and the repeat decides, so a server answering slower than `Timeout` stays bypassed. But rueidis spreads commands over several connections (up to four to one server, by `GOMAXPROCS`, and one per cluster node) and the probe reconnects only the one it lands on, so size `Timeout` above a reconnect, or operations that land on the others keep failing. Nothing selects this backend yet: `cache.backend` accepts only `local` until [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s E4 adds `redis`, the `cache.redis.*` settings and the wiring. - **redis_codec.go** — the key schema (`tokenKeys`, `bumpKeys`, `valueKey`, which take a `Namespace`'s raw names and escape them) and the value frame: format, flags, expiry, the token list, the payload, zstd-compressed from `CompressMinBytes` when that is smaller. A stored value is capped at `MaxValueBytes` and a decoded one at eight times that, which refuses a zip bomb planted in a shared server; an unknown format, as a newer process writes during a rolling upgrade, is a miss. -- **breaker.go** — the circuit breaker's state machine: `BreakerThreshold` (5) consecutive failures open it, `trip` opens it at once, and while it is open every operation skips the server; once `BreakerOpenFor` (5 s) has passed, `allow` hands one caller the probe, and only the probe's success closes it. What feeds it is `redis.go`'s. `record` counts a transport failure or a timeout against the server, and trips the breaker on an error reply that means no bump can land: one saying the server takes no writes right now, as `refusesWork` lists them — `READONLY` (a demoted primary), `MASTERDOWN`, `OOM` (full under `maxmemory-policy noeviction`, Redis's default, so give the cache an evicting policy), `NOREPLICAS`, `MISCONF`, `LOADING`, `BUSY`, `CLUSTERDOWN` — or one refusing the credentials, as `rejectsCredentials` lists them — `WRONGPASS`, `NOAUTH`, what a new connection's handshake meets after a password rotation — which it logs at `ERROR`. Any other reply, an error reply about one key or command (`WRONGTYPE`, `NOPERM`, `TRYAGAIN`) or one the backend cannot use included, counts as a success, and a caller that gave up first counts as nothing. The probe (`probe`) is a write, `SET :probe`, so a server that answers but refuses writes stays bypassed. +- **breaker.go** — the circuit breaker's state machine: `BreakerThreshold` (5) consecutive failures open it, `trip` opens it at once, and while it is open every operation skips the server; once `BreakerOpenFor` (5 s) has passed, `allow` hands one caller the probe, and only the probe's success closes it. What feeds it is `redis.go`'s. `record` counts a transport failure or a timeout against the server, and trips the breaker on an error reply that means no bump can land: one saying the server takes no writes right now, as `refusesWork` lists them — `READONLY` (a demoted primary), `MASTERDOWN`, `OOM` (full under `maxmemory-policy noeviction`, Redis's default, so give the cache an evicting policy), `NOREPLICAS`, `MISCONF`, `LOADING`, `BUSY`, `CLUSTERDOWN` — or one refusing the credentials, as `rejectsCredentials` lists them — `WRONGPASS`, `NOAUTH`, what a new connection's handshake meets after a password rotation. Each opening is logged once, at `ERROR` for the credentials and `WARN` otherwise, not once per operation in flight; a failed probe opens it afresh, so a server that stays down logs once every `BreakerOpenFor`. Any other reply, an error reply about one key or command (`WRONGTYPE`, `NOPERM`, `TRYAGAIN`) or one the backend cannot use included, counts as a success, and a caller that gave up first counts as nothing. The probe (`probe`) is a write, `SET :probe`, so a server that answers but refuses writes stays bypassed. - **pending.go** — the invalidations owed (`pendingBumps`): kept per token key, repeats coalescing, and past `PendingMax` (100,000) keys collapsed to one tenant bump per affected tenant; `owesAny` tells `Lookup` which lookups to hold. `redis.go` delivers them: `Invalidate` sends its bumps 1,000 to a round trip and defers those the server does not take, and `drainLoop` retries them through `drain` until they land — the first one owed at once, then with backoff from 100 ms to 10 s, and while the breaker is open at each probe, which the loop starts when due, so a process that makes no lookups (ingest only) recovers as soon as one that does. Landing late is still correct: a fresh token orphans the pre-write entries and any fill made meanwhile. `Close` makes one last attempt at the bumps still owed, past the breaker (an open one is why they are owed); what that attempt cannot deliver is lost, and the entries they would orphan are served until their TTL — the same failure as a worker stopping between an insert and its invalidation. - **metrics.go** — the shared backend's instruments (meter `wavehouse-cache`, `backend="redis"`, no tenant attribute): `wavehouse_cache_lookups_total` by `result` (`hit`, `miss`, `stale` — filed under since-bumped tokens, `bypass` — server skipped, or held by a bump this process owes, `error`), `wavehouse_cache_op_duration_seconds` by `op` (`lookup`, `set`, `invalidate`), `wavehouse_cache_breaker_open` (1 while bypassed, including before the first connection), `wavehouse_cache_invalidations_total` by `result` (`ok` counts every bump that lands, retried ones included, and `deferred` each bump an invalidation could not deliver when made, a repeat of one already owed included; a failed retry is not counted again, so the two overlap), `wavehouse_cache_invalidations_pending` (bumps owed: entries they would orphan may be served stale meanwhile), `wavehouse_cache_value_bytes` (stored size), `wavehouse_cache_oversize_total` and `wavehouse_cache_set_failures_total` by `reason` (`oom`, `timeout`, `other`). - **cachetest** (`internal/testutil/cachetest`) — `Run(t, factory, Options)`, the conformance suite every `Cache` backend runs, each case on a fresh cache from the factory: what a hit, a miss and each kind of bump mean, independent of where entries and versions live. `Options` describes what a backend can do beyond the `Cache` contract, and leaving one unset skips the cases it enables: `MaxValueBytes` opens the oversize case; `NewPair` returns two instances over one shared store, for the cross-instance cases a shared backend must run; `Entries` counts a cache's entries, and without it the zero-snapshot case — no `Lookup` reads the key a zero `Snapshot` would land under, so only a count shows that one stored nothing — is skipped. From 0281ca31efe0327b9428b8087f68a4abef2df103 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 14:36:05 -0400 Subject: [PATCH 74/79] docs(cache): correct the Close bound, deferred count and probe wording Close can wait up to 4s, not 3s: a dial in flight, then the 1s final drain. deferred counts each deferral once, not failed retries. The probe write closes the breaker only within timeout. The breaker's WARN on opening is documented alongside the credentials ERROR. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01FyrXjhR7iDg33paioLQHFq --- docs/src/content/docs/configuration.mdx | 2 +- docs/src/content/docs/deployment.md | 2 +- internal/config/cache_redis.go | 5 +++-- 3 files changed, 5 insertions(+), 4 deletions(-) diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index c8878afee..dab012d38 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -168,7 +168,7 @@ The `redis` backend's settings, read only when `cache.backend` is `redis`. It is | `cache.redis.compress_min_bytes` | `WH_CACHE_REDIS_COMPRESS_MIN_BYTES` | `1024` | Results at least this large are zstd-compressed when that makes them smaller. `0` never compresses. | | `cache.redis.version_ttl` | `WH_CACHE_REDIS_VERSION_TTL` | `168h` | How long a table's or tenant's version token outlives its last write, so the tokens of dropped tables and removed tenants eventually expire. At least `2s`. An expired token only causes misses. | -**When the server is unreachable or misbehaves, the cache is bypassed; queries are not.** A failure or a timeout makes the lookup a miss and the fill a no-op. Five in a row, one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), or one refusing the credentials (`WRONGPASS`, `NOAUTH` — what an already-open connection meets once the password is rotated; logged at `ERROR`, once per opening and once per refused probe) open a circuit breaker that skips the server entirely until a probe write, every 5 s, succeeds. `NOPERM` does not open it: it names one key or command an ACL user cannot use, not every operation, so it should not bypass the cache for every tenant. Queries then go straight to ClickHouse, still coalesced per instance by `singleflight`. An invalidation the server did not take is kept and retried until it lands, and until then the instance that owes it bypasses the lookups it would orphan. `/readyz` does not depend on the cache. +**When the server is unreachable or misbehaves, the cache is bypassed; queries are not.** A failure or a timeout makes the lookup a miss and the fill a no-op. Five in a row, one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), or one refusing the credentials (`WRONGPASS`, `NOAUTH` — what an already-open connection meets once the password is rotated) open a circuit breaker that skips the server entirely until a probe write, every 5 s, succeeds within `timeout`. Each opening is logged once, at `ERROR` for the credentials and `WARN` otherwise; a failed probe opens it again, so a server that stays down logs every 5 s. `NOPERM` does not open it: it names one key or command an ACL user cannot use, not every operation, so it should not bypass the cache for every tenant. Queries then go straight to ClickHouse, still coalesced per instance by `singleflight`. An invalidation the server did not take is kept and retried until it lands, and until then the instance that owes it bypasses the lookups it would orphan. `/readyz` does not depend on the cache. **Boot does not wait for the server.** A malformed block (an address without a port, `mode: cluster` with `db` other than `0`, an unreadable or unparsable TLS file) refuses boot. A server that cannot be reached, or that refuses the credentials, does not: the process boots with the cache bypassed and keeps reconnecting, with backoff up to 30 s. A rejected credential (`WRONGPASS`, `NOAUTH`, or `NOPERM` for an ACL user missing a connection command) is logged at `ERROR` on every attempt; any other failure at `WARN`. This is deliberate: a rotated Redis password must not crash-loop every instance at once. Watch `wavehouse_cache_breaker_open`, which reads `1` while the cache is bypassed, including before the first connection. diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 0ba84aa5c..1d0c5ff16 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -443,7 +443,7 @@ The query-result cache is the layer that can be shared today. With the default ` **Coalescing stays per instance.** `singleflight` collapses identical concurrent queries within each instance, so a cold hot query costs at most one ClickHouse query per instance, not one per request. -**Metrics** (meter `wavehouse-cache`, every series labeled `backend="redis"`, no tenant label): `wavehouse_cache_lookups_total{result}` (`hit`, `miss`, `stale`, `bypass`, `error`), `wavehouse_cache_op_duration_seconds{op}` (`lookup`, `set`, `invalidate`), `wavehouse_cache_breaker_open` (1 while the cache is bypassed), `wavehouse_cache_invalidations_total{result}` (`ok` counts every bump that lands, retried ones included, and `deferred` each time one is put off, a repeat of one already owed included — the two overlap, not a split), `wavehouse_cache_invalidations_pending`, `wavehouse_cache_value_bytes`, `wavehouse_cache_oversize_total` and `wavehouse_cache_set_failures_total{reason}` (`oom`, `timeout`, `other`). Two signals are worth alerting on: `wavehouse_cache_breaker_open` at 1, or `wavehouse_cache_invalidations_pending` above 0, for more than a few minutes. +**Metrics** (meter `wavehouse-cache`, every series labeled `backend="redis"`, no tenant label): `wavehouse_cache_lookups_total{result}` (`hit`, `miss`, `stale`, `bypass`, `error`), `wavehouse_cache_op_duration_seconds{op}` (`lookup`, `set`, `invalidate`), `wavehouse_cache_breaker_open` (1 while the cache is bypassed), `wavehouse_cache_invalidations_total{result}` (`ok` counts every bump that lands, retried ones included, and `deferred` each bump an invalidation could not deliver when made, a repeat of one already owed included; a failed retry is not counted again — the two overlap, not a split), `wavehouse_cache_invalidations_pending`, `wavehouse_cache_value_bytes`, `wavehouse_cache_oversize_total` and `wavehouse_cache_set_failures_total{reason}` (`oom`, `timeout`, `other`). Two signals are worth alerting on: `wavehouse_cache_breaker_open` at 1, or `wavehouse_cache_invalidations_pending` above 0, for more than a few minutes. For local development, `docker compose -f deployments/compose/dependencies.yaml --profile redis up -d` starts a Redis on `localhost:6379` with no persistence. diff --git a/internal/config/cache_redis.go b/internal/config/cache_redis.go index be065266f..456006670 100644 --- a/internal/config/cache_redis.go +++ b/internal/config/cache_redis.go @@ -23,8 +23,9 @@ const ( // maxRedisTimeout caps cache.redis.timeout and cache.redis.dial_timeout. // Boot and the cache's Close each wait out a dial in flight: a connect and // a handshake, each bounded by dial_timeout, then for a cluster a topology -// read bounded by the larger of the two. At the caps that is at most 3s, -// inside the 5s budget Close shares with the stores released after it. +// read bounded by the larger of the two. At the caps that is at most 3s; +// Close then spends up to 1s delivering owed invalidations, so at most 4s, +// inside the 5s budget it shares with the stores released after it. const maxRedisTimeout = time.Second // CacheRedisConfig configures cache.backend=redis: one Redis-compatible server From 212ad133c9af84cd7202fc27f5b45b4d32789dcb Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 14:52:56 -0400 Subject: [PATCH 75/79] fix(cache): log a failed probe's reopening at DEBUG, not WARN A closed breaker opening still logs one WARN, or ERROR for rejected credentials. A failed probe reopening an already-open breaker now logs at DEBUG, so a server that stays down is one line for the outage rather than one every BreakerOpenFor. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01FyrXjhR7iDg33paioLQHFq --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- internal/cache/breaker.go | 44 ++++++++++++++++++--------- internal/cache/breaker_test.go | 14 ++++----- internal/cache/redis.go | 33 +++++++++++--------- internal/cache/redis_test.go | 20 +++++++++--- 6 files changed, 73 insertions(+), 42 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 792316245..ccbc214f2 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -11,7 +11,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added - **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `internal/ingest/worker.go`, `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `.github/workflows/{ci.yml,README.md}`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md,index.mdx,why-wavehouse.md,sdk/reference.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`) and `dial_timeout` (`1s`), each at most `1s` since boot and shutdown each wait out a connection attempt they bound, `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a valid port, more than one address in `standalone` mode, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The ingest worker's log of an invalidation that did not land drops from `ERROR` to `WARN`, since the shared backend defers and retries it: an outage would otherwise log an `ERROR` for every batch. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end; the e2e per-suite exclude now names what its run still can't reach instead — `internal/cache/pending.go` (the retry of an invalidation the server did not take, which needs an outage), `internal/config/cache_redis.go` (the block's own rejection paths) and `internal/cache/(local|version_manager).go` (the `local` backend, which e2e no longer runs) — and the unit and integration suites keep covering them. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits; a third runs the first one's hit, insert and fresh-miss lifecycle on the suite's own `cache.backend: local` app, since e2e no longer exercises that backend. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. -- **A Redis-compatible shared cache backend** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/cache.go`, `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`, `CONTRIBUTING.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart without persistence can only cause misses — `maxmemory-policy allkeys-lru` is safe. Restoring an RDB or AOF snapshot, or a backup, is a rollback instead (a restart after a crash that reloads the server's last save included, which stock Redis and Valkey make by default): the old tokens return with their values, so invalidations made since are undone until those entries' TTL. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like) or the credentials (`WRONGPASS` or `NOAUTH` after a password rotation), open a circuit breaker — logged once per opening, at `ERROR` for the credentials and `WARN` otherwise, which a failed probe repeats every 5 s while the server stays down — that skips the server until a probe write succeeds within the per-operation timeout (connections are replaced every minute, so after a failover behind a stable address the process reaches the new primary, and delivers the bumps it owes, within about a minute; the probe also allows up to twice the dial timeout for a reconnect, but the client's other connections must reconnect within the per-operation timeout, so that timeout should still exceed a reconnect), and deferred invalidations are retried until they land — the first at once, and at the probe's cadence while the breaker is open — collapsing to one tenant-wide bump per tenant past 100,000 keys; until one lands, the process that owes it bypasses the lookups it would orphan. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. `cache.backend: redis` selects it (the entry above). Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, snapshot-rollback, compression, stored-size, server-stops-answering, refused-writes (a demoted primary, a full `noeviction` server), rotated-credentials, slow-reconnect (over TLS), slow-server, failover-behind-a-stable-address and owed-invalidation cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. +- **A Redis-compatible shared cache backend** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/cache.go`, `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`, `CONTRIBUTING.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart without persistence can only cause misses — `maxmemory-policy allkeys-lru` is safe. Restoring an RDB or AOF snapshot, or a backup, is a rollback instead (a restart after a crash that reloads the server's last save included, which stock Redis and Valkey make by default): the old tokens return with their values, so invalidations made since are undone until those entries' TTL. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like) or the credentials (`WRONGPASS` or `NOAUTH` after a password rotation), open a circuit breaker — logged once per opening, at `ERROR` for the credentials and `WARN` otherwise, while a failed probe reopening it every 5 s logs only at `DEBUG` — that skips the server until a probe write succeeds within the per-operation timeout (connections are replaced every minute, so after a failover behind a stable address the process reaches the new primary, and delivers the bumps it owes, within about a minute; the probe also allows up to twice the dial timeout for a reconnect, but the client's other connections must reconnect within the per-operation timeout, so that timeout should still exceed a reconnect), and deferred invalidations are retried until they land — the first at once, and at the probe's cadence while the breaker is open — collapsing to one tenant-wide bump per tenant past 100,000 keys; until one lands, the process that owes it bypasses the lookups it would orphan. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. `cache.backend: redis` selects it (the entry above). Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, snapshot-rollback, compression, stored-size, server-stops-answering, refused-writes (a demoted primary, a full `noeviction` server), rotated-credentials, slow-reconnect (over TLS), slow-server, failover-behind-a-stable-address and owed-invalidation cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. - **Leases for work that must run in one process at a time, and the sweeper runs under one** (`internal/coord/` (new: `coord.go`, `local.go`, `elect.go`, `coordtest/`, + tests), `internal/app/{app,wire}.go` (+ tests), `internal/config/backends.go`, `internal/ingest/sweeper.go`, `.github/labeler.yml`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development,ingest-pipeline}.md`, `docs/src/content/docs/configuration.mdx`, `config.yaml`, `AGENTS.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `coord.Coordinator` hands out named leases (`TryAcquire` → a `Term` with a strictly increasing fencing `Token`, a `Done` channel and `Resign`; `ErrHeld` while another holder's term is live), and `coord.RunElected` runs a loop only while its process holds the lease, resigning when the loop returns and campaigning again every 2s. `coord.Local` is the in-process implementation, and `coordtest.Conformance` is the suite every implementation runs — the NATS KV backend that lets several replicas share one queue comes next. The sweeper now runs through `RunElected` under the `sweeper` lease; with the in-process coordinator the one process always holds it, so nothing changes for a single-process deployment beyond one `coord: elected` log line at startup. `coord.backend` now selects the coordinator (`local`, the only value), so a `config.Config` built without `config.Load` must name it as well as the other three layers' backends. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index f1b07a397..0c78a2c53 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -117,7 +117,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: one version per tenant, per (tenant, table) and per (tenant, table, scope), each keyed by its name alone and bumped in place, so the index holds one entry per live tenant, table and scope however often each is bumped ([#262](https://github.com/Wave-RF/WaveHouse/issues/262)). A query key folds the tenant's version and, for each dependency, its tenant's, table's and scope's, so bumping a table (a scopeless write) orphans every scope of it, and bumping one scope orphans that scope and the whole-table view — scope is reserved and empty today, so every write is the whole-table bump — all without touching the pool. Every field — the caller's query key, the tenant id, and each dependency's table and scope — is escaped and joined by `internal/keyenc` where the key is built, so a dot, a space or a `%` in a name is never read as a separator: each dependency renders as `..
.
..`, and the whole entry key is `|.||…`. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results. A tenant's version is a *generation*, unique within the process and handed out by the first key built for the tenant; `BumpTenant` (behind `InvalidateTenant`) drops the tenant's whole index, so the next key gets a fresh generation no cached entry folds, orphaning every cached result of the tenant in one step — a pipe result with no dependencies, and a table no bump ever keyed, included — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). `LocalCache.Prune` does the same for every tenant no longer served, which `internal/app` runs after each settings reload, so a tenant removed or rejected stops holding its index. A table bump drops the table's scope versions with it, since every key they were folded into also folds the old table version; and any bump (of a table, a scope or the tenant) under a tenant with no index is a no-op that records nothing, since the next key built for it gets a fresh generation no cached entry folds — so neither the `sharedTables` fan-out nor an insert still in flight for a tenant just pruned brings its index back. The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. - **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET` and `MGET`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot; the table and scope are escaped with `keyenc` after the fixed prefix, so a `:` in a name is never read as the separator. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, the hash over their escaped forms, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token (a string of another length, a list, a hash) is replaced, so a token lost to eviction, expiry, `FLUSHALL` or a restart without persistence is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. Restoring an RDB or AOF snapshot, or a backup, is not a loss but a rollback — a restart after a crash that reloads the server's last save included, which stock Redis and Valkey make by default: the old tokens come back with the values filed under them, so what was invalidated since is served again until its TTL, as after a failover to a replica that missed the bumps. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. While this process owes a bump on any of a lookup's tokens, that lookup is a bypass that files nothing: the bump would orphan whatever it found (other processes, which cannot know, serve those entries until it lands). `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing, each attempt bounded by `DialTimeout` (1 s), and a cluster client's topology read after the handshake by the larger of it and `Timeout`, which is also how long a connection waits on a silent server before it is redialed. Every connection is replaced after a minute (`clientOption`; rueidis retries what was in flight), because one that outlives a failover behind a stable address stays on the demoted node, which answers but refuses writes; the replacement re-resolves the address, so a bypassed process reaches the new primary, and delivers the bumps it owes, within about that long. rueidis dials a replacement lazily, under the context of the operation that lands on it, bounding the dial (TLS included) by `DialTimeout` and then the handshake by `DialTimeout` again, so a reconnect slower than `Timeout` fails that operation. The breaker's probe gives its write twice `DialTimeout` for a reconnect on top of `Timeout`, so such a reconnect still closes an open breaker. The allowance is for a reconnect only: a probe write slower than `Timeout` is repeated under `Timeout`, and the repeat decides, so a server answering slower than `Timeout` stays bypassed. But rueidis spreads commands over several connections (up to four to one server, by `GOMAXPROCS`, and one per cluster node) and the probe reconnects only the one it lands on, so size `Timeout` above a reconnect, or operations that land on the others keep failing. `cache.backend: redis` selects it: `internal/app`'s `wireCache` maps the boot config's `cache.redis` block onto `RedisConfig`, reading the TLS files, and releases it with the other components. - **redis_codec.go** — the key schema (`tokenKeys`, `bumpKeys`, `valueKey`, which take a `Namespace`'s raw names and escape them) and the value frame: format, flags, expiry, the token list, the payload, zstd-compressed from `CompressMinBytes` when that is smaller. A stored value is capped at `MaxValueBytes` and a decoded one at eight times that, which refuses a zip bomb planted in a shared server; an unknown format, as a newer process writes during a rolling upgrade, is a miss. -- **breaker.go** — the circuit breaker's state machine: `BreakerThreshold` (5) consecutive failures open it, `trip` opens it at once, and while it is open every operation skips the server; once `BreakerOpenFor` (5 s) has passed, `allow` hands one caller the probe, and only the probe's success closes it. What feeds it is `redis.go`'s. `record` counts a transport failure or a timeout against the server, and trips the breaker on an error reply that means no bump can land: one saying the server takes no writes right now, as `refusesWork` lists them — `READONLY` (a demoted primary), `MASTERDOWN`, `OOM` (full under `maxmemory-policy noeviction`, Redis's default, so give the cache an evicting policy), `NOREPLICAS`, `MISCONF`, `LOADING`, `BUSY`, `CLUSTERDOWN` — or one refusing the credentials, as `rejectsCredentials` lists them — `WRONGPASS`, `NOAUTH`, what a new connection's handshake meets after a password rotation. Each opening is logged once, at `ERROR` for the credentials and `WARN` otherwise, not once per operation in flight; a failed probe opens it afresh, so a server that stays down logs once every `BreakerOpenFor`. Any other reply, an error reply about one key or command (`WRONGTYPE`, `NOPERM`, `TRYAGAIN`) or one the backend cannot use included, counts as a success, and a caller that gave up first counts as nothing. The probe (`probe`) is a write, `SET :probe`, so a server that answers but refuses writes stays bypassed. +- **breaker.go** — the circuit breaker's state machine: `BreakerThreshold` (5) consecutive failures open it, `trip` opens it at once, and while it is open every operation skips the server; once `BreakerOpenFor` (5 s) has passed, `allow` hands one caller the probe, and only the probe's success closes it. What feeds it is `redis.go`'s. `record` counts a transport failure or a timeout against the server, and trips the breaker on an error reply that means no bump can land: one saying the server takes no writes right now, as `refusesWork` lists them — `READONLY` (a demoted primary), `MASTERDOWN`, `OOM` (full under `maxmemory-policy noeviction`, Redis's default, so give the cache an evicting policy), `NOREPLICAS`, `MISCONF`, `LOADING`, `BUSY`, `CLUSTERDOWN` — or one refusing the credentials, as `rejectsCredentials` lists them — `WRONGPASS`, `NOAUTH`, what a new connection's handshake meets after a password rotation. A closed breaker opening is logged once, at `ERROR` for the credentials and `WARN` otherwise, not once per operation in flight; a failed probe opens it afresh and logs at `DEBUG`, so a server that stays down logs one `WARN` or `ERROR` for the whole outage. Any other reply, an error reply about one key or command (`WRONGTYPE`, `NOPERM`, `TRYAGAIN`) or one the backend cannot use included, counts as a success, and a caller that gave up first counts as nothing. The probe (`probe`) is a write, `SET :probe`, so a server that answers but refuses writes stays bypassed. - **pending.go** — the invalidations owed (`pendingBumps`): kept per token key, repeats coalescing, and past `PendingMax` (100,000) keys collapsed to one tenant bump per affected tenant; `owesAny` tells `Lookup` which lookups to hold. `redis.go` delivers them: `Invalidate` sends its bumps 1,000 to a round trip and defers those the server does not take, and `drainLoop` retries them through `drain` until they land — the first one owed at once, then with backoff from 100 ms to 10 s, and while the breaker is open at each probe, which the loop starts when due, so a process that makes no lookups (ingest only) recovers as soon as one that does. Landing late is still correct: a fresh token orphans the pre-write entries and any fill made meanwhile. `Close` makes one last attempt at the bumps still owed, past the breaker (an open one is why they are owed); what that attempt cannot deliver is lost, and the entries they would orphan are served until their TTL — the same failure as a worker stopping between an insert and its invalidation. - **metrics.go** — the shared backend's instruments (meter `wavehouse-cache`, `backend="redis"`, no tenant attribute): `wavehouse_cache_lookups_total` by `result` (`hit`, `miss`, `stale` — filed under since-bumped tokens, `bypass` — server skipped, or held by a bump this process owes, `error`), `wavehouse_cache_op_duration_seconds` by `op` (`lookup`, `set`, `invalidate`), `wavehouse_cache_breaker_open` (1 while bypassed, including before the first connection), `wavehouse_cache_invalidations_total` by `result` (`ok` counts every bump that lands, retried ones included, and `deferred` each bump an invalidation could not deliver when made, a repeat of one already owed included; a failed retry is not counted again, so the two overlap), `wavehouse_cache_invalidations_pending` (bumps owed: entries they would orphan may be served stale meanwhile), `wavehouse_cache_value_bytes` (stored size), `wavehouse_cache_oversize_total` and `wavehouse_cache_set_failures_total` by `reason` (`oom`, `timeout`, `other`). - **cachetest** (`internal/testutil/cachetest`) — `Run(t, factory, Options)`, the conformance suite every `Cache` backend runs, each case on a fresh cache from the factory: what a hit, a miss and each kind of bump mean, independent of where entries and versions live. `Options` describes what a backend can do beyond the `Cache` contract, and leaving one unset skips the cases it enables: `MaxValueBytes` opens the oversize case; `NewPair` returns two instances over one shared store, for the cross-instance cases a shared backend must run; `Entries` counts a cache's entries, and without it the zero-snapshot case — no `Lookup` reads the key a zero `Snapshot` would land under, so only a count shows that one stored nothing — is skipped. diff --git a/internal/cache/breaker.go b/internal/cache/breaker.go index f3e065a53..d9d79af05 100644 --- a/internal/cache/breaker.go +++ b/internal/cache/breaker.go @@ -53,31 +53,47 @@ func (b *breaker) success() { b.failures, b.open, b.probing = 0, false, false } +// opening is what a failure or trip did to the breaker, so a caller logs an +// outage once rather than once per operation in flight or per failed probe. +type opening int + +const ( + unchanged opening = iota // still closed, or already open + opened // a closed breaker opened + reopened // a failed probe opened it for another period +) + +// openLocked opens the breaker and reports which opening that was. +func (b *breaker) openLocked() opening { + o := unchanged + switch { + case !b.open: + o = opened + case b.probing: + o = reopened + } + b.open, b.openedAt, b.probing = true, b.now(), false + return o +} + // failure records a call the server did not answer in time, and opens the -// breaker at the threshold — or at once, for a failed probe. It reports -// whether that opened a closed or probing breaker, as trip does. -func (b *breaker) failure() bool { +// breaker at the threshold — or at once, for a failed probe. +func (b *breaker) failure() opening { b.mu.Lock() defer b.mu.Unlock() b.failures++ if b.probing || b.failures >= b.threshold { - opened := !b.open || b.probing - b.open, b.openedAt, b.probing = true, b.now(), false - return opened + return b.openLocked() } - return false + return unchanged } // trip opens the breaker at once, for a reply that says the server cannot -// do the work: one is as conclusive as any number. It reports whether the -// breaker was closed or probing, so a caller logs once per opening and once -// per refused probe rather than once per operation in flight. -func (b *breaker) trip() bool { +// do the work: one is as conclusive as any number. +func (b *breaker) trip() opening { b.mu.Lock() defer b.mu.Unlock() - opened := !b.open || b.probing - b.open, b.openedAt, b.probing = true, b.now(), false - return opened + return b.openLocked() } func (b *breaker) isOpen() bool { diff --git a/internal/cache/breaker_test.go b/internal/cache/breaker_test.go index 273a871c0..c6e0afbbe 100644 --- a/internal/cache/breaker_test.go +++ b/internal/cache/breaker_test.go @@ -26,11 +26,11 @@ func TestBreaker(t *testing.T) { b.failure() b.success() // a success resets the run b.failure() - assert.False(t, b.failure()) + assert.Equal(t, unchanged, b.failure()) assert.False(t, b.isOpen(), "two in a row is below the threshold") - assert.True(t, b.failure(), "it opened") + assert.Equal(t, opened, b.failure()) assert.True(t, b.isOpen()) - assert.False(t, b.failure(), "already open") + assert.Equal(t, unchanged, b.failure(), "already open") ok, probe = allow() assert.False(t, ok, "open: skip the server") @@ -44,7 +44,7 @@ func TestBreaker(t *testing.T) { assert.False(t, ok) assert.False(t, probe, "one probe at a time") - assert.True(t, b.failure(), "the probe failed: open for another period") + assert.Equal(t, reopened, b.failure(), "the probe failed: open for another period") assert.True(t, b.isOpen()) clock.t = clock.t.Add(4 * time.Second) _, probe = allow() @@ -69,9 +69,9 @@ func TestBreaker_TripAndProbeSchedule(t *testing.T) { _, open := b.untilProbe() assert.False(t, open) - assert.True(t, b.trip(), "it opened") + assert.Equal(t, opened, b.trip()) assert.True(t, b.isOpen(), "no threshold for a refusal") - assert.False(t, b.trip(), "already open") + assert.Equal(t, unchanged, b.trip(), "already open") b.success() assert.True(t, b.isOpen(), "a success that is not the probe's leaves it open") d, open := b.untilProbe() @@ -87,7 +87,7 @@ func TestBreaker_TripAndProbeSchedule(t *testing.T) { require.True(t, probe) d, _ = b.untilProbe() assert.Equal(t, 5*time.Second, d, "while the probe runs, wait out a whole period") - assert.True(t, b.trip(), "the probe was refused too") + assert.Equal(t, reopened, b.trip(), "the probe was refused too") d, _ = b.untilProbe() assert.Equal(t, 5*time.Second, d) diff --git a/internal/cache/redis.go b/internal/cache/redis.go index 1889ee727..258387454 100644 --- a/internal/cache/redis.go +++ b/internal/cache/redis.go @@ -370,31 +370,36 @@ func (r *RedisCache) record(parent context.Context, err error) { if parent.Err() != nil { return } - if r.breaker.failure() { - slog.WarnContext(parent, "cache: redis not answering; bypassing the cache", - "addrs", r.cfg.Addrs, "error", err) - } + logOpening(parent, r.breaker.failure(), slog.LevelWarn, "cache: redis not answering; bypassing the cache", + "addrs", r.cfg.Addrs, "error", err) } -// recordReply is record for an error reply. The breaker opening is logged -// once, not once per operation in flight. +// recordReply is record for an error reply. func (r *RedisCache) recordReply(parent context.Context, msg string) { switch { case rejectsCredentials(msg): - if r.breaker.trip() { - slog.ErrorContext(parent, "cache: redis rejected the credentials; bypassing the cache until they work", - "addrs", r.cfg.Addrs, "error", msg) - } + logOpening(parent, r.breaker.trip(), slog.LevelError, "cache: redis rejected the credentials; bypassing the cache until they work", + "addrs", r.cfg.Addrs, "error", msg) case refusesWork(msg): - if r.breaker.trip() { - slog.WarnContext(parent, "cache: redis refusing writes; bypassing the cache", - "addrs", r.cfg.Addrs, "reply", msg) - } + logOpening(parent, r.breaker.trip(), slog.LevelWarn, "cache: redis refusing writes; bypassing the cache", + "addrs", r.cfg.Addrs, "reply", msg) default: r.breaker.success() } } +// logOpening logs a closed breaker opening at level, and a failed probe +// reopening it at DEBUG: a long outage is one line, not one per probe. +func logOpening(ctx context.Context, o opening, level slog.Level, msg string, args ...any) { + switch o { + case opened: + slog.Log(ctx, level, msg, args...) + case reopened: + slog.Log(ctx, slog.LevelDebug, msg, args...) + case unchanged: + } +} + // refusesWork reports whether an error reply says the server takes no // writes from anyone right now, so no bump can land: a replica (READONLY, // or MASTERDOWN, which refuses reads too), memory full under noeviction diff --git a/internal/cache/redis_test.go b/internal/cache/redis_test.go index 5f9d5cec4..c830873dd 100644 --- a/internal/cache/redis_test.go +++ b/internal/cache/redis_test.go @@ -163,15 +163,16 @@ func TestRedis_Record(t *testing.T) { assert.True(t, r.breaker.isOpen()) } -// Each opening of the breaker logs one WARN naming why; the operations that -// fail while it is open log nothing. Not parallel: it captures the default -// logger. +// A closed breaker opening logs one WARN naming why; the operations that +// fail while it is open log nothing, and a failed probe reopening it logs at +// DEBUG. Not parallel: it captures the default logger. func TestRedis_BreakerOpeningLogsOnce(t *testing.T) { - buf := logtest.Capture(t, slog.LevelWarn) + buf := logtest.Capture(t, slog.LevelDebug) clock := &fakeClock{t: time.Unix(0, 0)} r := &RedisCache{breaker: newBreaker(2, time.Second, clock.now)} live := context.Background() warns := func() int { return strings.Count(buf.String(), `"level":"WARN"`) } + debugs := func() int { return strings.Count(buf.String(), `"level":"DEBUG"`) } r.recordReply(live, "READONLY You can't write against a read only replica.") r.recordReply(live, "READONLY You can't write against a read only replica.") @@ -197,7 +198,16 @@ func TestRedis_BreakerOpeningLogsOnce(t *testing.T) { _, probe = r.breaker.allow() require.True(t, probe) r.record(live, context.DeadlineExceeded) - assert.Equal(t, 3, warns(), "a failed probe opens it again") + assert.True(t, r.breaker.isOpen(), "a failed probe opens it again") + assert.Equal(t, 2, warns(), "not at WARN") + assert.Equal(t, 1, debugs(), buf.String()) + + clock.t = clock.t.Add(time.Second) + _, probe = r.breaker.allow() + require.True(t, probe) + r.recordReply(live, "READONLY You can't write against a read only replica.") + assert.Equal(t, 2, warns(), "a refused probe is a reopening too") + assert.Equal(t, 2, debugs(), buf.String()) } func TestRejectsCredentials(t *testing.T) { From a7cd351566f7f57bd7a0cbbf34b2099b31dd4ad9 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 15:01:20 -0400 Subject: [PATCH 76/79] fix(cache): log a reopening at its own level when its cause changed A failed probe reopening the breaker logged at DEBUG whatever its cause, so an outage that turned into rejected credentials after a restart left only the first opening's WARN. A reopening for another cause than the one last logged is now logged at its own level. The configuration page still said a down server logs every 5 s; it now says one line. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01FyrXjhR7iDg33paioLQHFq --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/configuration.mdx | 2 +- internal/cache/redis.go | 27 ++++++++++++------ internal/cache/redis_test.go | 37 ++++++++++++++++--------- 5 files changed, 45 insertions(+), 25 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index ccbc214f2..5c69ebc50 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -11,7 +11,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added - **`cache.backend: redis` shares the query cache across instances** (`internal/config/{cache_redis,backends,config}.go` (+ tests), `internal/app/wire.go` (+ tests), `internal/ingest/worker.go`, `tests/integration/shared_cache_test.go`, `scripts/orchestrator/main.go`, `tests/e2e/fixtures/config.yaml`, `.testcoverage.yml`, `deployments/compose/dependencies.yaml`, `config.yaml`, `.github/workflows/{ci.yml,README.md}`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md,settings-directory.mdx,getting-started.md,pipes.mdx,api.md,development.md,index.mdx,why-wavehouse.md,sdk/reference.md}`, `README.md`, `AGENTS.md`): PR E4 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config's `cache.backend` now takes `redis`, configured by a new `cache.redis` block (`WH_CACHE_REDIS_*`): `addrs` (required), `mode` (`standalone` or `cluster`; `sentinel` refuses boot until [#656](https://github.com/Wave-RF/WaveHouse/issues/656)), `username`, `password` (a secret — set it through the environment), `db`, `tls.{enabled,ca_file,cert_file,key_file,server_name,insecure_skip_verify}`, `key_prefix` (`wh`), `timeout` (`100ms`) and `dial_timeout` (`1s`), each at most `1s` since boot and shutdown each wait out a connection attempt they bound, `max_value_bytes` (1 MiB), `compress_min_bytes` (`1024`; `0` never compresses) and `version_ttl` (`168h`). Every instance pointed at one server shares its cached results, and an insert on any instance invalidates every instance's. It is also the shared cache a split of [`roles`](https://github.com/Wave-RF/WaveHouse/pull/622) needs: the refusal of `api` without `ingest`, or the reverse, over `cache.backend: local` now names `redis`, though every split is still refused while the queue is embedded. A malformed block — no address, an address without a valid port, more than one address in `standalone` mode, a URL-style address (refused without repeating it, since it may carry a password), an unknown mode, `db` other than `0` in cluster mode, an unreadable TLS file, a TLS key set while `tls.enabled` is off — refuses boot; an unreachable server, or one that rejects the password, does not: the process boots with the cache bypassed and keeps reconnecting, logging a rejected password at `ERROR` on every attempt, so a rotated secret cannot crash-loop every instance at once. `insecure_skip_verify`, and a `cache.redis.addrs` set while `cache.backend` is `local`, are logged at `WARN` at boot. The ingest worker's log of an invalidation that did not land drops from `ERROR` to `WARN`, since the shared backend defers and retries it: an outage would otherwise log an `ERROR` for every batch. The e2e suite now runs against a Redis testcontainer with `cache.backend: redis`, so the shared backend is exercised end to end; the e2e per-suite exclude now names what its run still can't reach instead — `internal/cache/pending.go` (the retry of an invalidation the server did not take, which needs an outage), `internal/config/cache_redis.go` (the block's own rejection paths) and `internal/cache/(local|version_manager).go` (the `local` backend, which e2e no longer runs) — and the unit and integration suites keep covering them. An integration test boots two instances over one Redis and one ClickHouse: a result one fills is a hit for the other, and a row ingested through one is served fresh by the other on its next query, well inside the stale entry's TTL; another pauses Redis and checks queries keep succeeding from ClickHouse, then turn back to hits; a third runs the first one's hit, insert and fresh-miss lifecycle on the suite's own `cache.backend: local` app, since e2e no longer exercises that backend. `deployments/compose/dependencies.yaml` gains an optional `redis` profile for local multi-instance work. The deployment guide gains a "Multiple instances and the shared cache" section: what each instance keeps to itself, what a reader on another instance can see and when, `maxmemory-policy`, and the metrics to alert on. -- **A Redis-compatible shared cache backend** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/cache.go`, `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`, `CONTRIBUTING.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart without persistence can only cause misses — `maxmemory-policy allkeys-lru` is safe. Restoring an RDB or AOF snapshot, or a backup, is a rollback instead (a restart after a crash that reloads the server's last save included, which stock Redis and Valkey make by default): the old tokens return with their values, so invalidations made since are undone until those entries' TTL. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like) or the credentials (`WRONGPASS` or `NOAUTH` after a password rotation), open a circuit breaker — logged once per opening, at `ERROR` for the credentials and `WARN` otherwise, while a failed probe reopening it every 5 s logs only at `DEBUG` — that skips the server until a probe write succeeds within the per-operation timeout (connections are replaced every minute, so after a failover behind a stable address the process reaches the new primary, and delivers the bumps it owes, within about a minute; the probe also allows up to twice the dial timeout for a reconnect, but the client's other connections must reconnect within the per-operation timeout, so that timeout should still exceed a reconnect), and deferred invalidations are retried until they land — the first at once, and at the probe's cadence while the breaker is open — collapsing to one tenant-wide bump per tenant past 100,000 keys; until one lands, the process that owes it bypasses the lookups it would orphan. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. `cache.backend: redis` selects it (the entry above). Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, snapshot-rollback, compression, stored-size, server-stops-answering, refused-writes (a demoted primary, a full `noeviction` server), rotated-credentials, slow-reconnect (over TLS), slow-server, failover-behind-a-stable-address and owed-invalidation cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. +- **A Redis-compatible shared cache backend** (`internal/cache/{redis,redis_codec,breaker,pending,metrics}.go` (+ tests), `internal/cache/cache.go`, `internal/cache/redis_integration_test.go`, `Makefile`, `go.mod`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development}.md`, `AGENTS.md`, `CONTRIBUTING.md`): PR E3 of the distributed-deployment epic ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). `cache.RedisCache` keeps query results and their versions in one Redis, Valkey, Dragonfly, ElastiCache or MemoryDB server shared by every process, so an insert one process makes invalidates what every other process has cached. Versions are random tokens, one per tenant, table and scope, under the tenant's hash tag, with the table and scope escaped into the key (`internal/keyenc`) so no two names share a token; a bump sets a fresh one, and a value carries the tokens it was computed under, so a lookup is one pipelined round trip (`MGET` of the tokens plus `GET` of the value, no scripts) and a token lost to eviction, expiry, `FLUSHALL` or a restart without persistence can only cause misses — `maxmemory-policy allkeys-lru` is safe. Restoring an RDB or AOF snapshot, or a backup, is a rollback instead (a restart after a crash that reloads the server's last save included, which stock Redis and Valkey make by default): the old tokens return with their values, so invalidations made since are undone until those entries' TTL. Values of 1 KiB or more are zstd-compressed when that makes them smaller, and a value over 1 MiB stored is not cached. A server that fails or takes longer than the per-operation timeout (100 ms) is a miss, a skipped fill and a deferred invalidation, never a failed query; five failures in a row, or one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like) or the credentials (`WRONGPASS` or `NOAUTH` after a password rotation), open a circuit breaker — logged once per opening, at `ERROR` for the credentials and `WARN` otherwise, while a failed probe reopening it every 5 s logs only at `DEBUG` unless its cause changed — that skips the server until a probe write succeeds within the per-operation timeout (connections are replaced every minute, so after a failover behind a stable address the process reaches the new primary, and delivers the bumps it owes, within about a minute; the probe also allows up to twice the dial timeout for a reconnect, but the client's other connections must reconnect within the per-operation timeout, so that timeout should still exceed a reconnect), and deferred invalidations are retried until they land — the first at once, and at the probe's cadence while the breaker is open — collapsing to one tenant-wide bump per tenant past 100,000 keys; until one lands, the process that owes it bypasses the lookups it would orphan. New metrics: `wavehouse_cache_lookups_total{backend,result}`, `wavehouse_cache_op_duration_seconds{backend,op}`, `wavehouse_cache_breaker_open{backend}`, `wavehouse_cache_invalidations_total{backend,result}`, `wavehouse_cache_invalidations_pending{backend}`, `wavehouse_cache_value_bytes{backend}`, `wavehouse_cache_oversize_total{backend}`, `wavehouse_cache_set_failures_total{backend,reason}`. `cache.backend: redis` selects it (the entry above). Tested against Redis 8.10, Valkey 8.1, Dragonfly 2.0 and a Redis Cluster node by the conformance suite, plus lost-token, snapshot-rollback, compression, stored-size, server-stops-answering, refused-writes (a demoted primary, a full `noeviction` server), rotated-credentials, slow-reconnect (over TLS), slow-server, failover-behind-a-stable-address and owed-invalidation cases; `make test-integration` now also runs `internal/cache`'s integration-tagged tests. Adds `github.com/redis/rueidis` (Redis org, Apache-2.0; its only runtime dependency is `golang.org/x/sys`) and makes `github.com/klauspost/compress` a direct dependency. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. - **Leases for work that must run in one process at a time, and the sweeper runs under one** (`internal/coord/` (new: `coord.go`, `local.go`, `elect.go`, `coordtest/`, + tests), `internal/app/{app,wire}.go` (+ tests), `internal/config/backends.go`, `internal/ingest/sweeper.go`, `.github/labeler.yml`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development,ingest-pipeline}.md`, `docs/src/content/docs/configuration.mdx`, `config.yaml`, `AGENTS.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `coord.Coordinator` hands out named leases (`TryAcquire` → a `Term` with a strictly increasing fencing `Token`, a `Done` channel and `Resign`; `ErrHeld` while another holder's term is live), and `coord.RunElected` runs a loop only while its process holds the lease, resigning when the loop returns and campaigning again every 2s. `coord.Local` is the in-process implementation, and `coordtest.Conformance` is the suite every implementation runs — the NATS KV backend that lets several replicas share one queue comes next. The sweeper now runs through `RunElected` under the `sweeper` lease; with the in-process coordinator the one process always holds it, so nothing changes for a single-process deployment beyond one `coord: elected` log line at startup. `coord.backend` now selects the coordinator (`local`, the only value), so a `config.Config` built without `config.Load` must name it as well as the other three layers' backends. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 0c78a2c53..341095dc0 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -117,7 +117,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **version_manager.go** — `VersionManager`, the invalidation index behind `Invalidate` and `InvalidateTenant`: one version per tenant, per (tenant, table) and per (tenant, table, scope), each keyed by its name alone and bumped in place, so the index holds one entry per live tenant, table and scope however often each is bumped ([#262](https://github.com/Wave-RF/WaveHouse/issues/262)). A query key folds the tenant's version and, for each dependency, its tenant's, table's and scope's, so bumping a table (a scopeless write) orphans every scope of it, and bumping one scope orphans that scope and the whole-table view — scope is reserved and empty today, so every write is the whole-table bump — all without touching the pool. Every field — the caller's query key, the tenant id, and each dependency's table and scope — is escaped and joined by `internal/keyenc` where the key is built, so a dot, a space or a `%` in a name is never read as a separator: each dependency renders as `..
.
..`, and the whole entry key is `|.||…`. The tenant leads every key ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8): the same table under two tenants is two namespaces, so a bump through `Invalidate` under one tenant never touches — and a read under one tenant is never served — the other's results. A tenant's version is a *generation*, unique within the process and handed out by the first key built for the tenant; `BumpTenant` (behind `InvalidateTenant`) drops the tenant's whole index, so the next key gets a fresh generation no cached entry folds, orphaning every cached result of the tenant in one step — a pipe result with no dependencies, and a table no bump ever keyed, included — for a tenant back on a pool after an absence from the fan-out, or moved to another address or database (story 6). `LocalCache.Prune` does the same for every tenant no longer served, which `internal/app` runs after each settings reload, so a tenant removed or rejected stops holding its index. A table bump drops the table's scope versions with it, since every key they were folded into also folds the old table version; and any bump (of a table, a scope or the tenant) under a tenant with no index is a no-op that records nothing, since the next key built for it gets a fresh generation no cached entry folds — so neither the `sharedTables` fan-out nor an insert still in flight for a tenant just pruned brings its index back. The index is per tenant; the cross-tenant invalidation an insert into a shared table needs is not the index's but the wiring's: `internal/app` hands the ingest worker a cache (`sharedTables`) that repeats each bump under every tenant on the same ClickHouse address and database. - **redis.go** — `RedisCache`, the shared backend: one Redis-compatible server (Redis, Valkey, Dragonfly, ElastiCache, MemoryDB — only `GET`, `SET` and `MGET`, no scripts, no client tracking) holds every process's results and versions, so a bump one process makes orphans what every process cached. Versions are random 8-byte tokens in a flat key space, one per tenant (`:{t}:T`), table (`:B:
`) and scope (`:S:
:`, the empty scope being the whole-table view), all under the tenant's hash tag so they share one cluster slot; the table and scope are escaped with `keyenc` after the fixed prefix, so a `:` in a name is never read as the separator. A dependency folds three: the tenant's, its table's and its scope's; a scopeless write bumps `B:
`, a scoped one `S:
:` and `S:
:`, `InvalidateTenant` the tenant's — the lattice `VersionManager` encodes. `Lookup` pipelines an `MGET` of the tokens with a `GET` of the value in one round trip; the value (`:q::`, the hash over their escaped forms, no hash tag, so a tenant's values spread across shards) carries the tokens it was filed under and is a hit only while they are all current. A missing token is created (`SET NX`) and read back, never read as a value, and a token key holding anything but a token (a string of another length, a list, a hash) is replaced, so a token lost to eviction, expiry, `FLUSHALL` or a restart without persistence is a miss for everything under it, never a revival — any `maxmemory-policy` that evicts is safe. Restoring an RDB or AOF snapshot, or a backup, is not a loss but a rollback — a restart after a crash that reloads the server's last save included, which stock Redis and Valkey make by default: the old tokens come back with the values filed under them, so what was invalidated since is served again until its TTL, as after a failover to a replica that missed the bumps. A failure or a timeout past `Timeout` (100 ms) is a miss, a skipped fill and a deferred invalidation; queries never fail on the cache. While this process owes a bump on any of a lookup's tokens, that lookup is a bypass that files nothing: the bump would orphan whatever it found (other processes, which cannot know, serve those entries until it lands). `NewRedis` never fails on an unreachable server: the cache starts bypassed and keeps dialing, each attempt bounded by `DialTimeout` (1 s), and a cluster client's topology read after the handshake by the larger of it and `Timeout`, which is also how long a connection waits on a silent server before it is redialed. Every connection is replaced after a minute (`clientOption`; rueidis retries what was in flight), because one that outlives a failover behind a stable address stays on the demoted node, which answers but refuses writes; the replacement re-resolves the address, so a bypassed process reaches the new primary, and delivers the bumps it owes, within about that long. rueidis dials a replacement lazily, under the context of the operation that lands on it, bounding the dial (TLS included) by `DialTimeout` and then the handshake by `DialTimeout` again, so a reconnect slower than `Timeout` fails that operation. The breaker's probe gives its write twice `DialTimeout` for a reconnect on top of `Timeout`, so such a reconnect still closes an open breaker. The allowance is for a reconnect only: a probe write slower than `Timeout` is repeated under `Timeout`, and the repeat decides, so a server answering slower than `Timeout` stays bypassed. But rueidis spreads commands over several connections (up to four to one server, by `GOMAXPROCS`, and one per cluster node) and the probe reconnects only the one it lands on, so size `Timeout` above a reconnect, or operations that land on the others keep failing. `cache.backend: redis` selects it: `internal/app`'s `wireCache` maps the boot config's `cache.redis` block onto `RedisConfig`, reading the TLS files, and releases it with the other components. - **redis_codec.go** — the key schema (`tokenKeys`, `bumpKeys`, `valueKey`, which take a `Namespace`'s raw names and escape them) and the value frame: format, flags, expiry, the token list, the payload, zstd-compressed from `CompressMinBytes` when that is smaller. A stored value is capped at `MaxValueBytes` and a decoded one at eight times that, which refuses a zip bomb planted in a shared server; an unknown format, as a newer process writes during a rolling upgrade, is a miss. -- **breaker.go** — the circuit breaker's state machine: `BreakerThreshold` (5) consecutive failures open it, `trip` opens it at once, and while it is open every operation skips the server; once `BreakerOpenFor` (5 s) has passed, `allow` hands one caller the probe, and only the probe's success closes it. What feeds it is `redis.go`'s. `record` counts a transport failure or a timeout against the server, and trips the breaker on an error reply that means no bump can land: one saying the server takes no writes right now, as `refusesWork` lists them — `READONLY` (a demoted primary), `MASTERDOWN`, `OOM` (full under `maxmemory-policy noeviction`, Redis's default, so give the cache an evicting policy), `NOREPLICAS`, `MISCONF`, `LOADING`, `BUSY`, `CLUSTERDOWN` — or one refusing the credentials, as `rejectsCredentials` lists them — `WRONGPASS`, `NOAUTH`, what a new connection's handshake meets after a password rotation. A closed breaker opening is logged once, at `ERROR` for the credentials and `WARN` otherwise, not once per operation in flight; a failed probe opens it afresh and logs at `DEBUG`, so a server that stays down logs one `WARN` or `ERROR` for the whole outage. Any other reply, an error reply about one key or command (`WRONGTYPE`, `NOPERM`, `TRYAGAIN`) or one the backend cannot use included, counts as a success, and a caller that gave up first counts as nothing. The probe (`probe`) is a write, `SET :probe`, so a server that answers but refuses writes stays bypassed. +- **breaker.go** — the circuit breaker's state machine: `BreakerThreshold` (5) consecutive failures open it, `trip` opens it at once, and while it is open every operation skips the server; once `BreakerOpenFor` (5 s) has passed, `allow` hands one caller the probe, and only the probe's success closes it. What feeds it is `redis.go`'s. `record` counts a transport failure or a timeout against the server, and trips the breaker on an error reply that means no bump can land: one saying the server takes no writes right now, as `refusesWork` lists them — `READONLY` (a demoted primary), `MASTERDOWN`, `OOM` (full under `maxmemory-policy noeviction`, Redis's default, so give the cache an evicting policy), `NOREPLICAS`, `MISCONF`, `LOADING`, `BUSY`, `CLUSTERDOWN` — or one refusing the credentials, as `rejectsCredentials` lists them — `WRONGPASS`, `NOAUTH`, what a new connection's handshake meets after a password rotation. A closed breaker opening is logged once, at `ERROR` for the credentials and `WARN` otherwise, not once per operation in flight; a failed probe opens it afresh and logs at `DEBUG` (`logOpening`), unless its cause differs from the one last logged, which is logged at its own level; so a server that stays down logs one `WARN` or `ERROR` for the whole outage. Any other reply, an error reply about one key or command (`WRONGTYPE`, `NOPERM`, `TRYAGAIN`) or one the backend cannot use included, counts as a success, and a caller that gave up first counts as nothing. The probe (`probe`) is a write, `SET :probe`, so a server that answers but refuses writes stays bypassed. - **pending.go** — the invalidations owed (`pendingBumps`): kept per token key, repeats coalescing, and past `PendingMax` (100,000) keys collapsed to one tenant bump per affected tenant; `owesAny` tells `Lookup` which lookups to hold. `redis.go` delivers them: `Invalidate` sends its bumps 1,000 to a round trip and defers those the server does not take, and `drainLoop` retries them through `drain` until they land — the first one owed at once, then with backoff from 100 ms to 10 s, and while the breaker is open at each probe, which the loop starts when due, so a process that makes no lookups (ingest only) recovers as soon as one that does. Landing late is still correct: a fresh token orphans the pre-write entries and any fill made meanwhile. `Close` makes one last attempt at the bumps still owed, past the breaker (an open one is why they are owed); what that attempt cannot deliver is lost, and the entries they would orphan are served until their TTL — the same failure as a worker stopping between an insert and its invalidation. - **metrics.go** — the shared backend's instruments (meter `wavehouse-cache`, `backend="redis"`, no tenant attribute): `wavehouse_cache_lookups_total` by `result` (`hit`, `miss`, `stale` — filed under since-bumped tokens, `bypass` — server skipped, or held by a bump this process owes, `error`), `wavehouse_cache_op_duration_seconds` by `op` (`lookup`, `set`, `invalidate`), `wavehouse_cache_breaker_open` (1 while bypassed, including before the first connection), `wavehouse_cache_invalidations_total` by `result` (`ok` counts every bump that lands, retried ones included, and `deferred` each bump an invalidation could not deliver when made, a repeat of one already owed included; a failed retry is not counted again, so the two overlap), `wavehouse_cache_invalidations_pending` (bumps owed: entries they would orphan may be served stale meanwhile), `wavehouse_cache_value_bytes` (stored size), `wavehouse_cache_oversize_total` and `wavehouse_cache_set_failures_total` by `reason` (`oom`, `timeout`, `other`). - **cachetest** (`internal/testutil/cachetest`) — `Run(t, factory, Options)`, the conformance suite every `Cache` backend runs, each case on a fresh cache from the factory: what a hit, a miss and each kind of bump mean, independent of where entries and versions live. `Options` describes what a backend can do beyond the `Cache` contract, and leaving one unset skips the cases it enables: `MaxValueBytes` opens the oversize case; `NewPair` returns two instances over one shared store, for the cross-instance cases a shared backend must run; `Entries` counts a cache's entries, and without it the zero-snapshot case — no `Lookup` reads the key a zero `Snapshot` would land under, so only a count shows that one stored nothing — is skipped. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 4efecf703..2bea00fab 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -168,7 +168,7 @@ The `redis` backend's settings, read only when `cache.backend` is `redis`. It is | `cache.redis.compress_min_bytes` | `WH_CACHE_REDIS_COMPRESS_MIN_BYTES` | `1024` | Results at least this large are zstd-compressed when that makes them smaller. `0` never compresses. | | `cache.redis.version_ttl` | `WH_CACHE_REDIS_VERSION_TTL` | `168h` | How long a table's or tenant's version token outlives its last write, so the tokens of dropped tables and removed tenants eventually expire. At least `2s`. An expired token only causes misses. | -**When the server is unreachable or misbehaves, the cache is bypassed; queries are not.** A failure or a timeout makes the lookup a miss and the fill a no-op. Five in a row, one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), or one refusing the credentials (`WRONGPASS`, `NOAUTH` — what an already-open connection meets once the password is rotated) open a circuit breaker that skips the server entirely until a probe write, every 5 s, succeeds within `timeout`. Each opening is logged once, at `ERROR` for the credentials and `WARN` otherwise; a failed probe opens it again, so a server that stays down logs every 5 s. `NOPERM` does not open it: it names one key or command an ACL user cannot use, not every operation, so it should not bypass the cache for every tenant. Queries then go straight to ClickHouse, still coalesced per instance by `singleflight`. An invalidation the server did not take is kept and retried until it lands, and until then the instance that owes it bypasses the lookups it would orphan. `/readyz` does not depend on the cache. +**When the server is unreachable or misbehaves, the cache is bypassed; queries are not.** A failure or a timeout makes the lookup a miss and the fill a no-op. Five in a row, one reply refusing writes (`READONLY` from a demoted primary, `OOM` when full under `noeviction`, and the like), or one refusing the credentials (`WRONGPASS`, `NOAUTH` — what an already-open connection meets once the password is rotated) open a circuit breaker that skips the server entirely until a probe write, every 5 s, succeeds within `timeout`. A closed breaker opening is logged once, at `ERROR` for the credentials and `WARN` otherwise; a failed probe opens it again and logs at `DEBUG`, unless it failed for another cause than the one last logged (rejected credentials after a restart, say), which is logged at its own level. A server that stays down is one line for the outage. `NOPERM` does not open it: it names one key or command an ACL user cannot use, not every operation, so it should not bypass the cache for every tenant. Queries then go straight to ClickHouse, still coalesced per instance by `singleflight`. An invalidation the server did not take is kept and retried until it lands, and until then the instance that owes it bypasses the lookups it would orphan. `/readyz` does not depend on the cache. **Boot does not wait for the server.** A malformed block (an address without a port, `mode: cluster` with `db` other than `0`, an unreadable or unparsable TLS file) refuses boot. A server that cannot be reached, or that refuses the credentials, does not: the process boots with the cache bypassed and keeps reconnecting, with backoff up to 30 s. A rejected credential (`WRONGPASS`, `NOAUTH`, or `NOPERM` for an ACL user missing a connection command) is logged at `ERROR` on every attempt; any other failure at `WARN`. This is deliberate: a rotated Redis password must not crash-loop every instance at once. Watch `wavehouse_cache_breaker_open`, which reads `1` while the cache is bypassed, including before the first connection. diff --git a/internal/cache/redis.go b/internal/cache/redis.go index 258387454..849b30631 100644 --- a/internal/cache/redis.go +++ b/internal/cache/redis.go @@ -214,6 +214,8 @@ type RedisCache struct { wg sync.WaitGroup wake chan struct{} closeOnce sync.Once + + openCause atomic.Pointer[string] // message of the last opening logged at its own level } var _ Cache = (*RedisCache)(nil) @@ -370,7 +372,7 @@ func (r *RedisCache) record(parent context.Context, err error) { if parent.Err() != nil { return } - logOpening(parent, r.breaker.failure(), slog.LevelWarn, "cache: redis not answering; bypassing the cache", + r.logOpening(parent, r.breaker.failure(), slog.LevelWarn, "cache: redis not answering; bypassing the cache", "addrs", r.cfg.Addrs, "error", err) } @@ -378,10 +380,10 @@ func (r *RedisCache) record(parent context.Context, err error) { func (r *RedisCache) recordReply(parent context.Context, msg string) { switch { case rejectsCredentials(msg): - logOpening(parent, r.breaker.trip(), slog.LevelError, "cache: redis rejected the credentials; bypassing the cache until they work", + r.logOpening(parent, r.breaker.trip(), slog.LevelError, "cache: redis rejected the credentials; bypassing the cache until they work", "addrs", r.cfg.Addrs, "error", msg) case refusesWork(msg): - logOpening(parent, r.breaker.trip(), slog.LevelWarn, "cache: redis refusing writes; bypassing the cache", + r.logOpening(parent, r.breaker.trip(), slog.LevelWarn, "cache: redis refusing writes; bypassing the cache", "addrs", r.cfg.Addrs, "reply", msg) default: r.breaker.success() @@ -389,15 +391,22 @@ func (r *RedisCache) recordReply(parent context.Context, msg string) { } // logOpening logs a closed breaker opening at level, and a failed probe -// reopening it at DEBUG: a long outage is one line, not one per probe. -func logOpening(ctx context.Context, o opening, level slog.Level, msg string, args ...any) { +// reopening it at DEBUG: a long outage is one line, not one per probe. A +// reopening for another cause than the one last logged — rejected +// credentials after a restart, say — is logged at its own level. +func (r *RedisCache) logOpening(ctx context.Context, o opening, level slog.Level, msg string, args ...any) { switch o { - case opened: - slog.Log(ctx, level, msg, args...) - case reopened: - slog.Log(ctx, slog.LevelDebug, msg, args...) case unchanged: + return + case reopened: + if last := r.openCause.Load(); last != nil && *last == msg { + slog.Log(ctx, slog.LevelDebug, msg, args...) + return + } + case opened: } + r.openCause.Store(&msg) + slog.Log(ctx, level, msg, args...) } // refusesWork reports whether an error reply says the server takes no diff --git a/internal/cache/redis_test.go b/internal/cache/redis_test.go index c830873dd..8ea0eda26 100644 --- a/internal/cache/redis_test.go +++ b/internal/cache/redis_test.go @@ -165,7 +165,8 @@ func TestRedis_Record(t *testing.T) { // A closed breaker opening logs one WARN naming why; the operations that // fail while it is open log nothing, and a failed probe reopening it logs at -// DEBUG. Not parallel: it captures the default logger. +// DEBUG unless its cause changed. Not parallel: it captures the default +// logger. func TestRedis_BreakerOpeningLogsOnce(t *testing.T) { buf := logtest.Capture(t, slog.LevelDebug) clock := &fakeClock{t: time.Unix(0, 0)} @@ -173,6 +174,18 @@ func TestRedis_BreakerOpeningLogsOnce(t *testing.T) { live := context.Background() warns := func() int { return strings.Count(buf.String(), `"level":"WARN"`) } debugs := func() int { return strings.Count(buf.String(), `"level":"DEBUG"`) } + errs := func() int { return strings.Count(buf.String(), `"level":"ERROR"`) } + failProbe := func(fail func()) { + t.Helper() + clock.t = clock.t.Add(time.Second) + _, probe := r.breaker.allow() + require.True(t, probe) + fail() + require.True(t, r.breaker.isOpen()) + } + timeout := func() { r.record(live, context.DeadlineExceeded) } + readonly := func() { r.recordReply(live, "READONLY You can't write against a read only replica.") } + wrongpass := func() { r.recordReply(live, "WRONGPASS invalid username-password pair or user is disabled.") } r.recordReply(live, "READONLY You can't write against a read only replica.") r.recordReply(live, "READONLY You can't write against a read only replica.") @@ -194,20 +207,18 @@ func TestRedis_BreakerOpeningLogsOnce(t *testing.T) { assert.Equal(t, 2, warns(), buf.String()) assert.Contains(t, buf.String(), "not answering") - clock.t = clock.t.Add(time.Second) - _, probe = r.breaker.allow() - require.True(t, probe) - r.record(live, context.DeadlineExceeded) - assert.True(t, r.breaker.isOpen(), "a failed probe opens it again") - assert.Equal(t, 2, warns(), "not at WARN") + failProbe(timeout) + assert.Equal(t, 2, warns(), "same cause: not at WARN") assert.Equal(t, 1, debugs(), buf.String()) - clock.t = clock.t.Add(time.Second) - _, probe = r.breaker.allow() - require.True(t, probe) - r.recordReply(live, "READONLY You can't write against a read only replica.") - assert.Equal(t, 2, warns(), "a refused probe is a reopening too") - assert.Equal(t, 2, debugs(), buf.String()) + failProbe(wrongpass) + assert.Equal(t, 1, errs(), "a new cause is logged at its own level") + failProbe(wrongpass) + failProbe(readonly) + failProbe(readonly) + assert.Equal(t, 1, errs(), buf.String()) + assert.Equal(t, 3, warns(), buf.String()) + assert.Equal(t, 3, debugs(), buf.String()) } func TestRejectsCredentials(t *testing.T) { From 71477cb8733fd6a001816db2d88bcc8c0b9f09c1 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 15:01:28 -0400 Subject: [PATCH 77/79] test(e2e): a write pipe inserts on every call over the shared cache Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01FyrXjhR7iDg33paioLQHFq --- tests/e2e/sdk/admin.test.ts | 29 +++++++++++++++++++++++++++++ 1 file changed, 29 insertions(+) diff --git a/tests/e2e/sdk/admin.test.ts b/tests/e2e/sdk/admin.test.ts index ee3dfc8a3..4c36d5ea3 100644 --- a/tests/e2e/sdk/admin.test.ts +++ b/tests/e2e/sdk/admin.test.ts @@ -231,6 +231,35 @@ describe("Admin", () => { } }); + // A write pipe is never answered from the cache (#386): with the shared + // cache this stack runs, a cached [] would skip every repeat's insert. + it("runs a write pipe on every call", async () => { + const writePipe = `test_pipe_write_${Date.now()}`; + const uid = `e2e-write-${Date.now()}`; + await setPipes([ + ...readPipesFile(), + { + name: writePipe, + sql: `INSERT INTO default.${T.users} (user_id, name, email) VALUES ({{uid}}, 'e2e', 'e2e@example.com')`, + parameters: [{ name: "uid", type: "string", required: true }], + description: "E2E test pipe (write)", + allowed_roles: ["admin"], + }, + ]); + + for (let i = 0; i < 2; i++) { + const result = await wh.pipe(writePipe, { uid }); + expect(result.error).toBeNull(); + expect(result.data ?? []).toEqual([]); + } + + const count = await wh.sql( + `SELECT count() AS cnt FROM default.${T.users} WHERE user_id = '${uid}'`, + ); + expect(count.error).toBeNull(); + expect(Number((count.data as { cnt: number | string }[])[0].cnt)).toBe(2); + }); + it("drops a pipe removed from pipes.json", async () => { await setPipes(readPipesFile().filter((p) => p.name !== pipeName)); From 79d550e4f22321ec9d9c0273fec9d8b77a12fe7a Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 15:01:28 -0400 Subject: [PATCH 78/79] docs(deployment): write pipes no longer cached; name #394 instead Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01FyrXjhR7iDg33paioLQHFq --- docs/src/content/docs/deployment.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 1d0c5ff16..7f2e1e907 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -434,7 +434,7 @@ The query-result cache is the layer that can be shared today. With the default ` - **The server is unreachable from the inserting instance.** The invalidation is kept and retried until it lands (`wavehouse_cache_invalidations_pending` counts what is owed). Meanwhile other instances that can still reach the server keep serving the older results, for as long as the outage lasts and at most until each entry's TTL. An instance that stops while invalidations are still owed loses them, with the same bound. The same thing happens today when a process stops between an insert and its invalidation. - **A failover to a replica that had not yet received the latest token writes** can bring back entries filed under the older tokens, bounded by the replication lag at the moment of failover and those entries' TTL. WaveHouse never reads from replicas. Behind a stable address (a managed primary endpoint), an instance still connected to the demoted node has its writes refused, which bypasses its cache; connections are replaced every minute, so it reaches the new primary and delivers the invalidations it owes within about that long. The breaker's own probe write gets a longer budget for a reconnect — twice `dial_timeout` (the client bounds the dial and the handshake by it in turn) plus `timeout` for the write itself — but only closes the breaker when a write actually lands within `timeout`: one slower than that is repeated under `timeout` alone, and the repeat decides, so a server that merely answers slowly stays bypassed instead of flapping open and shut. Every other connection redials under `timeout` alone: size it above how long a reconnect actually takes, or operations that land on one of those keep failing after the probe has already succeeded. - **The server is full and `maxmemory-policy` is `noeviction`.** It refuses the token writes. The inserting instance keeps its invalidations and retries them, bypassing its cache meanwhile, but every other instance serves the results from before the insert until one lands, up to their TTL. -- **A pipe that writes** (an `INSERT` in `pipes.json`) has its result cached like a read, so a repeated identical call is answered from the cache and the write does not run again ([#386](https://github.com/Wave-RF/WaveHouse/issues/386)). With a shared cache that holds on every instance, until the entry's TTL. +- **A pipe that writes** (an `INSERT` in `pipes.json`) is neither cached nor coalesced: it runs on every call, on whichever instance takes it ([Pipes that write](/pipes#pipes-that-write)). It does not invalidate cached reads of the tables it writes ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)), so with a shared cache every instance serves those results from before the write until their TTL. - **Admin writes through `POST /v1/ops/query`** do not invalidate the cache ([#394](https://github.com/Wave-RF/WaveHouse/issues/394)). With a shared cache, the stale results they leave are served by every instance, not only one. **Sizing the server.** Every key WaveHouse writes has a TTL, and a version token lost to eviction, expiry or `FLUSHALL` can only cause misses, never bring back an entry it had invalidated. So set `maxmemory` and let the server evict: `maxmemory-policy allkeys-lru` (or `allkeys-lfu`, `volatile-lru`, `volatile-lfu`). Under `noeviction`, a full server refuses the writes: each refusal (a fill's is counted by `wavehouse_cache_set_failures_total{reason="oom"}`) bypasses the cache of the instance that got it, and invalidations are kept and retried, so the pre-insert results above stay served by the others: avoid `noeviction`. A stored result is capped at `cache.redis.max_value_bytes` (1 MiB compressed). A tenant's version tokens share one hash tag, so each lookup reads them in one `MGET` in cluster mode as well. The results themselves carry no hash tag and spread across shards. **Run it without persistence** (`save ""` and `appendonly no`): stock Redis and Valkey persist by default (periodic RDB save points), so a crash that is followed by a restart reloads the last save on its own — the same rollback as restoring a snapshot by hand, not a loss. Without persistence, a restart can only cause misses, the same as any other token loss. With it on (the default), a restart reloads whatever snapshot or AOF it last wrote, tokens and values it had already invalidated included, so a fresh instance can serve the pre-write rows filed under them as hits until their TTL (up to 1 h) expires. Treat restoring a snapshot, or a crash-restart on a server that still has its defaults, as a rollback, not a resume. From 99dc945416c1b35648f4923a87c7731cef96e827 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 15:09:19 -0400 Subject: [PATCH 79/79] test(e2e): a failed write pipe answers not retryable Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01FyrXjhR7iDg33paioLQHFq --- tests/e2e/sdk/admin.test.ts | 19 +++++++++++++++++++ 1 file changed, 19 insertions(+) diff --git a/tests/e2e/sdk/admin.test.ts b/tests/e2e/sdk/admin.test.ts index 4c36d5ea3..2d9a77c4c 100644 --- a/tests/e2e/sdk/admin.test.ts +++ b/tests/e2e/sdk/admin.test.ts @@ -260,6 +260,25 @@ describe("Admin", () => { expect(Number((count.data as { cnt: number | string }[])[0].cnt)).toBe(2); }); + // A failed write may have run, so it is never answered as retryable. + it("answers a failed write pipe as not retryable", async () => { + const badWrite = `test_pipe_bad_write_${Date.now()}`; + await setPipes([ + ...readPipesFile(), + { + name: badWrite, + sql: `WITH 1 AS x INSERT INTO default.${T.users} (no_such_column) SELECT x`, + description: "E2E test pipe (failing write)", + allowed_roles: ["admin"], + }, + ]); + + const result = await wh.pipe(badWrite); + expect(result.error?.status).toBe(400); + expect(result.error?.code).toBe("clickhouse.rejected"); + expect(result.error?.retryable).toBe(false); + }); + it("drops a pipe removed from pipes.json", async () => { await setPipes(readPipesFile().filter((p) => p.name !== pipeName));