From 3022b926392f9cdfa16931ada447607b94fe09b7 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Thu, 24 Sep 2026 23:27:32 -0400 Subject: [PATCH 01/38] feat(config): choose each layer's implementation at boot mq.backend, cache.backend, dedupe.backend and coord.backend select each layer's implementation; only today's in-process one exists per layer and it is the default. Validate refuses an unknown value, internal/app picks the implementation in one switch per layer, data_dir is probed only when a selected backend keeps state there, and boot logs Config.Warnings. Part of #613. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 1 + cmd/wavehouse/main.go | 19 ++- config.yaml | 10 ++ docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/configuration.mdx | 31 +++- docs/src/content/docs/settings-directory.mdx | 2 +- internal/app/app.go | 4 + internal/app/app_test.go | 27 +++- internal/app/wire.go | 63 ++++++-- internal/config/backends.go | 143 ++++++++++++++++ internal/config/backends_test.go | 161 +++++++++++++++++++ internal/config/config.go | 12 +- internal/config/config_test.go | 10 +- tests/integration/setup_test.go | 5 +- tests/integration/tenants_test.go | 5 +- 15 files changed, 453 insertions(+), 42 deletions(-) create mode 100644 internal/config/backends.go create mode 100644 internal/config/backends_test.go diff --git a/CHANGELOG.md b/CHANGELOG.md index 23c0c7152..9f2e821bc 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,6 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added +- **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block, which until that backend lands is an unknown key and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name every backend: the zero value is not the default, and `app.New` refuses it. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/cmd/wavehouse/main.go b/cmd/wavehouse/main.go index 042257e2b..3a1304f0b 100644 --- a/cmd/wavehouse/main.go +++ b/cmd/wavehouse/main.go @@ -171,14 +171,17 @@ func run(ctx context.Context) int { return 1 } - // data_dir must be writable before anything dials out, so the refusal - // (and, for the typical cause — a bind mount owned by root rather than - // UID 65532 — the remediation) lands at the top of the log rather than - // after ClickHouse discovery. NATS and Pebble still fail loud on their - // own if the directory changes underneath us. - if err := config.CheckDataDir(cfg.DataDir); err != nil { - logger.Error("check data_dir", "error", err) - return 1 + // data_dir, when a selected backend keeps state there, must be writable + // before anything dials out, so the refusal (and, for the typical cause — + // a bind mount owned by root rather than UID 65532 — the remediation) + // lands at the top of the log rather than after ClickHouse discovery. + // NATS and Pebble still fail loud on their own if the directory changes + // underneath us. + if cfg.NeedsDataDir() { + if err := config.CheckDataDir(cfg.DataDir); err != nil { + logger.Error("check data_dir", "error", err) + return 1 + } } a, err := app.New(ctx, app.Options{ diff --git a/config.yaml b/config.yaml index 53a435029..65c384eaa 100644 --- a/config.yaml +++ b/config.yaml @@ -43,9 +43,19 @@ clickhouse: password: "" max_total_conns: 0 # ceiling on open native connections across pools; 0 = none +# Each layer's implementation, chosen at boot. Only the in-process backend +# exists for each today, and it is the default. +mq: + backend: embedded # NATS JetStream under /nats +dedupe: + backend: pebble # Pebble under /pebble +coord: + backend: local + # In-process L1 cache size. The query time-bucket # (query.timestamp_bucket_seconds) is a settings key. cache: + backend: local l1_max_cost: 67108864 # Auth has no on/off switch — the JWT middleware always runs. A request with no diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 6eaf3d54e..938751ff8 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -90,7 +90,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability, ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe` builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ` hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 8a709426f..8d311adc5 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -17,7 +17,7 @@ WaveHouse is configured via a YAML file with environment variable overrides. All 2. Environment variables override any values from the YAML file. 3. If no config file exists, all values are read from environment variables. Every key has a default except `settings.dir` (`WH_SETTINGS_DIR`), which must be set either way. 4. Both sources are **strict**. A YAML key this page doesn't list — a typo, or a tunable that has moved to the settings directory (`dlq.enabled`, `clickhouse.addr`, `stream.*`, a leftover `policy:` or `pipes:` block, …) — refuses to boot and names every offending key, so nothing is read, ignored, and believed. A `WH_*` environment variable that binds to no key on this page (`WH_DEDUPE_ENABLED`, `WH_CH_ADDR`, a misspelling) refuses to boot the same way. Two variables have no YAML key and are exempt because they are not config keys at all but process-level settings `main` reads directly: `WH_CONFIG` (below), which locates the file, and `WH_LOG_LEVEL`. Only the `WH_` prefix is checked, since the environment always carries names that aren't WaveHouse's. One outside source does share the prefix. Kubernetes injects `{SERVICE}_SERVICE_HOST`, `{SERVICE}_PORT`, and similar link variables into every pod in a Service's own namespace, for each Service with a cluster IP that existed before the pod started (a headless Service injects nothing, and a Service in another namespace is harmless). The name is uppercased with `-` mapped to `_`, so a Service named `wh` produces `WH_SERVICE_HOST` and `WH_PORT`, one named `wh-foo` produces `WH_FOO_SERVICE_HOST` and `WH_FOO_PORT`, and either way the pod refuses to boot on its next restart. Set `enableServiceLinks: false` on the pod spec, or name the Service something else. The error says so. -5. Before anything dials out, `data_dir` is probed, and boot refuses on any of these: the value is empty; the path exists but is not a directory; the path, or any component above it, is a dangling symlink (a mount that never came up); the directory exists but the process cannot write to it; the directory is absent and its nearest existing ancestor is not writable, so it could not be created. The probe runs before ClickHouse discovery, so the refusal lands at the top of the log, and a permission denial — on the write probe, or on reaching the path at all through a parent without search permission — carries the UID-65532 remediation, since a bind mount owned by root is the typical cause. +5. Before anything dials out, `data_dir` is probed — when a selected [backend](#backends) keeps state there, as the in-process `mq` and `dedupe` backends do — and boot refuses on any of these: the value is empty; the path exists but is not a directory; the path, or any component above it, is a dangling symlink (a mount that never came up); the directory exists but the process cannot write to it; the directory is absent and its nearest existing ancestor is not writable, so it could not be created. The probe runs before ClickHouse discovery, so the refusal lands at the top of the log, and a permission denial — on the write probe, or on reaching the path at all through a parent without search permission — carries the UID-65532 remediation, since a bind mount owned by root is the typical cause. Boot is the validator for this half of configuration: there is no dry run, and a refused boot with the offending key, variable, or path named in the error is the loud signal. The hot-reloadable half has a dry run — `wavehouse validate` — because it is edited under a running server; boot config only ever takes effect through a restart, so the restart is where it is checked. @@ -37,6 +37,19 @@ This page is boot config only — what the platform operator owns (wiring, lifec | --- | --- | ------- | ----------- | | `data_dir` | `WH_DATA_DIR` | `./data` | Root directory for embedded state. NATS JetStream lives at `/nats`; Pebble, holding every tenant's dedupe store while any tenant has dedupe enabled, at `/pebble`. Subdirectory names are conventions, not config — one knob, one mount. **In a container this MUST resolve to a host-backed volume**; the relative default is for local binary use. WaveHouse logs a startup `WARN` when the directory is missing or empty (no prior state). See [Persistent Storage](/deployment#persistent-storage-required-for-containers). | +### Backends + +Each layer's implementation is chosen once, at boot. Today every layer has one backend, the in-process one, and it is the default, so a config that sets none of these keys runs as it always has. A value this build has no backend for refuses boot and names the valid ones. + +| YAML Key | Env Var | Default | Description | +| --- | --- | ------- | ----------- | +| `mq.backend` | `WH_MQ_BACKEND` | `embedded` | The message queue. `embedded`: NATS JetStream inside this process, under `/nats`. It listens on no port, so no other process can reach its queue. | +| `cache.backend` | `WH_CACHE_BACKEND` | `local` | The query-result cache. `local`: in this process, sized by `cache.l1_max_cost`. | +| `dedupe.backend` | `WH_DEDUPE_BACKEND` | `pebble` | Where ingest dedupe keeps the event ids it has seen. `pebble`: in this process, under `/pebble`, open while any tenant has dedupe on. | +| `coord.backend` | `WH_COORD_BACKEND` | `local` | Where the leases for work only one process may do at a time are held. `local`: in this process. | + +Settings for one backend go in a sub-block named after it, `.`, read only when that backend is selected. A sub-block for a backend this build does not have is an unknown key and refuses boot. `mq` and `dedupe` also appear in the settings directory's `config.json`, with different keys (`mq.max_bytes_gb`, `dedupe.enabled`, …): those are per-tenant tunables and stay there, and one written in `config.yaml` refuses boot as an unknown key. + ### Server | YAML Key | Env Var | Default | Description | @@ -97,6 +110,8 @@ WaveHouse's per-role caps are sent as per-query `SETTINGS` on its connection, so ### Message Queue (NATS) +This section describes the `embedded` [backend](#backends), the only one today. + Each tenant's queue has its own disk budget, `mq.max_bytes_gb`, a hot-reloadable key in the [Settings Directory](/settings-directory#message-queue) — there is no boot-config knob for it. **Durability.** The embedded server runs with JetStream `SyncAlways`, so every event is `fsync`'d to disk before `POST /v1/ingest` returns `200`. This makes your storage's `fsync` latency your ingest latency floor — see [Durability & Storage](/durability) to check whether your substrate can sustain it. There is no knob to relax this today ([#139](https://github.com/Wave-RF/WaveHouse/issues/139) tracks a configurable group-commit interval). @@ -191,9 +206,19 @@ clickhouse: # headers and pool sizes are settings (config.json) max_total_conns: 0 # ceiling on open native connections; 0 = none +mq: + backend: embedded # in-process NATS JetStream under /nats + cache: + backend: local l1_max_cost: 67108864 +dedupe: + backend: pebble # in-process Pebble under /pebble + +coord: + backend: local + auth: jwt_secret: change-me-in-production # jwks_url and role_claim are settings (config.json) operator_key: "" # non-JWT full-access operator credential (Authorization: Operator , or X-Operator-Key); empty disables @@ -239,7 +264,11 @@ WH_SERVER_SHUTDOWN_TIMEOUT=10 WH_CH_PASSWORD= WH_CH_MAX_TOTAL_CONNS=0 +WH_MQ_BACKEND=embedded +WH_CACHE_BACKEND=local WH_CACHE_L1_MAX_COST=67108864 +WH_DEDUPE_BACKEND=pebble +WH_COORD_BACKEND=local WH_AUTH_JWT_SECRET=change-me-in-production WH_AUTH_OPERATOR_KEY= diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index c0e7a19b7..561354627 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -179,7 +179,7 @@ The tenant tunables. Every key is required (a missing one is a validation error) } ``` -What stays in boot config is only what cannot change under a running process — resource sizing (`data_dir`, `cache.l1_max_cost`, `clickhouse.max_total_conns`), the listeners, the observability exporters — and the **secrets**: `clickhouse.password`, `auth.jwt_secret`, `auth.operator_key`. Secrets never belong in a tracked JSON file, so they stay in the environment and are combined with the wiring here on every (re)connect; rotating one is a restart. See [Configuration](/configuration). Everything else lives here and reloads. +What stays in boot config is only what cannot change under a running process — the implementation each layer runs on (`mq.backend`, `cache.backend`, `dedupe.backend`, `coord.backend`), resource sizing (`data_dir`, `cache.l1_max_cost`, `clickhouse.max_total_conns`), the listeners, the observability exporters — and the **secrets**: `clickhouse.password`, `auth.jwt_secret`, `auth.operator_key`. Secrets never belong in a tracked JSON file, so they stay in the environment and are combined with the wiring here on every (re)connect; rotating one is a restart. See [Configuration](/configuration). Everything else lives here and reloads. ## Deduplication diff --git a/internal/app/app.go b/internal/app/app.go index dc15bf5d7..726dad19a 100644 --- a/internal/app/app.go +++ b/internal/app/app.go @@ -170,6 +170,10 @@ func New(ctx context.Context, opts Options) (app *App, err error) { return nil, err } a.wireObservability(ctx) + // After observability, so an OTLP log pipeline carries them too. + for _, w := range a.cfg.Warnings() { + slog.Warn(w) + } if err := a.wireClickHouse(); err != nil { return nil, err } diff --git a/internal/app/app_test.go b/internal/app/app_test.go index 2ecd77d98..10658ddc3 100644 --- a/internal/app/app_test.go +++ b/internal/app/app_test.go @@ -95,7 +95,10 @@ func testConfig(t *testing.T, settingsDir string) *config.Config { return &config.Config{ DataDir: t.TempDir(), Server: config.Server{Port: closedPort(t), ShutdownTimeout: 2}, - Cache: config.Cache{L1MaxCost: 1 << 20}, + MQ: config.MQ{Backend: config.MQEmbedded}, + Cache: config.Cache{Backend: config.CacheLocal, L1MaxCost: 1 << 20}, + Dedupe: config.Dedupe{Backend: config.DedupePebble}, + Coord: config.Coord{Backend: config.CoordLocal}, Auth: config.Auth{JWTSecret: "unit-test-secret"}, Settings: config.Settings{Dir: settingsDir}, } @@ -535,6 +538,28 @@ func TestNew_NestedDedupeStoreFollowsEachTenant(t *testing.T) { assert.False(t, restored.Open(), "Close releases every open store") } +// Validate refuses a backend no layer has a case for, so the switch's default +// is reached only by a Config built by hand; it must refuse boot, not wire +// nothing. +func TestNew_RefusesALayerWithoutABackend(t *testing.T) { + for _, tc := range []struct { + key string + unset func(*config.Config) + }{ + {"dedupe.backend", func(c *config.Config) { c.Dedupe.Backend = "" }}, + {"mq.backend", func(c *config.Config) { c.MQ.Backend = "" }}, + {"cache.backend", func(c *config.Config) { c.Cache.Backend = "" }}, + } { + t.Run(tc.key, func(t *testing.T) { + guardGlobals(t) + cfg := testConfig(t, writeSettings(t, nil)) + tc.unset(cfg) + _, err := New(t.Context(), Options{Config: cfg}) + require.ErrorContains(t, err, tc.key+` "" has no wiring`) + }) + } +} + // A Pebble instance that cannot open follows the registry's own rule for the // shape: a flat directory refuses boot, like every other store, and a nested // one fails closed for every tenant with dedupe on, since they share the diff --git a/internal/app/wire.go b/internal/app/wire.go index 60cbdcd82..d53720253 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -471,7 +471,18 @@ func (a *App) wireDiscovery(ctx context.Context) { a.add(component{name: "schema discovery", close: d.close}) } -// wireDedupe builds the dedupe stores: one per tenant (#583 story 7), each +// wireDedupe builds the dedupe stores — the one place the implementation is +// chosen. +func (a *App) wireDedupe() error { + switch b := a.cfg.Dedupe.Backend; b { + case config.DedupePebble: + return a.wirePebbleDedupe() + default: + return unreachableBackend("dedupe.backend", b) + } +} + +// wirePebbleDedupe builds the dedupe stores: one per tenant (#583 story 7), each // following its own tenant's hot-reloadable dedupe.enabled, over the // embedded Pebble implementation, which is handed data_dir and decides the // rest: every tenant's seen ids in one instance there, open while any @@ -488,7 +499,7 @@ func (a *App) wireDiscovery(ctx context.Context) { // than silently publishing un-deduped, since the files asked for dedupe; // nested fails closed the same way at boot too, for every tenant with // dedupe on, the next reload retrying, so it never costs the process. -func (a *App) wireDedupe() error { +func (a *App) wirePebbleDedupe() error { nested := a.tenants.Nested() embedded := dedupe.NewEmbedded(a.cfg.DataDir) stores := dedupe.NewStores(embedded.Tenant) @@ -530,10 +541,20 @@ func (a *App) wireDedupe() error { return nil } -// wireMQ starts the MQ — the embedded NATS under data_dir/nats, the one -// place the implementation is chosen; everything after it sees mq.Broker — -// and hands it each served tenant's mq.max_bytes_gb, which opens that -// tenant's queue the first time. The budget is hot-reloadable: after every +// wireMQ starts the MQ — the one place the implementation is chosen; +// everything after it sees mq.Broker. +func (a *App) wireMQ() error { + switch b := a.cfg.MQ.Backend; b { + case config.MQEmbedded: + return a.wireEmbeddedMQ() + default: + return unreachableBackend("mq.backend", b) + } +} + +// wireEmbeddedMQ starts the embedded NATS under data_dir/nats and hands it +// each served tenant's mq.max_bytes_gb, which opens that tenant's queue the +// first time. The budget is hot-reloadable: after every // reload the registry applies, each served tenant's is handed over again, // and the MQ owns how it is split across the tenant's queues and keeps them // consistent (see mq.Broker.SetMaxBytes). A tenant no longer served keeps @@ -543,7 +564,7 @@ func (a *App) wireDedupe() error { // previous budget; a nested directory logs it at boot too, so it never costs // the process — the tenant's ingest answers 503 until a reload opens its // queue. The hook is registered before the boot apply, as the dedupe one is. -func (a *App) wireMQ() error { +func (a *App) wireEmbeddedMQ() error { dir := filepath.Join(a.cfg.DataDir, "nats") config.WarnIfFreshDataDir("nats", dir) var broker mq.Broker @@ -591,16 +612,28 @@ func (a *App) wireMQ() error { return nil } -// wireCache opens the L1 cache — the only tier in standalone mode. +// wireCache opens the query-result cache — the one place the implementation +// is chosen. func (a *App) wireCache() error { - l1, err := cache.NewLocal(a.cfg.Cache.L1MaxCost) - if err != nil { - return fmt.Errorf("cache init: %w", err) + switch b := a.cfg.Cache.Backend; b { + case config.CacheLocal: + l1, err := cache.NewLocal(a.cfg.Cache.L1MaxCost) + if err != nil { + return fmt.Errorf("cache init: %w", err) + } + a.cache = l1 + a.add(component{name: "cache", close: withoutContext(l1.Close)}) + return nil + default: + return unreachableBackend("cache.backend", b) } - // TODO: eventually this is where we can switch between ristretto, redis, tiered (both), etc - a.cache = l1 - a.add(component{name: "cache", close: withoutContext(l1.Close)}) - return nil +} + +// unreachableBackend is each layer switch's default case. config.Validate +// refuses a backend with no case, so reaching it means a Config built by hand +// without one (the zero value is not the default), or a case missing here. +func unreachableBackend[T ~string](key string, got T) error { + return fmt.Errorf("%s %q has no wiring: a Config built without config.Load must name every backend", key, got) } // wireSweeper adds the active sweeper — purges messages that are both diff --git a/internal/config/backends.go b/internal/config/backends.go new file mode 100644 index 000000000..8328cea61 --- /dev/null +++ b/internal/config/backends.go @@ -0,0 +1,143 @@ +package config + +import ( + "fmt" + "slices" + "strings" +) + +// Each layer's implementation is chosen here, once, at boot: `.backend` +// names it, and the default is today's in-process one. Settings for one +// backend go in `.`, a sub-block read only when that backend +// is selected. Adding a backend is its constant in the layer's list, a case +// in the layer's validate for its sub-block, and a case in the layer's +// wire function in internal/app — nothing else in Validate changes. + +// MQBackend names the message queue implementation. +type MQBackend string + +// MQEmbedded is the NATS JetStream server inside this process, under +// /nats. +const MQEmbedded MQBackend = "embedded" + +var mqBackends = []MQBackend{MQEmbedded} + +// MQ selects the message queue. The per-tenant byte budget, mq.max_bytes_gb, +// is a settings-directory key, not this block's. +type MQ struct { + Backend MQBackend `yaml:"backend" env:"WH_MQ_BACKEND" env-default:"embedded"` +} + +func (m MQ) validate() error { + return checkBackend("mq.backend", "WH_MQ_BACKEND", m.Backend, mqBackends) +} + +// CacheBackend names the query-result cache implementation. +type CacheBackend string + +// CacheLocal is the in-process Ristretto cache, sized by cache.l1_max_cost. +const CacheLocal CacheBackend = "local" + +var cacheBackends = []CacheBackend{CacheLocal} + +// Cache selects and sizes the query-result cache. The time-range bucket +// structured queries normalize to is a settings-directory key +// (query.timestamp_bucket_seconds) — query shaping, not process memory. +type Cache struct { + Backend CacheBackend `yaml:"backend" env:"WH_CACHE_BACKEND" env-default:"local"` + L1MaxCost int64 `yaml:"l1_max_cost" env:"WH_CACHE_L1_MAX_COST" env-default:"67108864"` +} + +func (c Cache) validate() error { + return checkBackend("cache.backend", "WH_CACHE_BACKEND", c.Backend, cacheBackends) +} + +// DedupeBackend names where ingest dedupe keeps the ids it has seen. +type DedupeBackend string + +// DedupePebble is the Pebble instance inside this process, under +// /pebble, opened while any tenant has dedupe on. +const DedupePebble DedupeBackend = "pebble" + +var dedupeBackends = []DedupeBackend{DedupePebble} + +// Dedupe selects the dedupe store. Whether a tenant dedupes, and on which +// field, are settings-directory keys, not this block's. +type Dedupe struct { + Backend DedupeBackend `yaml:"backend" env:"WH_DEDUPE_BACKEND" env-default:"pebble"` +} + +func (d Dedupe) validate() error { + return checkBackend("dedupe.backend", "WH_DEDUPE_BACKEND", d.Backend, dedupeBackends) +} + +// CoordBackend names the lease implementation singleton work (the sweeper) +// is elected through. +type CoordBackend string + +// CoordLocal holds leases in this process: correct while no other process +// shares its queue. +const CoordLocal CoordBackend = "local" + +var coordBackends = []CoordBackend{CoordLocal} + +// Coord selects the coordination layer. +type Coord struct { + Backend CoordBackend `yaml:"backend" env:"WH_COORD_BACKEND" env-default:"local"` +} + +func (c Coord) validate() error { + return checkBackend("coord.backend", "WH_COORD_BACKEND", c.Backend, coordBackends) +} + +// checkBackend refuses a backend this build has no implementation for, +// listing the ones it has. env repeats the struct tag's literal: a tag can't +// reference a constant. +func checkBackend[T ~string](key, env string, got T, valid []T) error { + if slices.Contains(valid, got) { + return nil + } + names := make([]string, len(valid)) + for i, v := range valid { + names[i] = string(v) + } + return fmt.Errorf("%s (%s) %q is not a backend this build has; valid: %s", key, env, got, strings.Join(names, ", ")) +} + +// validateBackends checks every layer's backend and its sub-block. +func (c *Config) validateBackends() error { + for _, check := range []func() error{c.MQ.validate, c.Cache.validate, c.Dedupe.validate, c.Coord.validate} { + if err := check(); err != nil { + return err + } + } + return nil +} + +// Distributed reports whether the message queue is shared with other +// processes. The embedded one listens on no port, so while it is selected +// every process is an island: nothing else can reach its queue. +func (c *Config) Distributed() bool { return c.MQ.Backend != MQEmbedded } + +// NeedsDataDir reports whether a selected backend keeps state under data_dir, +// and so whether boot must probe it (CheckDataDir). +func (c *Config) NeedsDataDir() bool { + return c.MQ.Backend == MQEmbedded || c.Dedupe.Backend == DedupePebble +} + +// Warnings returns what a valid configuration is still likely to get wrong, +// one line each, for boot to log at WARN. They are not errors because each is +// correct for a single replica, and one process cannot count its replicas. +func (c *Config) Warnings() []string { + if !c.Distributed() { + return nil + } + var out []string + if c.Cache.Backend == CacheLocal { + out = append(out, "cache.backend=local with a shared mq.backend is correct for one replica only: an event ingested on another replica never invalidates this one's cache, so its reads stay stale until the cached entry expires") + } + if c.Dedupe.Backend == DedupePebble { + out = append(out, "dedupe.backend=pebble with a shared mq.backend dedupes per replica only: an id seen by another replica is not seen by this one") + } + return out +} diff --git a/internal/config/backends_test.go b/internal/config/backends_test.go new file mode 100644 index 000000000..0e70dc7a3 --- /dev/null +++ b/internal/config/backends_test.go @@ -0,0 +1,161 @@ +package config + +import ( + "os" + "path/filepath" + "testing" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" +) + +// withDefaultBackends sets what Load's env-defaults would: a literal Config +// names no backend, and Validate refuses that. +func withDefaultBackends(c Config) *Config { + c.MQ.Backend, c.Cache.Backend = MQEmbedded, CacheLocal + c.Dedupe.Backend, c.Coord.Backend = DedupePebble, CoordLocal + return &c +} + +func defaultBackends() Config { + return *withDefaultBackends(Config{Server: Server{Port: 8080}, Settings: Settings{Dir: "./settings"}}) +} + +func TestLoad_BackendDefaults(t *testing.T) { + t.Parallel() + cfg, err := Load("nonexistent.yaml") + require.NoError(t, err) + assert.Equal(t, MQEmbedded, cfg.MQ.Backend) + assert.Equal(t, CacheLocal, cfg.Cache.Backend) + assert.Equal(t, DedupePebble, cfg.Dedupe.Backend) + assert.Equal(t, CoordLocal, cfg.Coord.Backend) + assert.False(t, cfg.Distributed()) + assert.True(t, cfg.NeedsDataDir()) + assert.Empty(t, cfg.Warnings()) +} + +func TestLoad_BackendsFromEnv(t *testing.T) { + t.Setenv("WH_MQ_BACKEND", "embedded") + t.Setenv("WH_CACHE_BACKEND", "local") + t.Setenv("WH_DEDUPE_BACKEND", "pebble") + t.Setenv("WH_COORD_BACKEND", "local") + cfg, err := Load("nonexistent.yaml") + require.NoError(t, err) + assert.Equal(t, MQEmbedded, cfg.MQ.Backend) + assert.Equal(t, CoordLocal, cfg.Coord.Backend) +} + +func TestLoad_BackendFromEnvRefusesAnUnknownValue(t *testing.T) { + t.Setenv("WH_MQ_BACKEND", "nats") + _, err := Load("nonexistent.yaml") + require.Error(t, err) + assert.Contains(t, err.Error(), `mq.backend (WH_MQ_BACKEND) "nats" is not a backend this build has; valid: embedded`) +} + +func TestLoad_BackendsFromYAML(t *testing.T) { + t.Parallel() + path := filepath.Join(t.TempDir(), "config.yaml") + require.NoError(t, os.WriteFile(path, []byte(` +mq: + backend: embedded +cache: + backend: local + l1_max_cost: 1024 +dedupe: + backend: pebble +coord: + backend: local +`), 0o600)) + cfg, err := Load(path) + require.NoError(t, err) + assert.Equal(t, MQEmbedded, cfg.MQ.Backend) + assert.Equal(t, CacheLocal, cfg.Cache.Backend) + assert.Equal(t, int64(1024), cfg.Cache.L1MaxCost) + assert.Equal(t, DedupePebble, cfg.Dedupe.Backend) + assert.Equal(t, CoordLocal, cfg.Coord.Backend) +} + +// A sub-block written before its backend exists, and a settings-directory +// key under a block both files share, are unknown keys — not read and ignored. +func TestLoad_BackendBlocksRefuseUnknownKeys(t *testing.T) { + t.Parallel() + path := filepath.Join(t.TempDir(), "config.yaml") + require.NoError(t, os.WriteFile(path, []byte(` +mq: + backend: embedded + max_bytes_gb: 5 + nats: + urls: nats://localhost:4222 +dedupe: + enabled: true +`), 0o600)) + _, err := Load(path) + require.Error(t, err) + assert.Contains(t, err.Error(), "dedupe.enabled, mq.max_bytes_gb, mq.nats") + assert.Contains(t, err.Error(), EnvSettingsDir) +} + +func TestUnboundEnv_KnowsTheBackendVariables(t *testing.T) { + t.Parallel() + assert.Empty(t, unboundEnv([]string{ + "WH_MQ_BACKEND=embedded", "WH_CACHE_BACKEND=local", + "WH_DEDUPE_BACKEND=pebble", "WH_COORD_BACKEND=local", + })) +} + +func TestValidate_UnknownBackend(t *testing.T) { + t.Parallel() + cases := []struct { + name string + set func(*Config) + want string + }{ + {"mq", func(c *Config) { c.MQ.Backend = "kafka" }, `mq.backend (WH_MQ_BACKEND) "kafka" is not a backend this build has; valid: embedded`}, + {"cache", func(c *Config) { c.Cache.Backend = "redis" }, `cache.backend (WH_CACHE_BACKEND) "redis" is not a backend this build has; valid: local`}, + {"dedupe", func(c *Config) { c.Dedupe.Backend = "dynamodb" }, `dedupe.backend (WH_DEDUPE_BACKEND) "dynamodb" is not a backend this build has; valid: pebble`}, + {"coord", func(c *Config) { c.Coord.Backend = "nats" }, `coord.backend (WH_COORD_BACKEND) "nats" is not a backend this build has; valid: local`}, + // The zero value, which a Config built without Load carries. + {"empty", func(c *Config) { c.MQ.Backend = "" }, `mq.backend (WH_MQ_BACKEND) "" is not a backend`}, + } + for _, tc := range cases { + t.Run(tc.name, func(t *testing.T) { + t.Parallel() + cfg := defaultBackends() + require.NoError(t, cfg.Validate()) + tc.set(&cfg) + err := cfg.Validate() + require.Error(t, err) + assert.Contains(t, err.Error(), tc.want) + }) + } +} + +// Every warning keys on a shared queue, which no backend offers yet, so the +// value is set directly: Warnings reads the choice, it doesn't validate it. +func TestWarnings_SharedQueue(t *testing.T) { + t.Parallel() + cfg := defaultBackends() + assert.Empty(t, cfg.Warnings()) + + cfg.MQ.Backend = "shared" + require.True(t, cfg.Distributed()) + got := cfg.Warnings() + require.Len(t, got, 2) + assert.Contains(t, got[0], "cache.backend=local") + assert.Contains(t, got[1], "dedupe.backend=pebble") + + cfg.Cache.Backend, cfg.Dedupe.Backend = "shared", "shared" + assert.Empty(t, cfg.Warnings()) +} + +func TestNeedsDataDir(t *testing.T) { + t.Parallel() + cfg := defaultBackends() + assert.True(t, cfg.NeedsDataDir()) + cfg.MQ.Backend = "shared" + assert.True(t, cfg.NeedsDataDir(), "pebble dedupe still keeps state under data_dir") + cfg.Dedupe.Backend = "shared" + assert.False(t, cfg.NeedsDataDir()) + cfg.MQ.Backend = MQEmbedded + assert.True(t, cfg.NeedsDataDir(), "the embedded mq keeps state under data_dir") +} diff --git a/internal/config/config.go b/internal/config/config.go index 68b0314b6..cc7566965 100644 --- a/internal/config/config.go +++ b/internal/config/config.go @@ -19,7 +19,10 @@ type Config struct { DataDir string `yaml:"data_dir" env:"WH_DATA_DIR" env-default:"./data"` Server Server `yaml:"server"` ClickHouse ClickHouse `yaml:"clickhouse"` + MQ MQ `yaml:"mq"` Cache Cache `yaml:"cache"` + Dedupe Dedupe `yaml:"dedupe"` + Coord Coord `yaml:"coord"` Auth Auth `yaml:"auth"` OTel OTel `yaml:"otel"` Prometheus Prometheus `yaml:"prometheus"` @@ -132,13 +135,6 @@ type ClickHouse struct { MaxTotalConns int `yaml:"max_total_conns" env:"WH_CH_MAX_TOTAL_CONNS" env-default:"0"` } -// Cache sizes the in-process L1 cache. The time-range bucket structured -// queries normalize to is a settings-directory key -// (query.timestamp_bucket_seconds) — query shaping, not process memory. -type Cache struct { - L1MaxCost int64 `yaml:"l1_max_cost" env:"WH_CACHE_L1_MAX_COST" env-default:"67108864"` -} - // Auth holds the authentication secrets. The verifier wiring — `jwks_url`, // `role_claim` — is the settings directory's `auth` block (hot-reloadable: // a change rebuilds the verifier). There is no on/off switch: the middleware always runs. A request @@ -224,7 +220,7 @@ func (c *Config) Validate() error { } } - return nil + return c.validateBackends() } // Load reads config from a YAML file (if it exists) with env var overrides. diff --git a/internal/config/config_test.go b/internal/config/config_test.go index 822d639d1..ee8b07cc5 100644 --- a/internal/config/config_test.go +++ b/internal/config/config_test.go @@ -203,7 +203,7 @@ func TestValidate_SampleRatesIgnoredWhenObservabilityDisabled(t *testing.T) { Logs: OTelLogs{SampleRate: -1}, }, } - assert.NoError(t, cfg.Validate()) + assert.NoError(t, withDefaultBackends(cfg).Validate()) } func TestValidate_SampleRatesIgnoredWhenSignalDisabled(t *testing.T) { @@ -219,7 +219,7 @@ func TestValidate_SampleRatesIgnoredWhenSignalDisabled(t *testing.T) { Logs: OTelLogs{Enabled: false, SampleRate: -1}, }, } - assert.NoError(t, cfg.Validate()) + assert.NoError(t, withDefaultBackends(cfg).Validate()) } func TestLoad_Defaults_PrometheusDisabled(t *testing.T) { @@ -336,7 +336,7 @@ func TestValidate_PrometheusV1PathAllowedOnSidecarPort(t *testing.T) { Settings: Settings{Dir: "./settings"}, Prometheus: Prometheus{Enabled: true, Path: "/v1/metrics", Port: 9091}, } - assert.NoError(t, cfg.Validate()) + assert.NoError(t, withDefaultBackends(cfg).Validate()) } func TestValidate_PrometheusOnly_NoOTel(t *testing.T) { @@ -348,7 +348,7 @@ func TestValidate_PrometheusOnly_NoOTel(t *testing.T) { Settings: Settings{Dir: "./settings"}, Prometheus: Prometheus{Enabled: true, Path: "/metrics", Port: 0}, } - assert.NoError(t, cfg.Validate()) + assert.NoError(t, withDefaultBackends(cfg).Validate()) } func TestValidate_PrometheusIgnoredWhenDisabled(t *testing.T) { @@ -365,7 +365,7 @@ func TestValidate_PrometheusIgnoredWhenDisabled(t *testing.T) { Port: 8080, }, } - assert.NoError(t, cfg.Validate()) + assert.NoError(t, withDefaultBackends(cfg).Validate()) } // TestEnvSettingsDir_MatchesStructTag pins the exported constant to the diff --git a/tests/integration/setup_test.go b/tests/integration/setup_test.go index a01064a29..ade560f7c 100644 --- a/tests/integration/setup_test.go +++ b/tests/integration/setup_test.go @@ -161,7 +161,10 @@ func setup() (int, func()) { DataDir: dataDir, Server: config.Server{ShutdownTimeout: 10}, ClickHouse: config.ClickHouse{Password: testCHPassword}, - Cache: config.Cache{L1MaxCost: 1 << 30}, // 1 GB + MQ: config.MQ{Backend: config.MQEmbedded}, + Cache: config.Cache{Backend: config.CacheLocal, L1MaxCost: 1 << 30}, // 1 GB + Dedupe: config.Dedupe{Backend: config.DedupePebble}, + Coord: config.Coord{Backend: config.CoordLocal}, Settings: config.Settings{Dir: settingsDir}, } a, err := app.New(ctx, app.Options{Config: cfg, Listener: ln}) diff --git a/tests/integration/tenants_test.go b/tests/integration/tenants_test.go index 40c5a700c..ca00f42f3 100644 --- a/tests/integration/tenants_test.go +++ b/tests/integration/tenants_test.go @@ -62,7 +62,10 @@ func TestNestedDirectory_PerTenantPoolsAndDiscovery(t *testing.T) { Server: config.Server{ShutdownTimeout: 10}, ClickHouse: config.ClickHouse{Password: testCHPassword}, Auth: config.Auth{OperatorKey: operatorKey}, - Cache: config.Cache{L1MaxCost: 1 << 20}, + MQ: config.MQ{Backend: config.MQEmbedded}, + Cache: config.Cache{Backend: config.CacheLocal, L1MaxCost: 1 << 20}, + Dedupe: config.Dedupe{Backend: config.DedupePebble}, + Coord: config.Coord{Backend: config.CoordLocal}, Settings: config.Settings{Dir: root}, } a, err := app.New(ctx, app.Options{Config: cfg, Listener: ln}) From fde17ba4578a2a03521e085a77fd2d3326a75c7d Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Thu, 24 Sep 2026 23:30:43 -0400 Subject: [PATCH 02/38] docs(config): say coord.backend is reserved; sync the boot-config lists Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- config.yaml | 2 +- docs/src/content/docs/architecture.md | 7 ++++--- docs/src/content/docs/configuration.mdx | 4 ++-- internal/app/wire.go | 2 +- internal/config/backends.go | 8 ++++---- internal/settings/settings.go | 8 +++++--- 7 files changed, 18 insertions(+), 15 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 9f2e821bc..6564775a8 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block, which until that backend lands is an unknown key and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name every backend: the zero value is not the default, and `app.New` refuses it. +- **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block, which until that backend lands is an unknown key and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name the `mq`, `cache` and `dedupe` backends: the zero value is not the default, and `app.New` refuses it. `coord.backend` is reserved: nothing reads it until the lease layer lands, and the sweeper still runs in every process. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/config.yaml b/config.yaml index 65c384eaa..5519b78cd 100644 --- a/config.yaml +++ b/config.yaml @@ -50,7 +50,7 @@ mq: dedupe: backend: pebble # Pebble under /pebble coord: - backend: local + backend: local # reserved: nothing is elected yet # In-process L1 cache size. The query time-bucket # (query.timestamp_bucket_seconds) is a settings key. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 938751ff8..425918076 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -89,7 +89,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring -- **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability, ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. +- **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability (after which each `Config.Warnings` line is logged at `WARN`), ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. - **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where they would hold the ack floor and stop the sweeper. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultSetting` reads the store tenant `0` last adopted (`App.defaultStore`, tracked by an `onDefaultAdopt` hook that runs only after a reload that adopted it), so a `0` folder that a reload rejects or removes leaves it as it was. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each served tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot and again after every reload, under the App's stop context, which opens that tenant's queue the first time; a queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out @@ -115,8 +115,9 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ ### `config/` — Configuration -- **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). -- **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load`, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. +- **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — the implementation each layer runs on, resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). +- **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load` when `Config.NeedsDataDir` says a selected backend keeps state there, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. +- **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` at the end of `Validate`), which refuses a value not on the list and names the ones that are. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns the valid combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), which `app.New` logs at `WARN`. - **strict.go** — `rejectUnknownKeys`, the YAML half: re-reads the file as a generic tree and walks it against the struct's `yaml` tags, listing every key the struct doesn't declare. cleanenv itself is lenient by design, which is exactly wrong for boot config once keys have moved to the settings directory. - **persistence.go** — `WarnIfFreshDataDir` logs the startup `WARN` when `data_dir` is missing or empty (on a redeploy, the sign that the volume didn't persist); `LogStorageInitError` attaches the UID-65532 `permissionHint` to a NATS or Pebble open failure that looks like a permission denial — the same hint string `CheckDataDir` uses. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 8d311adc5..ee264bf45 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -46,7 +46,7 @@ Each layer's implementation is chosen once, at boot. Today every layer has one b | `mq.backend` | `WH_MQ_BACKEND` | `embedded` | The message queue. `embedded`: NATS JetStream inside this process, under `/nats`. It listens on no port, so no other process can reach its queue. | | `cache.backend` | `WH_CACHE_BACKEND` | `local` | The query-result cache. `local`: in this process, sized by `cache.l1_max_cost`. | | `dedupe.backend` | `WH_DEDUPE_BACKEND` | `pebble` | Where ingest dedupe keeps the event ids it has seen. `pebble`: in this process, under `/pebble`, open while any tenant has dedupe on. | -| `coord.backend` | `WH_COORD_BACKEND` | `local` | Where the leases for work only one process may do at a time are held. `local`: in this process. | +| `coord.backend` | `WH_COORD_BACKEND` | `local` | Reserved for the leases that will elect work only one process may do at a time, such as the sweeper. Nothing is elected yet: every process runs its own sweeper, and `local`, the only value, changes nothing. | Settings for one backend go in a sub-block named after it, `.`, read only when that backend is selected. A sub-block for a backend this build does not have is an unknown key and refuses boot. `mq` and `dedupe` also appear in the settings directory's `config.json`, with different keys (`mq.max_bytes_gb`, `dedupe.enabled`, …): those are per-tenant tunables and stay there, and one written in `config.yaml` refuses boot as an unknown key. @@ -217,7 +217,7 @@ dedupe: backend: pebble # in-process Pebble under /pebble coord: - backend: local + backend: local # reserved: nothing is elected yet auth: jwt_secret: change-me-in-production # jwks_url and role_claim are settings (config.json) diff --git a/internal/app/wire.go b/internal/app/wire.go index d53720253..3cfcad4b5 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -633,7 +633,7 @@ func (a *App) wireCache() error { // refuses a backend with no case, so reaching it means a Config built by hand // without one (the zero value is not the default), or a case missing here. func unreachableBackend[T ~string](key string, got T) error { - return fmt.Errorf("%s %q has no wiring: a Config built without config.Load must name every backend", key, got) + return fmt.Errorf("%s %q has no wiring: a Config built without config.Load must name the backend of every layer it wires", key, got) } // wireSweeper adds the active sweeper — purges messages that are both diff --git a/internal/config/backends.go b/internal/config/backends.go index 8328cea61..f2ab93308 100644 --- a/internal/config/backends.go +++ b/internal/config/backends.go @@ -71,12 +71,12 @@ func (d Dedupe) validate() error { return checkBackend("dedupe.backend", "WH_DEDUPE_BACKEND", d.Backend, dedupeBackends) } -// CoordBackend names the lease implementation singleton work (the sweeper) -// is elected through. +// CoordBackend names where leases for singleton work (the sweeper) are held. +// Nothing reads it yet: the lease layer (#613) wires it. type CoordBackend string -// CoordLocal holds leases in this process: correct while no other process -// shares its queue. +// CoordLocal holds leases in this process, which is enough while no other +// process shares its queue. const CoordLocal CoordBackend = "local" var coordBackends = []CoordBackend{CoordLocal} diff --git a/internal/settings/settings.go b/internal/settings/settings.go index 55ec089d3..8db981b3a 100644 --- a/internal/settings/settings.go +++ b/internal/settings/settings.go @@ -59,9 +59,11 @@ type PipesFile struct { // TenantConfig is the shape of config.json: the behavioral tunables that // migrate out of boot config. Boot config (config.yaml/env) keeps only what -// cannot change under a running process — resource sizing (`data_dir`, -// `cache.l1_max_cost`), listeners, the observability -// exporters — and the secrets (`clickhouse.password`, `auth.jwt_secret`, +// cannot change under a running process — the implementation each layer +// runs on (`mq.backend`, `cache.backend`, `dedupe.backend`, +// `coord.backend`), resource sizing (`data_dir`, `cache.l1_max_cost`, +// `clickhouse.max_total_conns`), listeners, the observability exporters — +// and the secrets (`clickhouse.password`, `auth.jwt_secret`, // `auth.operator_key`), which never belong in a tracked JSON file. Every // block and every top-level key inside it is REQUIRED: the binary carries no // compiled defaults, so the adopted snapshot is exactly what the files say. From f129d5775f2ba700c4997d37e3999e5b7f062849 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Thu, 24 Sep 2026 23:36:56 -0400 Subject: [PATCH 03/38] docs(config): no backend has a sub-block yet; index backends.go in AGENTS.md Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- AGENTS.md | 2 +- CHANGELOG.md | 2 +- docs/src/content/docs/configuration.mdx | 2 +- 3 files changed, 3 insertions(+), 3 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 165957219..a7e4a28e9 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -34,7 +34,7 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, `...
.` for a namespace — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version that leads every namespace key of one tenant, orphaning every cached query keyed by its tables in one step (a pipe result names no table and keeps its TTL, [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the table-keyed cache (the structured-query results) of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) -- **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability — boot is the validator, there is no dry run +- **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability when a selected backend keeps state there (`NeedsDataDir`); `backends.go` holds each layer's `.backend` (only the in-process value today; `coord.backend` reserved) — boot is the validator, there is no dry run - **`dedupe/`** — `Deduplicator` interface → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) diff --git a/CHANGELOG.md b/CHANGELOG.md index 6564775a8..0aac12595 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block, which until that backend lands is an unknown key and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name the `mq`, `cache` and `dedupe` backends: the zero value is not the default, and `app.New` refuses it. `coord.backend` is reserved: nothing reads it until the lease layer lands, and the sweeper still runs in every process. +- **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block; no backend has settings yet, so every such sub-block is an unknown key for now and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name the `mq`, `cache` and `dedupe` backends: the zero value is not the default, and `app.New` refuses it. `coord.backend` is reserved: nothing reads it until the lease layer lands, and the sweeper still runs in every process. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index ee264bf45..193a6c21b 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -48,7 +48,7 @@ Each layer's implementation is chosen once, at boot. Today every layer has one b | `dedupe.backend` | `WH_DEDUPE_BACKEND` | `pebble` | Where ingest dedupe keeps the event ids it has seen. `pebble`: in this process, under `/pebble`, open while any tenant has dedupe on. | | `coord.backend` | `WH_COORD_BACKEND` | `local` | Reserved for the leases that will elect work only one process may do at a time, such as the sweeper. Nothing is elected yet: every process runs its own sweeper, and `local`, the only value, changes nothing. | -Settings for one backend go in a sub-block named after it, `.`, read only when that backend is selected. A sub-block for a backend this build does not have is an unknown key and refuses boot. `mq` and `dedupe` also appear in the settings directory's `config.json`, with different keys (`mq.max_bytes_gb`, `dedupe.enabled`, …): those are per-tenant tunables and stay there, and one written in `config.yaml` refuses boot as an unknown key. +Settings for one backend will go in a sub-block named after it, `.`, read only when that backend is selected. No backend has settings yet, so today any such sub-block, `mq.embedded` included, is an unknown key and refuses boot. `mq` and `dedupe` also appear in the settings directory's `config.json`, with different keys (`mq.max_bytes_gb`, `dedupe.enabled`, …): those are per-tenant tunables and stay there, and one written in `config.yaml` refuses boot as an unknown key. ### Server From 108499f4158a2820996f39b8e5688c6783c4a4db Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 01:04:55 -0400 Subject: [PATCH 04/38] feat(dedupe): DynamoDB backend, conformance-tested on dynamodb-local One shared table for every tenant, pk (binary) = the dedupe key, no sort key. Reserve is a conditional PutItem per key (ALL_OLD on failure answers Duplicate or InFlight without a read), Commit a BatchWriteItem with unprocessed-item retries, Release a DeleteItem conditional on the token. An item whose ex has passed is absent to Reserve whether or not TTL has deleted it. Throttles, server faults, timeouts and connection errors wrap ErrUnavailable; a breaker short-circuits Reserve after five in a second. Constructible and tested, not yet selectable at boot (F5). Part of #613. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- AGENTS.md | 2 +- CHANGELOG.md | 1 + docs/src/content/docs/architecture.md | 3 +- docs/src/content/docs/deployment.md | 72 +++ go.mod | 16 + go.sum | 32 ++ internal/dedupe/dynamodb.go | 605 ++++++++++++++++++++++ internal/dedupe/dynamodb_bench_test.go | 77 +++ internal/dedupe/dynamodb_test.go | 357 +++++++++++++ tests/integration/dedupe_dynamodb_test.go | 302 +++++++++++ tests/integration/setup_test.go | 34 ++ 11 files changed, 1499 insertions(+), 2 deletions(-) create mode 100644 internal/dedupe/dynamodb.go create mode 100644 internal/dedupe/dynamodb_bench_test.go create mode 100644 internal/dedupe/dynamodb_test.go create mode 100644 tests/integration/dedupe_dynamodb_test.go diff --git a/AGENTS.md b/AGENTS.md index 636156865..84b0a5eee 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -35,7 +35,7 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) - **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability — boot is the validator, there is no dry run -- **`dedupe/`** — `Deduplicator` interface (two-phase `Reserve`/`Commit`/`Release` over `Key{Table, ID}`; every backend passes the `dedupetest` conformance suite) → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant and table, pending claims in memory, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) +- **`dedupe/`** — `Deduplicator` interface (two-phase `Reserve`/`Commit`/`Release` over `Key{Table, ID}`; every backend passes the `dedupetest` conformance suite) → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant and table, pending claims in memory, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch); `Dynamo` (one shared DynamoDB table, conditional `PutItem` claims, built and conformance-tested against dynamodb-local but not yet selectable at boot), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) - **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, token encoding, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`); `internal/app` constructs it and hands everything else a `mq.Broker` diff --git a/CHANGELOG.md b/CHANGELOG.md index 56aa5a3ae..f31bb3c1e 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,6 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added +- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. `Reserve` is a conditional `PutItem` that is atomic across pods. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 740669323..3cee2210e 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -58,7 +58,7 @@ internal/ ├── chconn/ One ClickHouse pool per connection tuple among the served tenants, reconciled on reload under the ceiling ├── chsql/ Shared ClickHouse SQL helpers (identifier quoting, bind-safety) ├── config/ YAML + env var configuration loading -├── dedupe/ Optional deduplication (Reserve/Commit/Release; Pebble) +├── dedupe/ Optional deduplication (Reserve/Commit/Release; Pebble, DynamoDB) ├── discovery/ ClickHouse schema introspection and validation ├── ingest/ Batch buffering, DLQ, and Active Sweeper ├── mq/ MQ boundary: the only NATS/JetStream importer (owned message/consumer/stream types + embedded server) @@ -125,6 +125,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key layout every backend stores: a version byte, the tenant, a NUL, the table, a NUL, the id ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)). No tenant id holds a NUL and settings validation refuses one in a table name, so neither tenants nor tables share ids; an id over 1,024 bytes (or one starting `0xFF`) is stored as `0xFF` plus its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. +- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (binary) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims within a second short-circuit `Reserve` for a second. `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also refuses an unstorable key, reads a lease `<= 0` as `DefaultLease`, and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index bd47e10a8..2f97ad13b 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -421,6 +421,78 @@ WaveHouse discovers this schema on startup and refreshes it every `schema.refres The dedupe key now carries the table as well as the tenant ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)), so **an id deduped before the upgrade is not recognized after it**: a record carrying it is accepted once more. Nothing is migrated, and the old keys stay in `/pebble`, unread; nothing removes them yet ([#220](https://github.com/Wave-RF/WaveHouse/issues/220) tracks the sweep that will). Only a tenant with `dedupe.enabled` on is affected, and only by a record sent both before and after the upgrade — typically a producer retrying across the restart. To avoid duplicate rows, let retrying producers finish, or pause them, before upgrading. +## A shared dedupe table on DynamoDB + +:::note[Not selectable yet] +The DynamoDB dedupe backend is built and tested (`internal/dedupe/dynamodb.go`), but no boot key chooses it yet: every deployment still uses the embedded Pebble store. The `dedupe.backend` boot key lands with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot-config work. This section describes the table that backend expects, so the infrastructure can be ready first. +::: + +Pebble is per process, so two pods on it do not share seen ids. The DynamoDB backend keeps every tenant's ids in **one shared table**, and a conditional write makes a claim atomic across every pod that uses the table. WaveHouse **never creates this table in production**: the table belongs to your infrastructure code. The backend's `create_table` switch is refused unless an `endpoint` override is set, so it only works against [dynamodb-local](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/DynamoDBLocal.html). + +What the backend requires of the table: + +| Attribute | Type | Role | +|---|---|---| +| `pk` | Binary | Partition key, and the only key: tenant, table and id. No sort key. | +| `st` | Number | `1` = pending claim, `2` = committed. | +| `ex` | Number | Epoch seconds: the lease end while pending, the retention end once committed; absent = never expires. | +| `tk` | Binary | The claim token that `Release` matches. | + +Only `pk` is declared in the table definition. Turn TTL on for `ex`. Correctness never depends on TTL, because a claim whose `ex` has passed counts as absent whether or not DynamoDB has deleted it yet. Without TTL, though, expired items are never removed and storage keeps growing. The backend's table check, which boot will run once the backend is selectable, refuses a table whose key schema does not match and logs a warning if TTL is off. + +An example in Terraform. Its tags are the five that Wave RF's own deployments put on every AWS resource (`Name`, `Project`, `Environment`, `ManagedBy`, `CostCenter`, with lowercase-kebab values); use your own conventions in their place: + +```hcl +resource "aws_dynamodb_table" "wavehouse_dedupe" { + name = "wavehouse-dedupe-${var.environment}" + billing_mode = "PAY_PER_REQUEST" # provisioned + auto scaling once traffic is steady + hash_key = "pk" + deletion_protection_enabled = true + + attribute { + name = "pk" + type = "B" + } + + ttl { + attribute_name = "ex" + enabled = true + } + + server_side_encryption { + enabled = true + } + + tags = { + Name = "wavehouse-dedupe-${var.environment}" + Project = "wavehouse-cloud" + Environment = var.environment # prod | dev | ci | demo | benchmark + ManagedBy = "wavehouse-cloud/infra/stacks/prod-platform" + CostCenter = "data-plane" + } +} + +# The pods' role (EKS Pod Identity or IRSA). No Scan, no CreateTable. +data "aws_iam_policy_document" "wavehouse_dedupe" { + statement { + actions = [ + "dynamodb:PutItem", + "dynamodb:DeleteItem", + "dynamodb:BatchWriteItem", + "dynamodb:DescribeTable", + "dynamodb:DescribeTimeToLive", + ] + resources = [aws_dynamodb_table.wavehouse_dedupe.arn] + } +} +``` + +- **Credentials** come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; the environment or a profile locally), never from WaveHouse configuration. +- **Point-in-time recovery** is not needed. The table records which ids have been seen, so losing it produces duplicate rows, not lost events. +- **Cost:** every new event is two writes (the claim, then the commit), and a duplicate is one. On-demand, that is about $1.25 per million new events in us-east-1. Provisioned capacity with auto scaling is cheaper once traffic is steady. +- **One table serves every tenant,** so one tenant's burst can throttle the rest. A throttled or unreachable table fails the ingest request closed rather than publishing un-deduped. After five failed claims within one second, the backend stops calling the table for a second and fails requests immediately (`wavehouse_dedupe_dynamodb_short_circuits_total`). +- **Metrics:** `wavehouse_dedupe_dynamodb_requests_total{op,outcome}`, `wavehouse_dedupe_dynamodb_request_duration_seconds{op}`, `wavehouse_dedupe_dynamodb_unprocessed_items_total`. The table's own CloudWatch metrics `ThrottledRequests`, `SystemErrors` and `ConsumedWriteCapacityUnits` are worth alerting on too. + ## Upgrading across the v2 ingest envelope The NATS envelope changed shape in this release: the row now travels positionally, with `format`, `columns` and `row` replacing `data` — and the queue changed layout with it: boot deletes the earlier build's queue (below), so nothing an older version published reaches the new worker, which could not read it anyway (it carries no `format`, so there is no way to say which value belongs to which column). **Drain first** to keep what the old build had not yet inserted. diff --git a/go.mod b/go.mod index ca418e893..58b72f13b 100644 --- a/go.mod +++ b/go.mod @@ -18,6 +18,11 @@ require ( github.com/ClickHouse/clickhouse-go/v2 v2.48.0 github.com/MicahParks/jwkset v0.11.3 github.com/MicahParks/keyfunc/v3 v3.8.2 + github.com/aws/aws-sdk-go-v2 v1.47.1 + github.com/aws/aws-sdk-go-v2/config v1.33.6 + github.com/aws/aws-sdk-go-v2/credentials v1.20.6 + github.com/aws/aws-sdk-go-v2/service/dynamodb v1.69.1 + github.com/aws/smithy-go v1.28.1 github.com/cockroachdb/pebble v1.1.5 github.com/dgraph-io/ristretto/v2 v2.4.2 github.com/dustin/go-humanize v1.0.1 @@ -72,6 +77,17 @@ require ( github.com/andybalholm/brotli v1.2.2 // indirect github.com/antithesishq/antithesis-sdk-go v0.7.2-default-no-op // indirect github.com/aws/aws-sdk-go v1.49.4 // indirect + github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.20.1 // indirect + github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.4 // indirect + github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.4 // indirect + github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.4 // indirect + github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.19 // indirect + github.com/aws/aws-sdk-go-v2/service/internal/endpoint-discovery v1.13.4 // indirect + github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.4 // indirect + github.com/aws/aws-sdk-go-v2/service/signin v1.10.1 // indirect + github.com/aws/aws-sdk-go-v2/service/sso v1.38.1 // indirect + github.com/aws/aws-sdk-go-v2/service/ssooidc v1.43.1 // indirect + github.com/aws/aws-sdk-go-v2/service/sts v1.51.1 // indirect github.com/aymanbagabas/go-osc52/v2 v2.0.1 // indirect github.com/beorn7/perks v1.0.1 // indirect github.com/bitfield/gotestdox v0.2.2 // indirect diff --git a/go.sum b/go.sum index 71dc27272..04ac3aaf8 100644 --- a/go.sum +++ b/go.sum @@ -45,6 +45,38 @@ github.com/antithesishq/antithesis-sdk-go v0.7.2-default-no-op h1:p2zFsAzvhIpFya github.com/antithesishq/antithesis-sdk-go v0.7.2-default-no-op/go.mod h1:FQyySiasQQM8735Ddel3MRojmy4dA1IqCeyJ5jmPMbI= github.com/aws/aws-sdk-go v1.49.4 h1:qiXsqEeLLhdLgUIyfr5ot+N/dGPWALmtM1SetRmbUlY= github.com/aws/aws-sdk-go v1.49.4/go.mod h1:LF8svs817+Nz+DmiMQKTO3ubZ/6IaTpq3TjupRn3Eqk= +github.com/aws/aws-sdk-go-v2 v1.47.1 h1:uOIZnp4PK3ZhKI0dNrJrhTEsLxbpXHTAJlwoS1pvAtw= +github.com/aws/aws-sdk-go-v2 v1.47.1/go.mod h1:bttEH6JqnUL8LepvDVfdrds/fZ5bCIxzpe3abyUrhDU= +github.com/aws/aws-sdk-go-v2/config v1.33.6 h1:MBjkSTLczek/UgiK+EYPIoRTqE7gP8vtW3OFbFo7Nug= +github.com/aws/aws-sdk-go-v2/config v1.33.6/go.mod h1:grRAFzdAZJrwcbasJRg2MPvIrVjtlfXllHssN6+E1JE= +github.com/aws/aws-sdk-go-v2/credentials v1.20.6 h1:NpAFXCU7NzXNkdGK3zQTtsRJ+3v9tZQV0xcdRw8uBdw= +github.com/aws/aws-sdk-go-v2/credentials v1.20.6/go.mod h1:mcZCoiPnyMvP8VMNbygNX5lLqSlkYJIMPODylQMurOk= +github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.20.1 h1:8gALAAmacnIXh+z6VkdDanv4/IkG5APdg4DZLDTmLog= +github.com/aws/aws-sdk-go-v2/feature/ec2/imds v1.20.1/go.mod h1:Z7IJhJU+poOdJjUR2wpyY21ossQ1XS/R3Lk9Msq5kM4= +github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.4 h1:CLq4+8UHCI+ZZYl/EuJxXovaIVN2xeeT8JV+dsApQ5E= +github.com/aws/aws-sdk-go-v2/internal/configsources v1.5.4/go.mod h1:Wv4q5sAM04xAMkoOedxLx2inVf6K5FdxYp+A61L+q/0= +github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.4 h1:dD4MR81I7YkpEBRk6UP9rocC2QnT3qVuXwzlYTtfGEs= +github.com/aws/aws-sdk-go-v2/internal/endpoints/v2 v2.8.4/go.mod h1:EcXV1kAFd5XwSkDHlj94gnF3q5CkJyYiIJfH8N0VmrE= +github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.4 h1:7Wo47d/xn/7KttCSBd8EGYeZ7ULRFRkUHr6vkZPBzVQ= +github.com/aws/aws-sdk-go-v2/internal/v4a v1.5.4/go.mod h1:tDB2IVC1xC3vX8o+6uRlzhTxP3g1b77CZXFX/oD2FnQ= +github.com/aws/aws-sdk-go-v2/service/dynamodb v1.69.1 h1:bKwiQA6SKqFXBO+1IwP/hTwCU5RlqeitG4gVvSuMN8U= +github.com/aws/aws-sdk-go-v2/service/dynamodb v1.69.1/go.mod h1:Gm+i2GlUsFNlzoBq8VXF44XHbKANn3tV8nYBBp3rN8Q= +github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.19 h1:bAdDl/HkGCcGPoe25ToSHEw23VIxt6CT5fLcg111BKg= +github.com/aws/aws-sdk-go-v2/service/internal/accept-encoding v1.13.19/go.mod h1:KaUzbLxv4CeSxh6ZCl9B4m7CuFenS8kUEaDs+f/DQr4= +github.com/aws/aws-sdk-go-v2/service/internal/endpoint-discovery v1.13.4 h1:6HvmOQ1rBRrZ4qPJSWxd5szPKUsngXCwSw+V3UaJHmw= +github.com/aws/aws-sdk-go-v2/service/internal/endpoint-discovery v1.13.4/go.mod h1:zv2N29aiQUhG2XZNM9zgwCnAyVBdTBbcIpfNAlNmA20= +github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.4 h1:29SvnfGhXjTl8ONxFwbj2rs6lbhiFXD2CgFQmbT/bXY= +github.com/aws/aws-sdk-go-v2/service/internal/presigned-url v1.14.4/go.mod h1:wm04I5DMuNVvZHFe/dHnUxincvNbbK7AiNBbYsQivek= +github.com/aws/aws-sdk-go-v2/service/signin v1.10.1 h1:DzCCWLzcIRQ77F3DEUljud7bEjTgFOIKXP52NmVRyhU= +github.com/aws/aws-sdk-go-v2/service/signin v1.10.1/go.mod h1:xpo/geVldu8payT375WekctUzopG/hBU7miiqItMUlw= +github.com/aws/aws-sdk-go-v2/service/sso v1.38.1 h1:Umtl/0YZhng4xndfW3lKJrYYP7NLEjI6bGXVomwLcs0= +github.com/aws/aws-sdk-go-v2/service/sso v1.38.1/go.mod h1:rRD/dnm7q0HYE/I5TMaPgkWyyUGLcwuxHLABsLnQ3e0= +github.com/aws/aws-sdk-go-v2/service/ssooidc v1.43.1 h1:orIWdNiLgzrhu/11RcPPKO/SBzUUymbUQuZbSPImghg= +github.com/aws/aws-sdk-go-v2/service/ssooidc v1.43.1/go.mod h1:skwM/xsbR/1ReUTesv9BhpJp1VjajR7DWQnuVLwiXsQ= +github.com/aws/aws-sdk-go-v2/service/sts v1.51.1 h1:0HOqZXRvMytH6bFHVIc0oJX07sZjfhz0zXtjs6gdE8s= +github.com/aws/aws-sdk-go-v2/service/sts v1.51.1/go.mod h1:26zA0GhDrLo+yiLI2yXWxqB1PdsShfLikoI7GOEgugM= +github.com/aws/smithy-go v1.28.1 h1:R/nXH00c8qcfCzQVELtRw+eLQWtzv+VAIEFJ1/xxXlQ= +github.com/aws/smithy-go v1.28.1/go.mod h1:YE2RhdIuDbA5E5bTdciG9KrW3+TiEONeUWCqxX9i1Fc= github.com/aymanbagabas/go-osc52/v2 v2.0.1 h1:HwpRHbFMcZLEVr42D4p7XBqjyuxQH5SMiErDT4WkJ2k= github.com/aymanbagabas/go-osc52/v2 v2.0.1/go.mod h1:uYgXzlJ7ZpABp8OJ+exZzJJhRNQ2ASbcXHWsFqH8hp8= github.com/aymanbagabas/go-udiff v0.3.1 h1:LV+qyBQ2pqe0u42ZsUEtPiCaUoqgA9gYRDs3vj1nolY= diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go new file mode 100644 index 000000000..2ce30e6ea --- /dev/null +++ b/internal/dedupe/dynamodb.go @@ -0,0 +1,605 @@ +package dedupe + +import ( + "context" + "crypto/rand" + "errors" + "fmt" + "log/slog" + "strconv" + "sync" + "time" + + "github.com/aws/aws-sdk-go-v2/aws" + "github.com/aws/aws-sdk-go-v2/aws/retry" + "github.com/aws/aws-sdk-go-v2/config" + "github.com/aws/aws-sdk-go-v2/service/dynamodb" + "github.com/aws/aws-sdk-go-v2/service/dynamodb/types" + "github.com/aws/smithy-go" + "go.opentelemetry.io/otel" + "go.opentelemetry.io/otel/attribute" + "go.opentelemetry.io/otel/metric" + "golang.org/x/sync/errgroup" + + "github.com/Wave-RF/WaveHouse/internal/tenant" +) + +// The table's attributes. pk is the key from key.go and the only key +// attribute; ex is the table's TTL attribute. +const ( + attrKey = "pk" + attrState = "st" + attrExpiry = "ex" + attrToken = "tk" + + statePending = "1" + stateCommitted = "2" + + // A claim is live while now < ex; one whose ex has passed is absent + // to Reserve, whether or not TTL has deleted it yet. + condReserve = "attribute_not_exists(pk) OR ex <= :now" + condRelease = "tk = :tk AND st = :pending" + + // batchWriteMax is BatchWriteItem's per-call item limit. + batchWriteMax = 25 + // commitRounds bounds the BatchWriteItem rounds one chunk gets before + // its still-unprocessed items fail the Commit. + commitRounds = 8 + tokenBytes = 16 + + // opReserve is the operation the breaker watches: Release and Commit + // answers say nothing about whether a new Reserve would get through. + opReserve = "put_item" +) + +// DynamoConfig is the DynamoDB backend's wiring. Credentials are never here: +// the SDK's default chain finds them (EKS Pod Identity or IRSA in a pod, the +// environment or a profile locally). +type DynamoConfig struct { + // Table is the shared table every tenant's keys live in. Required. + Table string + // Region overrides the SDK chain's region (AWS_REGION) when set. + Region string + // Endpoint points the client at dynamodb-local. Tests and development + // only; it is also what unlocks CreateTable. + Endpoint string + // Timeout bounds each DynamoDB call, its SDK retries included. + // 0 = 250ms. + Timeout time.Duration + // MaxAttempts is the SDK retryer's attempts per call. 0 = 3. + MaxAttempts int + // RetryMode is "standard" (default) or "adaptive", which also rate-limits + // the client after throttles. + RetryMode string + // ReserveConcurrency bounds the parallel calls one Reserve, Commit or + // Release makes. 0 = 64. + ReserveConcurrency int +} + +func (c DynamoConfig) withDefaults() DynamoConfig { + if c.Timeout <= 0 { + c.Timeout = 250 * time.Millisecond + } + if c.MaxAttempts <= 0 { + c.MaxAttempts = 3 + } + if c.RetryMode == "" { + c.RetryMode = "standard" + } + if c.ReserveConcurrency <= 0 { + c.ReserveConcurrency = 64 + } + return c +} + +// dynamoAPI is the part of *dynamodb.Client the backend calls, so a unit test +// can inject throttles and unprocessed items. +type dynamoAPI interface { + PutItem(context.Context, *dynamodb.PutItemInput, ...func(*dynamodb.Options)) (*dynamodb.PutItemOutput, error) + BatchWriteItem(context.Context, *dynamodb.BatchWriteItemInput, ...func(*dynamodb.Options)) (*dynamodb.BatchWriteItemOutput, error) + DeleteItem(context.Context, *dynamodb.DeleteItemInput, ...func(*dynamodb.Options)) (*dynamodb.DeleteItemOutput, error) + DescribeTable(context.Context, *dynamodb.DescribeTableInput, ...func(*dynamodb.Options)) (*dynamodb.DescribeTableOutput, error) + DescribeTimeToLive(context.Context, *dynamodb.DescribeTimeToLiveInput, ...func(*dynamodb.Options)) (*dynamodb.DescribeTimeToLiveOutput, error) + CreateTable(context.Context, *dynamodb.CreateTableInput, ...func(*dynamodb.Options)) (*dynamodb.CreateTableOutput, error) + UpdateTimeToLive(context.Context, *dynamodb.UpdateTimeToLiveInput, ...func(*dynamodb.Options)) (*dynamodb.UpdateTimeToLiveOutput, error) +} + +// Dynamo is the DynamoDB implementation: every tenant's keys in one shared +// table, so pods sharing the table share seen ids and Reserve's conditional +// write is atomic across all of them. WaveHouse never creates the table in +// production; CreateTable is for dynamodb-local. +type Dynamo struct { + api dynamoAPI + cfg DynamoConfig + now func() time.Time + breaker *breaker + metrics dynamoMetrics + // commitBackoff is the wait before retrying the attempt'th round of + // unprocessed items. + commitBackoff func(attempt int) time.Duration +} + +// NewDynamo builds the backend over a client from the SDK's default config +// chain. extra is appended to the chain's options (a test's static +// credentials, say). It dials nothing: Check does. +func NewDynamo(ctx context.Context, cfg DynamoConfig, extra ...func(*config.LoadOptions) error) (*Dynamo, error) { + if cfg.Table == "" { + return nil, errors.New("dedupe: dynamodb table is required") + } + cfg = cfg.withDefaults() + retryer, err := newRetryer(cfg) + if err != nil { + return nil, err + } + opts := []func(*config.LoadOptions) error{config.WithRetryer(retryer)} + if cfg.Region != "" { + opts = append(opts, config.WithRegion(cfg.Region)) + } + awsCfg, err := config.LoadDefaultConfig(ctx, append(opts, extra...)...) + if err != nil { + return nil, fmt.Errorf("dedupe: aws config: %w", err) + } + client := dynamodb.NewFromConfig(awsCfg, func(o *dynamodb.Options) { + if cfg.Endpoint != "" { + o.BaseEndpoint = aws.String(cfg.Endpoint) + } + }) + return newDynamo(client, cfg), nil +} + +func newRetryer(cfg DynamoConfig) (func() aws.Retryer, error) { + standard := func(o *retry.StandardOptions) { + o.MaxAttempts = cfg.MaxAttempts + o.MaxBackoff = 200 * time.Millisecond + } + switch cfg.RetryMode { + case "standard": + return func() aws.Retryer { return retry.NewStandard(standard) }, nil + case "adaptive": + return func() aws.Retryer { + return retry.NewAdaptiveMode(func(o *retry.AdaptiveModeOptions) { + o.StandardOptions = append(o.StandardOptions, standard) + }) + }, nil + } + return nil, fmt.Errorf("dedupe: dynamodb retry_mode %q: want standard or adaptive", cfg.RetryMode) +} + +func newDynamo(api dynamoAPI, cfg DynamoConfig) *Dynamo { + now := time.Now + return &Dynamo{ + api: api, + cfg: cfg.withDefaults(), + now: now, + breaker: newBreaker(now), + metrics: newDynamoMetrics(), + commitBackoff: func(attempt int) time.Duration { + return min(25*time.Millisecond< 0 { + ex = &types.AttributeValueMemberN{Value: strconv.FormatInt(expiresAt(s.d.now(), retention), 10)} + } + // BatchWriteItem refuses a key twice in one call; a caller merging + // claims from two Reserves could hand one over twice. + seen := make(map[string]bool, len(claims)) + writes := make([]types.WriteRequest, 0, len(claims)) + for _, c := range claims { + pk := AppendKey(nil, s.prefix, c.Key) + if seen[string(pk)] { + continue + } + seen[string(pk)] = true + item := map[string]types.AttributeValue{ + attrKey: &types.AttributeValueMemberB{Value: pk}, + attrState: &types.AttributeValueMemberN{Value: stateCommitted}, + attrToken: &types.AttributeValueMemberB{Value: []byte(c.Token)}, + } + if ex != nil { + item[attrExpiry] = ex + } + writes = append(writes, types.WriteRequest{PutRequest: &types.PutRequest{Item: item}}) + } + g, gctx := errgroup.WithContext(ctx) + g.SetLimit(s.d.cfg.ReserveConcurrency) + for start := 0; start < len(writes); start += batchWriteMax { + chunk := writes[start:min(start+batchWriteMax, len(writes))] + g.Go(func() error { return s.commitChunk(gctx, chunk) }) + } + return g.Wait() +} + +func (s *dynamoStore) commitChunk(ctx context.Context, writes []types.WriteRequest) error { + for attempt := 0; ; attempt++ { + var unprocessed []types.WriteRequest + err := s.d.call(ctx, "batch_write_item", func(ctx context.Context) error { + out, err := s.d.api.BatchWriteItem(ctx, &dynamodb.BatchWriteItemInput{ + RequestItems: map[string][]types.WriteRequest{s.d.cfg.Table: writes}, + }) + if err == nil { + unprocessed = out.UnprocessedItems[s.d.cfg.Table] + } + return err + }) + if err != nil { + return err + } + if len(unprocessed) == 0 { + return nil + } + if attempt+1 >= commitRounds { + return fmt.Errorf("%w: dynamodb batch_write_item: %d items still unprocessed", ErrUnavailable, len(unprocessed)) + } + s.d.metrics.unprocessed.Add(ctx, int64(len(unprocessed))) + writes = unprocessed + select { + case <-ctx.Done(): + return ctx.Err() + case <-time.After(s.d.commitBackoff(attempt)): + } + } +} + +// Release deletes each claim's item only while it is still that claim's +// pending item; a failed condition means the key lapsed, was re-claimed or +// was committed, and is left alone. +func (s *dynamoStore) Release(ctx context.Context, claims []Claim) error { + if len(claims) == 0 { + return nil + } + g, gctx := errgroup.WithContext(ctx) + g.SetLimit(s.d.cfg.ReserveConcurrency) + for _, c := range claims { + g.Go(func() error { + err := s.d.call(gctx, "delete_item", func(ctx context.Context) error { + _, err := s.d.api.DeleteItem(ctx, &dynamodb.DeleteItemInput{ + TableName: &s.d.cfg.Table, + Key: map[string]types.AttributeValue{attrKey: &types.AttributeValueMemberB{Value: AppendKey(nil, s.prefix, c.Key)}}, + ConditionExpression: aws.String(condRelease), + ExpressionAttributeValues: map[string]types.AttributeValue{ + ":tk": &types.AttributeValueMemberB{Value: []byte(c.Token)}, + ":pending": &types.AttributeValueMemberN{Value: statePending}, + }, + }) + return err + }) + var gone *types.ConditionalCheckFailedException + if errors.As(err, &gone) { + return nil + } + return err + }) + } + return g.Wait() +} + +// Close is a no-op: the client is the Dynamo's, shared by every tenant. +func (s *dynamoStore) Close() error { return nil } + +// expiresAt is t+d in epoch seconds rounded up, so a claim or commit never +// ends before it was asked to: TTL attributes are whole seconds. +func expiresAt(t time.Time, d time.Duration) int64 { + end := t.Add(d) + sec := end.Unix() + if end.Nanosecond() > 0 { + sec++ + } + return sec +} + +func newToken() string { + b := make([]byte, tokenBytes) + _, _ = rand.Read(b) // crypto/rand.Read never fails + return string(b) +} + +// classify maps a DynamoDB error onto the contract: a condition failure is +// returned as is for the caller to read, anything retrying later can cure +// wraps ErrUnavailable (503), and the rest — a missing table, denied access, +// a malformed request — is a configuration bug (500). +func classify(op string, err error) error { + if err == nil { + return nil + } + var cond *types.ConditionalCheckFailedException + if errors.As(err, &cond) { + return err + } + if transient(err) { + return fmt.Errorf("%w: dynamodb %s: %w", ErrUnavailable, op, err) + } + return fmt.Errorf("dynamodb %s: %w", op, err) +} + +func transient(err error) bool { + if errors.Is(err, context.DeadlineExceeded) { + return true + } + if (retry.RetryableConnectionError{}).IsErrorRetryable(err) == aws.TrueTernary { + return true + } + var api smithy.APIError + if !errors.As(err, &api) { + return false + } + code := api.ErrorCode() + if _, ok := retry.DefaultThrottleErrorCodes[code]; ok { + return true + } + if _, ok := retry.DefaultRetryableErrorCodes[code]; ok { + return true + } + switch code { + case "InternalServerError", "ServiceUnavailable", "ReplicatedWriteConflictException": + return true + } + return api.ErrorFault() == smithy.FaultServer +} + +// breaker short-circuits Reserve for a second after breakerTrips consecutive +// unavailable answers inside a second, so a throttled or unreachable table +// fails requests fast instead of spending every one's full timeout. +type breaker struct { + mu sync.Mutex + now func() time.Time + fails int + since time.Time + openUntil time.Time +} + +const ( + breakerTrips = 5 + breakerWindow = time.Second + breakerCool = time.Second +) + +var errBreakerOpen = fmt.Errorf("%w: dynamodb is failing; short-circuited", ErrUnavailable) + +func newBreaker(now func() time.Time) *breaker { return &breaker{now: now} } + +func (b *breaker) allow() error { + b.mu.Lock() + defer b.mu.Unlock() + if b.now().Before(b.openUntil) { + return errBreakerOpen + } + return nil +} + +func (b *breaker) record(err error) { + b.mu.Lock() + defer b.mu.Unlock() + if !errors.Is(err, ErrUnavailable) { + b.fails = 0 + return + } + now := b.now() + if b.fails == 0 || now.Sub(b.since) > breakerWindow { + b.fails, b.since = 0, now + } + b.fails++ + if b.fails >= breakerTrips { + b.fails = 0 + b.openUntil = now.Add(breakerCool) + } +} + +type dynamoMetrics struct { + requests metric.Int64Counter + duration metric.Float64Histogram + unprocessed metric.Int64Counter + shorted metric.Int64Counter +} + +func newDynamoMetrics() dynamoMetrics { + meter := otel.Meter("wavehouse-dedupe") + requests, _ := meter.Int64Counter("wavehouse_dedupe_dynamodb_requests_total", + metric.WithDescription("DynamoDB dedupe requests by operation and outcome (ok, condition_failed, unavailable, error)")) + duration, _ := meter.Float64Histogram("wavehouse_dedupe_dynamodb_request_duration_seconds", + metric.WithDescription("DynamoDB dedupe request latency, SDK retries included"), metric.WithUnit("s")) + unprocessed, _ := meter.Int64Counter("wavehouse_dedupe_dynamodb_unprocessed_items_total", + metric.WithDescription("Commit items DynamoDB left unprocessed and the backend retried")) + shorted, _ := meter.Int64Counter("wavehouse_dedupe_dynamodb_short_circuits_total", + metric.WithDescription("Reserves refused without a request while DynamoDB was failing")) + return dynamoMetrics{requests: requests, duration: duration, unprocessed: unprocessed, shorted: shorted} +} + +func (m dynamoMetrics) record(ctx context.Context, op string, took time.Duration, err error) { + outcome := "ok" + var cond *types.ConditionalCheckFailedException + switch { + case err == nil: + case errors.As(err, &cond): + outcome = "condition_failed" + case errors.Is(err, ErrUnavailable): + outcome = "unavailable" + default: + outcome = "error" + } + ctx = context.WithoutCancel(ctx) + m.requests.Add(ctx, 1, metric.WithAttributes(attribute.String("op", op), attribute.String("outcome", outcome))) + m.duration.Record(ctx, took.Seconds(), metric.WithAttributes(attribute.String("op", op))) +} + +func (m dynamoMetrics) shortCircuit(ctx context.Context) { m.shorted.Add(ctx, 1) } diff --git a/internal/dedupe/dynamodb_bench_test.go b/internal/dedupe/dynamodb_bench_test.go new file mode 100644 index 000000000..ad7239edf --- /dev/null +++ b/internal/dedupe/dynamodb_bench_test.go @@ -0,0 +1,77 @@ +//go:build dynamobench + +// Manual latency benchmark for the DynamoDB backend, never run by CI. Point it +// at an existing table (the credentials and region come from the SDK chain): +// +// DEDUPE_BENCH_TABLE=wavehouse-dedupe-dev go test -tags dynamobench \ +// -run '^$' -bench Dynamo -benchtime 2000x ./internal/dedupe/ +// +// DEDUPE_BENCH_ENDPOINT=http://localhost:8000 runs it against dynamodb-local +// instead, creating the table there. +package dedupe + +import ( + "fmt" + "os" + "sync/atomic" + "testing" + "time" +) + +var benchSeq atomic.Uint64 + +func benchDynamo(b *testing.B) *Managed { + b.Helper() + cfg := DynamoConfig{Table: os.Getenv("DEDUPE_BENCH_TABLE"), Endpoint: os.Getenv("DEDUPE_BENCH_ENDPOINT")} + if cfg.Table == "" { + b.Skip("DEDUPE_BENCH_TABLE is not set") + } + if cfg.Endpoint != "" { + cfg.Timeout = 5 * time.Second // dynamodb-local is far slower than the service + } + d, err := NewDynamo(b.Context(), cfg) + if err != nil { + b.Fatal(err) + } + if cfg.Endpoint != "" { + if err := d.CreateTable(b.Context()); err != nil { + b.Fatal(err) + } + } + if err := d.Check(b.Context()); err != nil { + b.Fatal(err) + } + m := d.Tenant("bench") + if err := m.Apply(true); err != nil { + b.Fatal(err) + } + return m +} + +// benchKeys are n ids no run has used, with a short retention so the table +// forgets them. +func benchKeys(n int) []Key { + run := time.Now().UnixNano() + out := make([]Key, n) + for i := range out { + out[i] = Key{Table: "bench", ID: fmt.Sprintf("%d-%d", run, benchSeq.Add(1))} + } + return out +} + +func benchReserveCommit(b *testing.B, window int) { + m := benchDynamo(b) + b.ResetTimer() + for b.Loop() { + claims, err := m.Reserve(b.Context(), benchKeys(window), DefaultLease) + if err != nil { + b.Fatal(err) + } + if err := m.Commit(b.Context(), claims, time.Hour); err != nil { + b.Fatal(err) + } + } +} + +func BenchmarkDynamo_ReserveCommit1(b *testing.B) { benchReserveCommit(b, 1) } +func BenchmarkDynamo_ReserveCommit256(b *testing.B) { benchReserveCommit(b, 256) } diff --git a/internal/dedupe/dynamodb_test.go b/internal/dedupe/dynamodb_test.go new file mode 100644 index 000000000..8648c9382 --- /dev/null +++ b/internal/dedupe/dynamodb_test.go @@ -0,0 +1,357 @@ +package dedupe + +import ( + "context" + "errors" + "fmt" + "net" + "sync" + "sync/atomic" + "testing" + "time" + + "github.com/aws/aws-sdk-go-v2/aws" + "github.com/aws/aws-sdk-go-v2/service/dynamodb" + "github.com/aws/aws-sdk-go-v2/service/dynamodb/types" + "github.com/aws/smithy-go" + smithyhttp "github.com/aws/smithy-go/transport/http" + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" +) + +// fakeDynamo answers each operation through its func, or with success when +// that is nil. The DynamoDB semantics themselves are tested against +// dynamodb-local (tests/integration); this is for the error paths it cannot +// produce. +type fakeDynamo struct { + put func(*dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) + batch func(*dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) + del func(*dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) + describe func() (*dynamodb.DescribeTableOutput, error) + ttl func() (*dynamodb.DescribeTimeToLiveOutput, error) +} + +func (f *fakeDynamo) PutItem(_ context.Context, in *dynamodb.PutItemInput, _ ...func(*dynamodb.Options)) (*dynamodb.PutItemOutput, error) { + if f.put == nil { + return &dynamodb.PutItemOutput{}, nil + } + return f.put(in) +} + +func (f *fakeDynamo) BatchWriteItem(_ context.Context, in *dynamodb.BatchWriteItemInput, _ ...func(*dynamodb.Options)) (*dynamodb.BatchWriteItemOutput, error) { + if f.batch == nil { + return &dynamodb.BatchWriteItemOutput{}, nil + } + return f.batch(in) +} + +func (f *fakeDynamo) DeleteItem(_ context.Context, in *dynamodb.DeleteItemInput, _ ...func(*dynamodb.Options)) (*dynamodb.DeleteItemOutput, error) { + if f.del == nil { + return &dynamodb.DeleteItemOutput{}, nil + } + return f.del(in) +} + +func (f *fakeDynamo) DescribeTable(context.Context, *dynamodb.DescribeTableInput, ...func(*dynamodb.Options)) (*dynamodb.DescribeTableOutput, error) { + return f.describe() +} + +func (f *fakeDynamo) DescribeTimeToLive(context.Context, *dynamodb.DescribeTimeToLiveInput, ...func(*dynamodb.Options)) (*dynamodb.DescribeTimeToLiveOutput, error) { + return f.ttl() +} + +func (f *fakeDynamo) CreateTable(context.Context, *dynamodb.CreateTableInput, ...func(*dynamodb.Options)) (*dynamodb.CreateTableOutput, error) { + return nil, errors.New("not used") +} + +func (f *fakeDynamo) UpdateTimeToLive(context.Context, *dynamodb.UpdateTimeToLiveInput, ...func(*dynamodb.Options)) (*dynamodb.UpdateTimeToLiveOutput, error) { + return nil, errors.New("not used") +} + +func apiErr(code string, fault smithy.ErrorFault) error { + return &smithy.GenericAPIError{Code: code, Message: "injected", Fault: fault} +} + +func openFake(t *testing.T, f *fakeDynamo) (*Dynamo, Deduplicator) { + t.Helper() + d := newDynamo(f, DynamoConfig{Table: "dedupe"}) + d.commitBackoff = func(int) time.Duration { return 0 } + m := d.Tenant("acme") + require.NoError(t, m.Apply(true)) + t.Cleanup(func() { _ = m.Close() }) + return d, m +} + +func keys(ids ...string) []Key { + out := make([]Key, len(ids)) + for i, id := range ids { + out[i] = Key{Table: "events", ID: id} + } + return out +} + +func TestClassify(t *testing.T) { + t.Parallel() + for _, tc := range []struct { + name string + err error + unavailable bool + }{ + {"throttled", apiErr("ThrottlingException", smithy.FaultClient), true}, + {"over provisioned throughput", &types.ProvisionedThroughputExceededException{}, true}, + {"account request limit", apiErr("RequestLimitExceeded", smithy.FaultClient), true}, + {"internal error", &types.InternalServerError{}, true}, + {"unknown server fault", apiErr("Whatever", smithy.FaultServer), true}, + {"request timeout", apiErr("RequestTimeoutException", smithy.FaultClient), true}, + {"multi-region write conflict", &types.ReplicatedWriteConflictException{}, true}, + {"deadline", fmt.Errorf("op: %w", context.DeadlineExceeded), true}, + {"connection refused", &smithyhttp.RequestSendError{Err: &net.OpError{Op: "dial", Err: errors.New("refused")}}, true}, + {"missing table", &types.ResourceNotFoundException{}, false}, + {"access denied", apiErr("AccessDeniedException", smithy.FaultClient), false}, + {"validation", apiErr("ValidationException", smithy.FaultClient), false}, + {"caller went away", context.Canceled, false}, + } { + t.Run(tc.name, func(t *testing.T) { + t.Parallel() + err := classify("put_item", tc.err) + assert.ErrorIs(t, err, tc.err, "the cause stays reachable") + assert.Equal(t, tc.unavailable, errors.Is(err, ErrUnavailable)) + }) + } + assert.NoError(t, classify("put_item", nil)) + ccf := &types.ConditionalCheckFailedException{} + assert.Same(t, error(ccf), classify("put_item", ccf), "a condition failure is an answer, not an error") +} + +func TestDynamo_ReserveReadsTheHeldItem(t *testing.T) { + t.Parallel() + _, m := openFake(t, &fakeDynamo{put: func(in *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { + id := string(in.Item[attrKey].(*types.AttributeValueMemberB).Value) + switch id[len(id)-1] { + case 'd': + return nil, &types.ConditionalCheckFailedException{Item: map[string]types.AttributeValue{attrState: &types.AttributeValueMemberN{Value: stateCommitted}}} + case 'f': + return nil, &types.ConditionalCheckFailedException{Item: map[string]types.AttributeValue{attrState: &types.AttributeValueMemberN{Value: statePending}}} + } + assert.Equal(t, condReserve, aws.ToString(in.ConditionExpression)) + assert.Equal(t, types.ReturnValuesOnConditionCheckFailureAllOld, in.ReturnValuesOnConditionCheckFailure) + return &dynamodb.PutItemOutput{}, nil + }}) + claims, err := m.Reserve(t.Context(), keys("new", "old", "inf"), time.Minute) + require.NoError(t, err) + assert.Equal(t, []Status{Claimed, Duplicate, InFlight}, []Status{claims[0].Status, claims[1].Status, claims[2].Status}) + assert.Len(t, claims[0].Token, tokenBytes) + assert.Empty(t, claims[1].Token) +} + +func TestDynamo_FailedReserveReleasesEveryPutThatMayHaveLanded(t *testing.T) { + t.Parallel() + var mu sync.Mutex + putTokens := map[string]string{} + var released []string + _, m := openFake(t, &fakeDynamo{ + put: func(in *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { + id := string(in.Item[attrKey].(*types.AttributeValueMemberB).Value) + mu.Lock() + putTokens[id] = string(in.Item[attrToken].(*types.AttributeValueMemberB).Value) + mu.Unlock() + switch id[len(id)-3:] { + case "dup": + return nil, &types.ConditionalCheckFailedException{Item: map[string]types.AttributeValue{attrState: &types.AttributeValueMemberN{Value: stateCommitted}}} + case "bad": + return nil, &types.InternalServerError{} + } + return &dynamodb.PutItemOutput{}, nil + }, + del: func(in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { + id := string(in.Key[attrKey].(*types.AttributeValueMemberB).Value) + mu.Lock() + defer mu.Unlock() + assert.Equal(t, putTokens[id], string(in.ExpressionAttributeValues[":tk"].(*types.AttributeValueMemberB).Value), "released by the token it was put with") + released = append(released, id[len(id)-3:]) + return &dynamodb.DeleteItemOutput{}, nil + }, + }) + _, err := m.Reserve(t.Context(), keys("ok1", "dup", "bad", "ok2"), time.Minute) + require.ErrorIs(t, err, ErrUnavailable) + assert.ElementsMatch(t, []string{"ok1", "bad", "ok2"}, released, "the failed put may have landed; the duplicate was never ours") +} + +func TestDynamo_CommitRetriesUnprocessedItems(t *testing.T) { + t.Parallel() + var calls atomic.Int64 + var mu sync.Mutex + written := map[string]int{} + heldBack := map[string]bool{} + _, m := openFake(t, &fakeDynamo{batch: func(in *dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) { + calls.Add(1) + reqs := in.RequestItems["dedupe"] + assert.LessOrEqual(t, len(reqs), batchWriteMax) + // Leave the last item of every call unprocessed once. + mu.Lock() + defer mu.Unlock() + var left []types.WriteRequest + for i, r := range reqs { + pk := string(r.PutRequest.Item[attrKey].(*types.AttributeValueMemberB).Value) + assert.Equal(t, stateCommitted, r.PutRequest.Item[attrState].(*types.AttributeValueMemberN).Value) + assert.Contains(t, r.PutRequest.Item, attrExpiry) + if i == len(reqs)-1 && !heldBack[pk] && len(reqs) > 1 { + heldBack[pk] = true + left = append(left, r) + continue + } + written[pk]++ + } + return &dynamodb.BatchWriteItemOutput{UnprocessedItems: map[string][]types.WriteRequest{"dedupe": left}}, nil + }}) + ids := make([]string, 60) + for i := range ids { + ids[i] = fmt.Sprint(i) + } + claims := make([]Claim, 0, len(ids)+1) + for _, k := range keys(ids...) { + claims = append(claims, Claim{Key: k, Status: Claimed, Token: "t"}) + } + claims = append(claims, claims[0]) + require.NoError(t, m.Commit(t.Context(), claims, time.Hour)) + assert.Len(t, written, 60, "a key handed over twice is written once") + for pk, n := range written { + assert.Equal(t, 1, n, "%q", pk) + } + assert.Equal(t, int64(6), calls.Load(), "3 chunks, each retried once") +} + +func TestDynamo_CommitGivesUpOnItemsThatStayUnprocessed(t *testing.T) { + t.Parallel() + _, m := openFake(t, &fakeDynamo{batch: func(in *dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) { + return &dynamodb.BatchWriteItemOutput{UnprocessedItems: in.RequestItems}, nil + }}) + err := m.Commit(t.Context(), []Claim{{Key: keys("a")[0], Status: Claimed, Token: "t"}}, 0) + require.ErrorIs(t, err, ErrUnavailable) +} + +func TestDynamo_ReleaseTreatsAFailedConditionAsDone(t *testing.T) { + t.Parallel() + _, m := openFake(t, &fakeDynamo{del: func(in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { + assert.Equal(t, condRelease, aws.ToString(in.ConditionExpression)) + id := string(in.Key[attrKey].(*types.AttributeValueMemberB).Value) + if id[len(id)-1] == 'x' { + return nil, &types.ResourceNotFoundException{} + } + return nil, &types.ConditionalCheckFailedException{} + }}) + claim := func(id string) []Claim { return []Claim{{Key: keys(id)[0], Status: Claimed, Token: "t"}} } + require.NoError(t, m.Release(t.Context(), claim("gone"))) + err := m.Release(t.Context(), claim("x")) + require.Error(t, err) + assert.False(t, errors.Is(err, ErrUnavailable)) +} + +func TestDynamo_BreakerShortCircuitsReserve(t *testing.T) { + t.Parallel() + var puts atomic.Int64 + var down atomic.Bool + down.Store(true) + d, m := openFake(t, &fakeDynamo{put: func(*dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { + puts.Add(1) + if down.Load() { + return nil, &types.ProvisionedThroughputExceededException{} + } + return &dynamodb.PutItemOutput{}, nil + }}) + now := time.Unix(1_000_000, 0) + var clock sync.Mutex + d.breaker.now = func() time.Time { clock.Lock(); defer clock.Unlock(); return now } + for range breakerTrips { + _, err := m.Reserve(t.Context(), keys("a"), time.Minute) + require.ErrorIs(t, err, ErrUnavailable) + } + _, err := m.Reserve(t.Context(), keys("a"), time.Minute) + require.ErrorIs(t, err, errBreakerOpen) + assert.Equal(t, int64(breakerTrips), puts.Load(), "the open breaker sent nothing") + + down.Store(false) + clock.Lock() + now = now.Add(breakerCool) + clock.Unlock() + c, err := m.Reserve(t.Context(), keys("a"), time.Minute) + require.NoError(t, err, "it closes after the cool-down") + assert.Equal(t, Claimed, c[0].Status) +} + +func TestBreaker_FailuresSpreadOutDoNotTrip(t *testing.T) { + t.Parallel() + now := time.Unix(1_000_000, 0) + b := newBreaker(func() time.Time { return now }) + fail := fmt.Errorf("%w: x", ErrUnavailable) + for range 3 * breakerTrips { + b.record(fail) + now = now.Add(breakerWindow/(breakerTrips-1) + time.Millisecond) + } + require.NoError(t, b.allow()) + now = now.Add(2 * breakerWindow) + for range breakerTrips - 1 { + b.record(fail) + } + b.record(nil) + b.record(fail) + require.NoError(t, b.allow(), "a success resets the count") +} + +func TestDynamo_Check(t *testing.T) { + t.Parallel() + good := &dynamodb.DescribeTableOutput{Table: &types.TableDescription{ + KeySchema: []types.KeySchemaElement{{AttributeName: aws.String("pk"), KeyType: types.KeyTypeHash}}, + AttributeDefinitions: []types.AttributeDefinition{{AttributeName: aws.String("pk"), AttributeType: types.ScalarAttributeTypeB}}, + }} + ttlOn := &dynamodb.DescribeTimeToLiveOutput{TimeToLiveDescription: &types.TimeToLiveDescription{ + AttributeName: aws.String("ex"), TimeToLiveStatus: types.TimeToLiveStatusEnabled, + }} + check := func(table *dynamodb.DescribeTableOutput, ttl *dynamodb.DescribeTimeToLiveOutput, ttlErr error) error { + f := &fakeDynamo{ + describe: func() (*dynamodb.DescribeTableOutput, error) { return table, nil }, + ttl: func() (*dynamodb.DescribeTimeToLiveOutput, error) { return ttl, ttlErr }, + } + return newDynamo(f, DynamoConfig{Table: "dedupe"}).Check(t.Context()) + } + require.NoError(t, check(good, ttlOn, nil)) + require.NoError(t, check(good, &dynamodb.DescribeTimeToLiveOutput{}, nil), "no TTL is a warning") + require.ErrorIs(t, check(good, nil, &types.InternalServerError{}), ErrUnavailable) + + withRange := &dynamodb.DescribeTableOutput{Table: &types.TableDescription{ + KeySchema: []types.KeySchemaElement{ + {AttributeName: aws.String("pk"), KeyType: types.KeyTypeHash}, + {AttributeName: aws.String("sk"), KeyType: types.KeyTypeRange}, + }, + }} + assert.ErrorContains(t, check(withRange, ttlOn, nil), "key schema") +} + +func TestDynamo_Config(t *testing.T) { + t.Parallel() + c := DynamoConfig{}.withDefaults() + assert.Equal(t, DynamoConfig{Timeout: 250 * time.Millisecond, MaxAttempts: 3, RetryMode: "standard", ReserveConcurrency: 64}, c) + for _, mode := range []string{"standard", "adaptive"} { + r, err := newRetryer(DynamoConfig{RetryMode: mode, MaxAttempts: 4}) + require.NoError(t, err) + assert.Equal(t, 4, r().MaxAttempts()) + } + _, err := newRetryer(DynamoConfig{RetryMode: "legacy"}) + require.Error(t, err) + + _, err = NewDynamo(t.Context(), DynamoConfig{}) + require.ErrorContains(t, err, "table is required") + _, err = NewDynamo(t.Context(), DynamoConfig{Table: "t", RetryMode: "legacy"}) + require.Error(t, err) + d, err := NewDynamo(t.Context(), DynamoConfig{Table: "t", Region: "us-east-1"}) + require.NoError(t, err) + require.ErrorIs(t, d.CreateTable(t.Context()), ErrCreateTableNeedsEndpoint, "never against real AWS") +} + +func TestExpiresAt(t *testing.T) { + t.Parallel() + base := time.Unix(100, 0) + assert.Equal(t, int64(101), expiresAt(base, time.Second)) + assert.Equal(t, int64(102), expiresAt(base, 1500*time.Millisecond), "rounded up: never ends early") + assert.Equal(t, int64(102), expiresAt(base.Add(time.Nanosecond), time.Second)) +} diff --git a/tests/integration/dedupe_dynamodb_test.go b/tests/integration/dedupe_dynamodb_test.go new file mode 100644 index 000000000..192734a97 --- /dev/null +++ b/tests/integration/dedupe_dynamodb_test.go @@ -0,0 +1,302 @@ +//go:build integration + +package tests + +import ( + "context" + "errors" + "fmt" + "io" + "net/http" + "strconv" + "strings" + "sync" + "sync/atomic" + "testing" + "time" + + "github.com/aws/aws-sdk-go-v2/aws" + "github.com/aws/aws-sdk-go-v2/config" + "github.com/aws/aws-sdk-go-v2/credentials" + "github.com/aws/aws-sdk-go-v2/service/dynamodb" + "github.com/aws/aws-sdk-go-v2/service/dynamodb/types" + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + + "github.com/Wave-RF/WaveHouse/internal/dedupe" + "github.com/Wave-RF/WaveHouse/internal/dedupe/dedupetest" +) + +var dynamoTables atomic.Uint64 + +// newDynamoTable names a fresh table on dynamodb-local for one test. +func newDynamoTable() string { + return fmt.Sprintf("dedupe_%d", dynamoTables.Add(1)) +} + +// dynamoClient is one client — one pod's view — over table on +// dynamodb-local, through the production constructor. +func dynamoClient(t *testing.T, table string, cfg dedupe.DynamoConfig, extra ...func(*config.LoadOptions) error) *dedupe.Dynamo { + t.Helper() + cfg.Table, cfg.Endpoint, cfg.Region = table, env(t).dynamoEndpoint, "us-east-1" + if cfg.Timeout == 0 { + // dynamodb-local under a parallel suite is slower than the real thing. + cfg.Timeout = 5 * time.Second + } + opts := append([]func(*config.LoadOptions) error{ + config.WithCredentialsProvider(credentials.NewStaticCredentialsProvider("local", "local", "")), + }, extra...) + d, err := dedupe.NewDynamo(t.Context(), cfg, opts...) + require.NoError(t, err) + return d +} + +// rawDynamo is a plain client, for reading and planting items directly. +func rawDynamo(t *testing.T) *dynamodb.Client { + t.Helper() + return dynamodb.New(dynamodb.Options{ + Region: "us-east-1", + BaseEndpoint: aws.String(env(t).dynamoEndpoint), + Credentials: credentials.NewStaticCredentialsProvider("local", "local", ""), + }) +} + +// faultyHTTP answers matching requests itself instead of sending them. +type faultyHTTP struct { + next *http.Client + fault func(target string) (*http.Response, error, bool) +} + +func (f *faultyHTTP) Do(r *http.Request) (*http.Response, error) { + if resp, err, ok := f.fault(r.Header.Get("X-Amz-Target")); ok { + return resp, err + } + return f.next.Do(r) +} + +// awsError is a DynamoDB JSON error response. +func awsError(status int, code string) *http.Response { + body := fmt.Sprintf(`{"__type":"com.amazonaws.dynamodb.v20120810#%s","message":"injected"}`, code) + return &http.Response{ + StatusCode: status, + Header: http.Header{"Content-Type": {"application/x-amz-json-1.0"}}, + Body: io.NopCloser(strings.NewReader(body)), + } +} + +const putItem = "DynamoDB_20120810.PutItem" + +func TestDedupeDynamo_Conformance(t *testing.T) { + t.Parallel() + dedupetest.Run(t, func(t *testing.T) dedupetest.Harness { + table := newDynamoTable() + // failAfter < 0 is off; otherwise the put after that many fails once. + var failAfter, puts atomic.Int64 + failAfter.Store(-1) + fault := config.WithHTTPClient(&faultyHTTP{next: http.DefaultClient, fault: func(target string) (*http.Response, error, bool) { + if target != putItem || failAfter.Load() < 0 || puts.Add(1) <= failAfter.Load() { + return nil, nil, false + } + failAfter.Store(-1) + return awsError(http.StatusBadRequest, "ValidationException"), nil, true + }}) + d := dynamoClient(t, table, dedupe.DynamoConfig{}, fault) + require.NoError(t, d.CreateTable(t.Context())) + require.NoError(t, d.Check(t.Context())) + return dedupetest.Harness{ + Factory: d.Tenant, + Peer: dynamoClient(t, table, dedupe.DynamoConfig{}).Tenant, + FailNextReserve: func(n int) { + puts.Store(0) + failAfter.Store(int64(n)) + }, + } + }) +} + +// 32 clients — 32 pods — race one id: DynamoDB's condition, not anything in +// process, is what lets exactly one through. +func TestDedupeDynamo_ThirtyTwoClientsOneID(t *testing.T) { + t.Parallel() + table := newDynamoTable() + first := dynamoClient(t, table, dedupe.DynamoConfig{}) + require.NoError(t, first.CreateTable(t.Context())) + const n = 32 + stores := make([]*dedupe.Managed, n) + for i := range stores { + stores[i] = dynamoClient(t, table, dedupe.DynamoConfig{}).Tenant("acme") + require.NoError(t, stores[i].Apply(true)) + } + k := []dedupe.Key{{Table: "events", ID: "e1"}} + race := func() map[dedupe.Status][]dedupe.Claim { + got := make([]dedupe.Claim, n) + start := make(chan struct{}) + var wg sync.WaitGroup + for i, s := range stores { + wg.Go(func() { + <-start + c, err := s.Reserve(context.Background(), k, time.Minute) + if assert.NoError(t, err) { + got[i] = c[0] + } + }) + } + close(start) + wg.Wait() + by := map[dedupe.Status][]dedupe.Claim{} + for _, c := range got { + by[c.Status] = append(by[c.Status], c) + } + return by + } + by := race() + require.Len(t, by[dedupe.Claimed], 1, "exactly one client claims the id") + assert.Len(t, by[dedupe.InFlight], n-1) + require.NoError(t, stores[0].Commit(t.Context(), by[dedupe.Claimed], 0)) + assert.Len(t, race()[dedupe.Duplicate], n, "and every client then sees it committed") +} + +func TestDedupeDynamo_Throttled(t *testing.T) { + t.Parallel() + table := newDynamoTable() + require.NoError(t, dynamoClient(t, table, dedupe.DynamoConfig{}).CreateTable(t.Context())) + for _, code := range []string{"ThrottlingException", "ProvisionedThroughputExceededException", "RequestLimitExceeded"} { + t.Run(code, func(t *testing.T) { + t.Parallel() + var sent atomic.Int64 + d := dynamoClient(t, table, dedupe.DynamoConfig{MaxAttempts: 2}, config.WithHTTPClient(&faultyHTTP{ + next: http.DefaultClient, + fault: func(target string) (*http.Response, error, bool) { + if target != putItem { + return nil, nil, false + } + sent.Add(1) + return awsError(http.StatusBadRequest, code), nil, true + }, + })) + m := d.Tenant("acme") + require.NoError(t, m.Apply(true)) + _, err := m.Reserve(t.Context(), []dedupe.Key{{Table: "events", ID: "e1"}}, time.Minute) + require.ErrorIs(t, err, dedupe.ErrUnavailable, "a throttle is worth retrying: 503") + assert.Equal(t, int64(2), sent.Load(), "the SDK retried it once first") + }) + } +} + +func TestDedupeDynamo_Unreachable(t *testing.T) { + t.Parallel() + d, err := dedupe.NewDynamo(t.Context(), dedupe.DynamoConfig{ + Table: "dedupe", Region: "us-east-1", Endpoint: "http://127.0.0.1:1", MaxAttempts: 1, + }, config.WithCredentialsProvider(credentials.NewStaticCredentialsProvider("local", "local", ""))) + require.NoError(t, err) + m := d.Tenant("acme") + require.NoError(t, m.Apply(true)) + k := []dedupe.Key{{Table: "events", ID: "e1"}} + for range 5 { + _, err = m.Reserve(t.Context(), k, time.Minute) + require.ErrorIs(t, err, dedupe.ErrUnavailable) + } + _, err = m.Reserve(t.Context(), k, time.Minute) + require.ErrorIs(t, err, dedupe.ErrUnavailable) + assert.Contains(t, err.Error(), "short-circuited", "five failures in a second open the breaker") + assert.ErrorIs(t, d.Check(t.Context()), dedupe.ErrUnavailable) +} + +func TestDedupeDynamo_ConfigErrorsAreNotUnavailable(t *testing.T) { + t.Parallel() + d := dynamoClient(t, "no_such_table", dedupe.DynamoConfig{}) + m := d.Tenant("acme") + require.NoError(t, m.Apply(true)) + _, err := m.Reserve(t.Context(), []dedupe.Key{{Table: "events", ID: "e1"}}, time.Minute) + require.Error(t, err) + var missing *types.ResourceNotFoundException + assert.ErrorAs(t, err, &missing) + assert.False(t, errors.Is(err, dedupe.ErrUnavailable), "a missing table is a config bug: 500, not 503") + require.Error(t, d.Check(t.Context())) +} + +func TestDedupeDynamo_Check(t *testing.T) { + t.Parallel() + raw := rawDynamo(t) + table := newDynamoTable() + _, err := raw.CreateTable(t.Context(), &dynamodb.CreateTableInput{ + TableName: aws.String(table), + BillingMode: types.BillingModePayPerRequest, + AttributeDefinitions: []types.AttributeDefinition{{AttributeName: aws.String("pk"), AttributeType: types.ScalarAttributeTypeS}}, + KeySchema: []types.KeySchemaElement{{AttributeName: aws.String("pk"), KeyType: types.KeyTypeHash}}, + }) + require.NoError(t, err) + assert.ErrorContains(t, dynamoClient(t, table, dedupe.DynamoConfig{}).Check(t.Context()), "must be binary") + + fresh := newDynamoTable() + d := dynamoClient(t, fresh, dedupe.DynamoConfig{}) + require.NoError(t, d.CreateTable(t.Context())) + require.NoError(t, d.CreateTable(t.Context()), "an existing table is left alone") + ttl, err := raw.DescribeTimeToLive(t.Context(), &dynamodb.DescribeTimeToLiveInput{TableName: aws.String(fresh)}) + require.NoError(t, err) + assert.Equal(t, "ex", aws.ToString(ttl.TimeToLiveDescription.AttributeName)) + assert.Equal(t, types.TimeToLiveStatusEnabled, ttl.TimeToLiveDescription.TimeToLiveStatus) +} + +// Expiry is the item's ex, in epoch seconds, and never depends on TTL having +// deleted the item. +func TestDedupeDynamo_Expiry(t *testing.T) { + t.Parallel() + raw := rawDynamo(t) + table := newDynamoTable() + d := dynamoClient(t, table, dedupe.DynamoConfig{}) + require.NoError(t, d.CreateTable(t.Context())) + m := d.Tenant("acme") + require.NoError(t, m.Apply(true)) + pk := func(id string) []byte { + return dedupe.AppendKey(nil, dedupe.KeyPrefix("acme"), dedupe.Key{Table: "events", ID: id}) + } + item := func(id string) map[string]types.AttributeValue { + out, err := raw.GetItem(t.Context(), &dynamodb.GetItemInput{ + TableName: aws.String(table), ConsistentRead: aws.Bool(true), + Key: map[string]types.AttributeValue{"pk": &types.AttributeValueMemberB{Value: pk(id)}}, + }) + require.NoError(t, err) + return out.Item + } + num := func(av types.AttributeValue) int64 { + n, err := strconv.ParseInt(av.(*types.AttributeValueMemberN).Value, 10, 64) + require.NoError(t, err) + return n + } + + before := time.Now() + claims, err := m.Reserve(t.Context(), []dedupe.Key{{Table: "events", ID: "kept"}, {Table: "events", ID: "brief"}}, 30*time.Second) + require.NoError(t, err) + pending := item("kept") + assert.Equal(t, "1", pending["st"].(*types.AttributeValueMemberN).Value) + assert.InDelta(t, before.Add(30*time.Second).Unix(), num(pending["ex"]), 2, "a pending item's ex is its lease end") + + require.NoError(t, m.Commit(t.Context(), claims[:1], 0)) + require.NoError(t, m.Commit(t.Context(), claims[1:], time.Hour)) + assert.NotContains(t, item("kept"), "ex", "retention 0 writes no ex, so TTL never takes it") + assert.InDelta(t, before.Add(time.Hour).Unix(), num(item("brief")["ex"]), 2, "a commit's ex is its retention end") + + // TTL deletes lazily; an item whose ex has passed is absent all the same. + for _, st := range []string{"1", "2"} { + _, err = raw.PutItem(t.Context(), &dynamodb.PutItemInput{TableName: aws.String(table), Item: map[string]types.AttributeValue{ + "pk": &types.AttributeValueMemberB{Value: pk("stale-" + st)}, + "st": &types.AttributeValueMemberN{Value: st}, + "ex": &types.AttributeValueMemberN{Value: strconv.FormatInt(time.Now().Add(-time.Minute).Unix(), 10)}, + "tk": &types.AttributeValueMemberB{Value: []byte("old")}, + }}) + require.NoError(t, err) + } + got, err := m.Reserve(t.Context(), []dedupe.Key{{Table: "events", ID: "stale-1"}, {Table: "events", ID: "stale-2"}, {Table: "events", ID: "kept"}}, time.Minute) + require.NoError(t, err) + assert.Equal(t, []dedupe.Status{dedupe.Claimed, dedupe.Claimed, dedupe.Duplicate}, + []dedupe.Status{got[0].Status, got[1].Status, got[2].Status}) +} + +func TestDedupeDynamo_CreateTableNeedsEndpoint(t *testing.T) { + t.Parallel() + d, err := dedupe.NewDynamo(t.Context(), dedupe.DynamoConfig{Table: "dedupe", Region: "us-east-1"}, + config.WithCredentialsProvider(credentials.NewStaticCredentialsProvider("local", "local", ""))) + require.NoError(t, err) + require.ErrorIs(t, d.CreateTable(t.Context()), dedupe.ErrCreateTableNeedsEndpoint) +} diff --git a/tests/integration/setup_test.go b/tests/integration/setup_test.go index a01064a29..9d1666908 100644 --- a/tests/integration/setup_test.go +++ b/tests/integration/setup_test.go @@ -51,6 +51,9 @@ type testEnv struct { embeddedMQ mq.Broker baseURL string // the wired API server, e.g. http://127.0.0.1:41234 registry *discovery.SchemaRegistry + // dynamoEndpoint is dynamodb-local, for the DynamoDB dedupe backend's + // tests; the wired app does not use it. + dynamoEndpoint string } var sharedEnv *testEnv @@ -137,6 +140,13 @@ func setup() (int, func()) { _ = ch.container.Terminate(context.Background()) }) + ddb, endpoint, err := startDynamoDBLocal(ctx) + if err != nil { + fmt.Fprintf(os.Stderr, "integration setup: dynamodb-local: %v\n", err) + return 1, cleanup + } + cleanups.push(func() { _ = ddb.Terminate(context.Background()) }) + settingsDir, err := writeTestSettings(ch) if err != nil { fmt.Fprintf(os.Stderr, "integration setup: settings: %v\n", err) @@ -199,6 +209,8 @@ func setup() (int, func()) { embeddedMQ: a.MQ(), baseURL: baseURL, registry: a.Registry(), + + dynamoEndpoint: endpoint, } return 0, cleanup } @@ -383,6 +395,28 @@ func startClickHouse(ctx context.Context) (*chInstance, error) { return ch, nil } +// startDynamoDBLocal starts dynamodb-local in memory (no volume) and +// returns it with its endpoint URL. +func startDynamoDBLocal(ctx context.Context) (testcontainers.Container, string, error) { + container, err := testcontainers.GenericContainer(ctx, testcontainers.GenericContainerRequest{ + ContainerRequest: testcontainers.ContainerRequest{ + Image: "amazon/dynamodb-local:3.3.1", + Cmd: []string{"-jar", "DynamoDBLocal.jar", "-inMemory"}, + ExposedPorts: []string{"8000/tcp"}, + WaitingFor: wait.ForListeningPort("8000/tcp").WithStartupTimeout(60 * time.Second), + }, + Started: true, + }) + if err != nil { + return nil, "", fmt.Errorf("start container: %w", err) + } + endpoint, err := container.PortEndpoint(ctx, "8000/tcp", "http") + if err != nil { + return container, "", fmt.Errorf("endpoint: %w", err) + } + return container, endpoint, nil +} + func waitForNativeReady(ctx context.Context, conn driver.Conn, timeout time.Duration) error { pingCtx, cancel := context.WithTimeout(ctx, timeout) defer cancel() From 6f7944aa8ac566c0dd15e2069af42324abe8381e Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 01:53:16 -0400 Subject: [PATCH 05/38] fix(dedupe): attempt every Commit and Release chunk; keep the breaker honest Review round 1. Commit and Release no longer cancel their siblings on the first failure: the records are already published, and a claim left behind holds its id for a lease. A failed Reserve sends no put after the first failure and releases only puts it sent; a sibling cancelled by that failure no longer resets the breaker (and is counted as outcome "canceled"). Docs: TTL reclaims only lapsed claims until retention lands (#220), no future boot-key names, the breaker counts consecutive unavailable claims. The e2e coverage suite excludes dynamodb.go, which the e2e binary never runs; unit and integration cover it and the merged total counts it. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- .testcoverage.yml | 5 ++ AGENTS.md | 2 +- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/deployment.md | 10 +-- internal/dedupe/dynamodb.go | 86 ++++++++++++--------- internal/dedupe/dynamodb_test.go | 103 ++++++++++++++++++++++++-- 7 files changed, 162 insertions(+), 48 deletions(-) diff --git a/.testcoverage.yml b/.testcoverage.yml index aff1a694d..270520828 100644 --- a/.testcoverage.yml +++ b/.testcoverage.yml @@ -73,3 +73,8 @@ exclude: - ^internal/settings/ - ^cmd/wavehouse/validate\.go$ - ^cmd/wavehouse/bootstrap\.go$ + # The DynamoDB dedupe backend: the e2e binary runs Pebble dedupe, so + # this file measured 0% there and pulled e2e to 58.6%. The unit + # (fake API) and integration (dynamodb-local) suites cover it, and the + # merged total still counts it. + - ^internal/dedupe/dynamodb\.go$ diff --git a/AGENTS.md b/AGENTS.md index 84b0a5eee..29e5e72c3 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -35,7 +35,7 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) - **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability — boot is the validator, there is no dry run -- **`dedupe/`** — `Deduplicator` interface (two-phase `Reserve`/`Commit`/`Release` over `Key{Table, ID}`; every backend passes the `dedupetest` conformance suite) → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant and table, pending claims in memory, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch); `Dynamo` (one shared DynamoDB table, conditional `PutItem` claims, built and conformance-tested against dynamodb-local but not yet selectable at boot), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) +- **`dedupe/`** — `Deduplicator` interface (two-phase `Reserve`/`Commit`/`Release` over `Key{Table, ID}`; every backend passes the `dedupetest` conformance suite) → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant and table, pending claims in memory, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges) or `Dynamo` (one shared DynamoDB table, conditional `PutItem` claims; built and conformance-tested against dynamodb-local but not yet selectable at boot), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) - **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, token encoding, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`); `internal/app` constructs it and hands everything else a `mq.Broker` diff --git a/CHANGELOG.md b/CHANGELOG.md index f31bb3c1e..7e11d26d4 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. `Reserve` is a conditional `PutItem` that is atomic across pods. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. +- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. `Reserve` is a conditional `PutItem` that is atomic across pods. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 3cee2210e..c4d7fcf89 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -125,7 +125,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key layout every backend stores: a version byte, the tenant, a NUL, the table, a NUL, the id ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)). No tenant id holds a NUL and settings validation refuses one in a table name, so neither tenants nor tables share ids; an id over 1,024 bytes (or one starting `0xFF`) is stored as `0xFF` plus its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. -- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (binary) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims within a second short-circuit `Reserve` for a second. `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. +- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (binary) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also refuses an unstorable key, reads a lease `<= 0` as `DefaultLease`, and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 2f97ad13b..9bea802e7 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -424,10 +424,10 @@ The dedupe key now carries the table as well as the tenant ([#222](https://githu ## A shared dedupe table on DynamoDB :::note[Not selectable yet] -The DynamoDB dedupe backend is built and tested (`internal/dedupe/dynamodb.go`), but no boot key chooses it yet: every deployment still uses the embedded Pebble store. The `dedupe.backend` boot key lands with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot-config work. This section describes the table that backend expects, so the infrastructure can be ready first. +The DynamoDB dedupe backend is built and tested (`internal/dedupe/dynamodb.go`), but no boot key chooses it yet: every deployment still uses the embedded Pebble store. A boot key to select it lands with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot-config work. This section describes the table that backend expects, so the infrastructure can be ready first. ::: -Pebble is per process, so two pods on it do not share seen ids. The DynamoDB backend keeps every tenant's ids in **one shared table**, and a conditional write makes a claim atomic across every pod that uses the table. WaveHouse **never creates this table in production**: the table belongs to your infrastructure code. The backend's `create_table` switch is refused unless an `endpoint` override is set, so it only works against [dynamodb-local](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/DynamoDBLocal.html). +Pebble is per process, so two pods on it do not share seen ids. The DynamoDB backend keeps every tenant's ids in **one shared table**, and a conditional write makes a claim atomic across every pod that uses the table. WaveHouse **never creates this table in production**: the table belongs to your infrastructure code. The backend refuses to create a table unless it is pointed at a custom endpoint, so table creation only works against [dynamodb-local](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/DynamoDBLocal.html). What the backend requires of the table: @@ -438,7 +438,7 @@ What the backend requires of the table: | `ex` | Number | Epoch seconds: the lease end while pending, the retention end once committed; absent = never expires. | | `tk` | Binary | The claim token that `Release` matches. | -Only `pk` is declared in the table definition. Turn TTL on for `ex`. Correctness never depends on TTL, because a claim whose `ex` has passed counts as absent whether or not DynamoDB has deleted it yet. Without TTL, though, expired items are never removed and storage keeps growing. The backend's table check, which boot will run once the backend is selectable, refuses a table whose key schema does not match and logs a warning if TTL is off. +Only `pk` is declared in the table definition. Turn TTL on for `ex`. Correctness never depends on TTL, because a claim whose `ex` has passed counts as absent whether or not DynamoDB has deleted it yet; TTL only reclaims the storage. **Today TTL removes only lapsed claims:** ingest commits every id with no retention, so a committed item carries no `ex` and is kept forever, and the table grows by one item (about 200 bytes) per distinct id. Per-tenant retention is [#220](https://github.com/Wave-RF/WaveHouse/issues/220). The backend's table check, which boot will run once the backend is selectable, refuses a table whose key schema does not match and logs a warning if TTL is off. An example in Terraform. Its tags are the five that Wave RF's own deployments put on every AWS resource (`Name`, `Project`, `Environment`, `ManagedBy`, `CostCenter`, with lowercase-kebab values); use your own conventions in their place: @@ -489,8 +489,8 @@ data "aws_iam_policy_document" "wavehouse_dedupe" { - **Credentials** come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; the environment or a profile locally), never from WaveHouse configuration. - **Point-in-time recovery** is not needed. The table records which ids have been seen, so losing it produces duplicate rows, not lost events. -- **Cost:** every new event is two writes (the claim, then the commit), and a duplicate is one. On-demand, that is about $1.25 per million new events in us-east-1. Provisioned capacity with auto scaling is cheaper once traffic is steady. -- **One table serves every tenant,** so one tenant's burst can throttle the rest. A throttled or unreachable table fails the ingest request closed rather than publishing un-deduped. After five failed claims within one second, the backend stops calling the table for a second and fails requests immediately (`wavehouse_dedupe_dynamodb_short_circuits_total`). +- **Cost:** every new event is two writes (the claim, then the commit), and a duplicate is one. On-demand, that is about $1.25 per million new events in us-east-1. Provisioned capacity with auto scaling is cheaper once traffic is steady. Storage is the other line: every distinct id stays in the table (see TTL above), at DynamoDB's per-GB-month rate. +- **One table serves every tenant,** so one tenant's burst can throttle the rest. A throttled or unreachable table fails the ingest request closed rather than publishing un-deduped. After five throttled or unreachable claims in a row within one second, the backend stops calling the table for a second and fails every tenant's dedupe requests immediately (`wavehouse_dedupe_dynamodb_short_circuits_total`). A duplicate or in-flight answer is not a failure and resets the count. - **Metrics:** `wavehouse_dedupe_dynamodb_requests_total{op,outcome}`, `wavehouse_dedupe_dynamodb_request_duration_seconds{op}`, `wavehouse_dedupe_dynamodb_unprocessed_items_total`. The table's own CloudWatch metrics `ThrottledRequests`, `SystemErrors` and `ConsumedWriteCapacityUnits` are worth alerting on too. ## Upgrading across the v2 ingest envelope diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index 2ce30e6ea..08fd0c664 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -260,7 +260,9 @@ func (d *Dynamo) call(ctx context.Context, op string, do func(context.Context) e start := time.Now() err := classify(op, do(ctx)) d.metrics.record(ctx, op, time.Since(start), err) - if op == opReserve { + // A request cancelled because a sibling failed says nothing about the + // table, and must not reset the breaker's count. + if op == opReserve && !errors.Is(err, context.Canceled) { d.breaker.record(err) } return err @@ -289,12 +291,19 @@ func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Durati exp := expiresAt(now, lease) claims := make([]Claim, len(keys)) tried := make([]Claim, len(keys)) + sent := make([]bool, len(keys)) + // The first failure cancels the puts not yet sent: the Reserve fails + // either way, and a throttled table should not take the rest. g, gctx := errgroup.WithContext(ctx) g.SetLimit(s.d.cfg.ReserveConcurrency) for i, k := range keys { token := newToken() tried[i] = Claim{Key: k, Status: Claimed, Token: token} g.Go(func() error { + if err := gctx.Err(); err != nil { + return err + } + sent[i] = true status, err := s.reserve(gctx, k, token, nowSec, exp) if err != nil { return err @@ -307,11 +316,11 @@ func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Durati }) } if err := g.Wait(); err != nil { - // A put that errored or never answered may still have landed; its - // token is known, and releasing a key it does not hold is a no-op. + // A put that was sent and errored may still have landed; its token + // is known, and releasing a key it does not hold is a no-op. var undo []Claim for i, c := range claims { - if c.Status == Claimed || c.Status == 0 { + if sent[i] && (c.Status == Claimed || c.Status == 0) { undo = append(undo, tried[i]) } } @@ -381,13 +390,12 @@ func (s *dynamoStore) Commit(ctx context.Context, claims []Claim, retention time } writes = append(writes, types.WriteRequest{PutRequest: &types.PutRequest{Item: item}}) } - g, gctx := errgroup.WithContext(ctx) - g.SetLimit(s.d.cfg.ReserveConcurrency) - for start := 0; start < len(writes); start += batchWriteMax { - chunk := writes[start:min(start+batchWriteMax, len(writes))] - g.Go(func() error { return s.commitChunk(gctx, chunk) }) - } - return g.Wait() + // Every chunk is attempted whatever another's fate: these records are + // already published, and an uncommitted id lets a retry publish again. + chunks := (len(writes) + batchWriteMax - 1) / batchWriteMax + return forEach(chunks, s.d.cfg.ReserveConcurrency, func(i int) error { + return s.commitChunk(ctx, writes[i*batchWriteMax:min((i+1)*batchWriteMax, len(writes))]) + }) } func (s *dynamoStore) commitChunk(ctx context.Context, writes []types.WriteRequest) error { @@ -428,30 +436,27 @@ func (s *dynamoStore) Release(ctx context.Context, claims []Claim) error { if len(claims) == 0 { return nil } - g, gctx := errgroup.WithContext(ctx) - g.SetLimit(s.d.cfg.ReserveConcurrency) - for _, c := range claims { - g.Go(func() error { - err := s.d.call(gctx, "delete_item", func(ctx context.Context) error { - _, err := s.d.api.DeleteItem(ctx, &dynamodb.DeleteItemInput{ - TableName: &s.d.cfg.Table, - Key: map[string]types.AttributeValue{attrKey: &types.AttributeValueMemberB{Value: AppendKey(nil, s.prefix, c.Key)}}, - ConditionExpression: aws.String(condRelease), - ExpressionAttributeValues: map[string]types.AttributeValue{ - ":tk": &types.AttributeValueMemberB{Value: []byte(c.Token)}, - ":pending": &types.AttributeValueMemberN{Value: statePending}, - }, - }) - return err + // Every claim is attempted: one left behind holds its id for a lease. + return forEach(len(claims), s.d.cfg.ReserveConcurrency, func(i int) error { + c := claims[i] + err := s.d.call(ctx, "delete_item", func(ctx context.Context) error { + _, err := s.d.api.DeleteItem(ctx, &dynamodb.DeleteItemInput{ + TableName: &s.d.cfg.Table, + Key: map[string]types.AttributeValue{attrKey: &types.AttributeValueMemberB{Value: AppendKey(nil, s.prefix, c.Key)}}, + ConditionExpression: aws.String(condRelease), + ExpressionAttributeValues: map[string]types.AttributeValue{ + ":tk": &types.AttributeValueMemberB{Value: []byte(c.Token)}, + ":pending": &types.AttributeValueMemberN{Value: statePending}, + }, }) - var gone *types.ConditionalCheckFailedException - if errors.As(err, &gone) { - return nil - } return err }) - } - return g.Wait() + var gone *types.ConditionalCheckFailedException + if errors.As(err, &gone) { + return nil + } + return err + }) } // Close is a no-op: the client is the Dynamo's, shared by every tenant. @@ -459,6 +464,19 @@ func (s *dynamoStore) Close() error { return nil } // expiresAt is t+d in epoch seconds rounded up, so a claim or commit never // ends before it was asked to: TTL attributes are whole seconds. +// forEach runs do for every index, at most limit at once, and joins the +// errors: one failure never stops the rest. +func forEach(n, limit int, do func(i int) error) error { + errs := make([]error, n) + var g errgroup.Group + g.SetLimit(limit) + for i := range n { + g.Go(func() error { errs[i] = do(i); return nil }) + } + _ = g.Wait() + return errors.Join(errs...) +} + func expiresAt(t time.Time, d time.Duration) int64 { end := t.Add(d) sec := end.Unix() @@ -575,7 +593,7 @@ type dynamoMetrics struct { func newDynamoMetrics() dynamoMetrics { meter := otel.Meter("wavehouse-dedupe") requests, _ := meter.Int64Counter("wavehouse_dedupe_dynamodb_requests_total", - metric.WithDescription("DynamoDB dedupe requests by operation and outcome (ok, condition_failed, unavailable, error)")) + metric.WithDescription("DynamoDB dedupe requests by operation and outcome (ok, condition_failed, unavailable, canceled, error)")) duration, _ := meter.Float64Histogram("wavehouse_dedupe_dynamodb_request_duration_seconds", metric.WithDescription("DynamoDB dedupe request latency, SDK retries included"), metric.WithUnit("s")) unprocessed, _ := meter.Int64Counter("wavehouse_dedupe_dynamodb_unprocessed_items_total", @@ -592,6 +610,8 @@ func (m dynamoMetrics) record(ctx context.Context, op string, took time.Duration case err == nil: case errors.As(err, &cond): outcome = "condition_failed" + case errors.Is(err, context.Canceled): + outcome = "canceled" case errors.Is(err, ErrUnavailable): outcome = "unavailable" default: diff --git a/internal/dedupe/dynamodb_test.go b/internal/dedupe/dynamodb_test.go index 8648c9382..e62761267 100644 --- a/internal/dedupe/dynamodb_test.go +++ b/internal/dedupe/dynamodb_test.go @@ -24,18 +24,18 @@ import ( // dynamodb-local (tests/integration); this is for the error paths it cannot // produce. type fakeDynamo struct { - put func(*dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) + put func(context.Context, *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) batch func(*dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) del func(*dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) describe func() (*dynamodb.DescribeTableOutput, error) ttl func() (*dynamodb.DescribeTimeToLiveOutput, error) } -func (f *fakeDynamo) PutItem(_ context.Context, in *dynamodb.PutItemInput, _ ...func(*dynamodb.Options)) (*dynamodb.PutItemOutput, error) { +func (f *fakeDynamo) PutItem(ctx context.Context, in *dynamodb.PutItemInput, _ ...func(*dynamodb.Options)) (*dynamodb.PutItemOutput, error) { if f.put == nil { return &dynamodb.PutItemOutput{}, nil } - return f.put(in) + return f.put(ctx, in) } func (f *fakeDynamo) BatchWriteItem(_ context.Context, in *dynamodb.BatchWriteItemInput, _ ...func(*dynamodb.Options)) (*dynamodb.BatchWriteItemOutput, error) { @@ -74,7 +74,12 @@ func apiErr(code string, fault smithy.ErrorFault) error { func openFake(t *testing.T, f *fakeDynamo) (*Dynamo, Deduplicator) { t.Helper() - d := newDynamo(f, DynamoConfig{Table: "dedupe"}) + return openFakeWith(t, f, DynamoConfig{Table: "dedupe"}) +} + +func openFakeWith(t *testing.T, f *fakeDynamo, cfg DynamoConfig) (*Dynamo, Deduplicator) { + t.Helper() + d := newDynamo(f, cfg) d.commitBackoff = func(int) time.Duration { return 0 } m := d.Tenant("acme") require.NoError(t, m.Apply(true)) @@ -125,7 +130,7 @@ func TestClassify(t *testing.T) { func TestDynamo_ReserveReadsTheHeldItem(t *testing.T) { t.Parallel() - _, m := openFake(t, &fakeDynamo{put: func(in *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { + _, m := openFake(t, &fakeDynamo{put: func(_ context.Context, in *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { id := string(in.Item[attrKey].(*types.AttributeValueMemberB).Value) switch id[len(id)-1] { case 'd': @@ -150,7 +155,7 @@ func TestDynamo_FailedReserveReleasesEveryPutThatMayHaveLanded(t *testing.T) { putTokens := map[string]string{} var released []string _, m := openFake(t, &fakeDynamo{ - put: func(in *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { + put: func(_ context.Context, in *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { id := string(in.Item[attrKey].(*types.AttributeValueMemberB).Value) mu.Lock() putTokens[id] = string(in.Item[attrToken].(*types.AttributeValueMemberB).Value) @@ -252,7 +257,7 @@ func TestDynamo_BreakerShortCircuitsReserve(t *testing.T) { var puts atomic.Int64 var down atomic.Bool down.Store(true) - d, m := openFake(t, &fakeDynamo{put: func(*dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { + d, m := openFake(t, &fakeDynamo{put: func(context.Context, *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { puts.Add(1) if down.Load() { return nil, &types.ProvisionedThroughputExceededException{} @@ -355,3 +360,87 @@ func TestExpiresAt(t *testing.T) { assert.Equal(t, int64(102), expiresAt(base, 1500*time.Millisecond), "rounded up: never ends early") assert.Equal(t, int64(102), expiresAt(base.Add(time.Nanosecond), time.Second)) } + +func idOf(av types.AttributeValue) string { + b := av.(*types.AttributeValueMemberB).Value + return string(b[len(b)-2:]) +} + +func TestDynamo_CommitAttemptsEveryChunk(t *testing.T) { + t.Parallel() + var mu sync.Mutex + written := 0 + _, m := openFake(t, &fakeDynamo{batch: func(in *dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) { + reqs := in.RequestItems["dedupe"] + if idOf(reqs[0].PutRequest.Item[attrKey]) == "00" { + return nil, &types.InternalServerError{} + } + mu.Lock() + written += len(reqs) + mu.Unlock() + return &dynamodb.BatchWriteItemOutput{}, nil + }}) + var claims []Claim + for i := range 3 * batchWriteMax { + claims = append(claims, Claim{Key: keys(fmt.Sprintf("%02d", i))[0], Status: Claimed, Token: "t"}) + } + require.ErrorIs(t, m.Commit(t.Context(), claims, 0), ErrUnavailable) + assert.Equal(t, 2*batchWriteMax, written, "a failed chunk does not cancel the others: their records are published") +} + +func TestDynamo_ReleaseAttemptsEveryClaim(t *testing.T) { + t.Parallel() + var deletes atomic.Int64 + _, m := openFakeWith(t, &fakeDynamo{del: func(in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { + deletes.Add(1) + if idOf(in.Key[attrKey]) == "k0" { + return nil, &types.ProvisionedThroughputExceededException{} + } + return &dynamodb.DeleteItemOutput{}, nil + }}, DynamoConfig{Table: "dedupe", ReserveConcurrency: 1}) + var claims []Claim + for _, k := range keys("k0", "k1", "k2", "k3") { + claims = append(claims, Claim{Key: k, Status: Claimed, Token: "t"}) + } + require.ErrorIs(t, m.Release(t.Context(), claims), ErrUnavailable) + assert.Equal(t, int64(4), deletes.Load()) +} + +// One throttled put in a multi-key Reserve: the unsent puts are neither sent +// nor released, and the cancelled siblings do not reset the breaker. +func TestDynamo_FailedMultiKeyReserve(t *testing.T) { + t.Parallel() + var mu sync.Mutex + var put, released []string + _, m := openFakeWith(t, &fakeDynamo{ + put: func(ctx context.Context, in *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { + id := idOf(in.Item[attrKey]) + mu.Lock() + put = append(put, id) + mu.Unlock() + if id == "k0" { + return nil, &types.ProvisionedThroughputExceededException{} + } + <-ctx.Done() + return nil, ctx.Err() + }, + del: func(in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { + mu.Lock() + released = append(released, idOf(in.Key[attrKey])) + mu.Unlock() + return nil, &types.ConditionalCheckFailedException{} + }, + }, DynamoConfig{Table: "dedupe", ReserveConcurrency: 2}) + ks := keys("k0", "k1", "k2", "k3", "k4", "k5", "k6", "k7") + for range breakerTrips { + _, err := m.Reserve(t.Context(), ks, time.Minute) + require.ErrorIs(t, err, ErrUnavailable) + require.NotErrorIs(t, err, errBreakerOpen) + } + _, err := m.Reserve(t.Context(), ks, time.Minute) + require.ErrorIs(t, err, errBreakerOpen, "the cancelled siblings did not reset the count") + mu.Lock() + defer mu.Unlock() + assert.ElementsMatch(t, put, released, "exactly the sent puts are released") + assert.Less(t, len(put), breakerTrips*len(ks), "unsent puts were never sent") +} From ff047d2d374a1cadce7357c53571e2cae24c4cb9 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 02:16:31 -0400 Subject: [PATCH 06/38] test(dedupe): make the Dynamo failure tests fail without their fix Review round 2. The fakes' BatchWriteItem and DeleteItem honour their context, and the Commit test runs one chunk at a time, so reverting Commit/Release to cancel-on-first-error fails both tests (checked against 108499f4). The failed-Reserve test asserts the sent puts are released rather than every key, which the unsent-put cutoff made flaky. forEach no longer sits between expiresAt and its doc comment. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- internal/dedupe/dynamodb.go | 4 +-- internal/dedupe/dynamodb_test.go | 43 +++++++++++++++++++++----------- 2 files changed, 30 insertions(+), 17 deletions(-) diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index 08fd0c664..eb1c83838 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -462,8 +462,6 @@ func (s *dynamoStore) Release(ctx context.Context, claims []Claim) error { // Close is a no-op: the client is the Dynamo's, shared by every tenant. func (s *dynamoStore) Close() error { return nil } -// expiresAt is t+d in epoch seconds rounded up, so a claim or commit never -// ends before it was asked to: TTL attributes are whole seconds. // forEach runs do for every index, at most limit at once, and joins the // errors: one failure never stops the rest. func forEach(n, limit int, do func(i int) error) error { @@ -477,6 +475,8 @@ func forEach(n, limit int, do func(i int) error) error { return errors.Join(errs...) } +// expiresAt is t+d in epoch seconds rounded up, so a claim or commit never +// ends before it was asked to: TTL attributes are whole seconds. func expiresAt(t time.Time, d time.Duration) int64 { end := t.Add(d) sec := end.Unix() diff --git a/internal/dedupe/dynamodb_test.go b/internal/dedupe/dynamodb_test.go index e62761267..dc74d9d1d 100644 --- a/internal/dedupe/dynamodb_test.go +++ b/internal/dedupe/dynamodb_test.go @@ -25,8 +25,8 @@ import ( // produce. type fakeDynamo struct { put func(context.Context, *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) - batch func(*dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) - del func(*dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) + batch func(context.Context, *dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) + del func(context.Context, *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) describe func() (*dynamodb.DescribeTableOutput, error) ttl func() (*dynamodb.DescribeTimeToLiveOutput, error) } @@ -38,18 +38,24 @@ func (f *fakeDynamo) PutItem(ctx context.Context, in *dynamodb.PutItemInput, _ . return f.put(ctx, in) } -func (f *fakeDynamo) BatchWriteItem(_ context.Context, in *dynamodb.BatchWriteItemInput, _ ...func(*dynamodb.Options)) (*dynamodb.BatchWriteItemOutput, error) { +func (f *fakeDynamo) BatchWriteItem(ctx context.Context, in *dynamodb.BatchWriteItemInput, _ ...func(*dynamodb.Options)) (*dynamodb.BatchWriteItemOutput, error) { + if err := ctx.Err(); err != nil { + return nil, err + } if f.batch == nil { return &dynamodb.BatchWriteItemOutput{}, nil } - return f.batch(in) + return f.batch(ctx, in) } -func (f *fakeDynamo) DeleteItem(_ context.Context, in *dynamodb.DeleteItemInput, _ ...func(*dynamodb.Options)) (*dynamodb.DeleteItemOutput, error) { +func (f *fakeDynamo) DeleteItem(ctx context.Context, in *dynamodb.DeleteItemInput, _ ...func(*dynamodb.Options)) (*dynamodb.DeleteItemOutput, error) { + if err := ctx.Err(); err != nil { + return nil, err + } if f.del == nil { return &dynamodb.DeleteItemOutput{}, nil } - return f.del(in) + return f.del(ctx, in) } func (f *fakeDynamo) DescribeTable(context.Context, *dynamodb.DescribeTableInput, ...func(*dynamodb.Options)) (*dynamodb.DescribeTableOutput, error) { @@ -168,7 +174,7 @@ func TestDynamo_FailedReserveReleasesEveryPutThatMayHaveLanded(t *testing.T) { } return &dynamodb.PutItemOutput{}, nil }, - del: func(in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { + del: func(_ context.Context, in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { id := string(in.Key[attrKey].(*types.AttributeValueMemberB).Value) mu.Lock() defer mu.Unlock() @@ -179,7 +185,14 @@ func TestDynamo_FailedReserveReleasesEveryPutThatMayHaveLanded(t *testing.T) { }) _, err := m.Reserve(t.Context(), keys("ok1", "dup", "bad", "ok2"), time.Minute) require.ErrorIs(t, err, ErrUnavailable) - assert.ElementsMatch(t, []string{"ok1", "bad", "ok2"}, released, "the failed put may have landed; the duplicate was never ours") + var sent []string + for id := range putTokens { + if id[len(id)-3:] != "dup" { + sent = append(sent, id[len(id)-3:]) + } + } + assert.ElementsMatch(t, sent, released, "every sent put but the duplicate, which was never ours") + assert.Contains(t, released, "bad", "the failed put may have landed") } func TestDynamo_CommitRetriesUnprocessedItems(t *testing.T) { @@ -188,7 +201,7 @@ func TestDynamo_CommitRetriesUnprocessedItems(t *testing.T) { var mu sync.Mutex written := map[string]int{} heldBack := map[string]bool{} - _, m := openFake(t, &fakeDynamo{batch: func(in *dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) { + _, m := openFake(t, &fakeDynamo{batch: func(_ context.Context, in *dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) { calls.Add(1) reqs := in.RequestItems["dedupe"] assert.LessOrEqual(t, len(reqs), batchWriteMax) @@ -228,7 +241,7 @@ func TestDynamo_CommitRetriesUnprocessedItems(t *testing.T) { func TestDynamo_CommitGivesUpOnItemsThatStayUnprocessed(t *testing.T) { t.Parallel() - _, m := openFake(t, &fakeDynamo{batch: func(in *dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) { + _, m := openFake(t, &fakeDynamo{batch: func(_ context.Context, in *dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) { return &dynamodb.BatchWriteItemOutput{UnprocessedItems: in.RequestItems}, nil }}) err := m.Commit(t.Context(), []Claim{{Key: keys("a")[0], Status: Claimed, Token: "t"}}, 0) @@ -237,7 +250,7 @@ func TestDynamo_CommitGivesUpOnItemsThatStayUnprocessed(t *testing.T) { func TestDynamo_ReleaseTreatsAFailedConditionAsDone(t *testing.T) { t.Parallel() - _, m := openFake(t, &fakeDynamo{del: func(in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { + _, m := openFake(t, &fakeDynamo{del: func(_ context.Context, in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { assert.Equal(t, condRelease, aws.ToString(in.ConditionExpression)) id := string(in.Key[attrKey].(*types.AttributeValueMemberB).Value) if id[len(id)-1] == 'x' { @@ -370,7 +383,7 @@ func TestDynamo_CommitAttemptsEveryChunk(t *testing.T) { t.Parallel() var mu sync.Mutex written := 0 - _, m := openFake(t, &fakeDynamo{batch: func(in *dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) { + _, m := openFakeWith(t, &fakeDynamo{batch: func(_ context.Context, in *dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) { reqs := in.RequestItems["dedupe"] if idOf(reqs[0].PutRequest.Item[attrKey]) == "00" { return nil, &types.InternalServerError{} @@ -379,7 +392,7 @@ func TestDynamo_CommitAttemptsEveryChunk(t *testing.T) { written += len(reqs) mu.Unlock() return &dynamodb.BatchWriteItemOutput{}, nil - }}) + }}, DynamoConfig{Table: "dedupe", ReserveConcurrency: 1}) var claims []Claim for i := range 3 * batchWriteMax { claims = append(claims, Claim{Key: keys(fmt.Sprintf("%02d", i))[0], Status: Claimed, Token: "t"}) @@ -391,7 +404,7 @@ func TestDynamo_CommitAttemptsEveryChunk(t *testing.T) { func TestDynamo_ReleaseAttemptsEveryClaim(t *testing.T) { t.Parallel() var deletes atomic.Int64 - _, m := openFakeWith(t, &fakeDynamo{del: func(in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { + _, m := openFakeWith(t, &fakeDynamo{del: func(_ context.Context, in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { deletes.Add(1) if idOf(in.Key[attrKey]) == "k0" { return nil, &types.ProvisionedThroughputExceededException{} @@ -424,7 +437,7 @@ func TestDynamo_FailedMultiKeyReserve(t *testing.T) { <-ctx.Done() return nil, ctx.Err() }, - del: func(in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { + del: func(_ context.Context, in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { mu.Lock() released = append(released, idOf(in.Key[attrKey])) mu.Unlock() From ced118c0f0d2e697b13fda9d892f75c21189bc8c Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 03:41:22 -0400 Subject: [PATCH 07/38] feat(app): choose the DynamoDB dedupe backend at boot dedupe.backend: dynamodb selects the shared table (F3's dedupe.Dynamo), configured by a dedupe.dynamodb block; dedupe.lease and dedupe.reserve_concurrency join the dedupe boot block. Boot checks the table, creating it first only with create_table on dynamodb-local; a failed check refuses a flat boot and fails switched-on tenants closed over a nested directory until a reload's check passes (Factory.Gated). The lease is capped at the embedded queue's 2m duplicate window. Part of #613 (PR F5). Co-Authored-By: Claude Opus 5.5 (1M context) --- AGENTS.md | 4 +- CHANGELOG.md | 3 +- config.yaml | 16 +- docs/src/content/docs/api.md | 4 +- docs/src/content/docs/architecture.md | 8 +- docs/src/content/docs/configuration.mdx | 51 +++++- docs/src/content/docs/deployment.md | 20 ++- docs/src/content/docs/settings-directory.mdx | 6 +- internal/app/app.go | 2 +- internal/app/dedupe_dynamodb_test.go | 164 +++++++++++++++++ internal/app/wire.go | 78 +++++++- internal/config/backends.go | 87 ++++++++- internal/config/backends_test.go | 143 ++++++++++++++- internal/dedupe/stores.go | 19 ++ internal/dedupe/stores_test.go | 19 ++ tests/integration/dedupe_dynamodb_app_test.go | 169 ++++++++++++++++++ 16 files changed, 758 insertions(+), 35 deletions(-) create mode 100644 internal/app/dedupe_dynamodb_test.go create mode 100644 tests/integration/dedupe_dynamodb_app_test.go diff --git a/AGENTS.md b/AGENTS.md index 8239899d5..790a728cd 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -34,8 +34,8 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`cache/`** — `Cache` interface → `LocalCache` (Ristretto: one pool for every tenant) + `VersionManager` (the invalidation index). Every key leads with the tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 8) — `:query:` for a result and its singleflight, `..
.
.` for a namespace — so no cached read or coalesced flight crosses tenants, a bump through `Invalidate` names one tenant's namespaces and no other's, and `InvalidateTenant` advances the tenant version that leads every namespace key of one tenant, orphaning every cached query keyed by its tables in one step (a pipe result names no table and keeps its TTL, [#343](https://github.com/Wave-RF/WaveHouse/pull/343)); the one crossing is the wiring's, above the package: `internal/app` hands the ingest worker the cache through `sharedTables`, which repeats each of the worker's bumps under every tenant on the same ClickHouse address and database (`chconn.Pools.SharingTables`, whatever their user or tls block — they read the same tables), and orphans the table-keyed cache (the structured-query results) of a tenant back on a pool after an absence, since it was out of that fan-out while away, or moved to another address or database, since it now reads other tables (story 6) - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) -- **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability when a selected backend keeps state there (`NeedsDataDir`); `backends.go` holds each layer's `.backend` (only the in-process value today; `coord.backend` reserved) — boot is the validator, there is no dry run -- **`dedupe/`** — `Deduplicator` interface (two-phase `Reserve`/`Commit`/`Release` over `Key{Table, ID}`; every backend passes the `dedupetest` conformance suite) → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant and table, pending claims in memory, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges) or `Dynamo` (one shared DynamoDB table, conditional `PutItem` claims; built and conformance-tested against dynamodb-local but not yet selectable at boot), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) +- **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability when a selected backend keeps state there (`NeedsDataDir`); `backends.go` holds each layer's `.backend` (the in-process value by default; `dedupe.backend` also takes `dynamodb`, with its `dedupe.dynamodb` sub-block; `coord.backend` reserved) — boot is the validator, there is no dry run +- **`dedupe/`** — `Deduplicator` interface (two-phase `Reserve`/`Commit`/`Release` over `Key{Table, ID}`; every backend passes the `dedupetest` conformance suite) → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant and table, pending claims in memory, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges) or `Dynamo` (one shared DynamoDB table, conditional `PutItem` claims; conformance-tested against dynamodb-local, selected by `dedupe.backend: dynamodb`; boot checks the table and never creates it outside dynamodb-local), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) - **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, token encoding, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`); `internal/app` constructs it and hands everything else a `mq.Broker` diff --git a/CHANGELOG.md b/CHANGELOG.md index e60f59a6c..f74c6cb50 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,8 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. `Reserve` is a conditional `PutItem` that is atomic across pods. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. +- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/backends.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/stores.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `2m`, the embedded queue's duplicate window), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` Binary alone; TTL off on `ex` is a warning) whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until a reload's check passes. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. +- **A DynamoDB dedupe backend** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. `Reserve` is a conditional `PutItem` that is atomic across pods. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. `dedupe.backend: dynamodb` selects it (entry above). New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block; no backend has settings yet, so every such sub-block is an unknown key for now and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name the `mq`, `cache` and `dedupe` backends: the zero value is not the default, and `app.New` refuses it. `coord.backend` is reserved: nothing reads it until the lease layer lands, and the sweeper still runs in every process. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. diff --git a/config.yaml b/config.yaml index 5519b78cd..7d4cbd450 100644 --- a/config.yaml +++ b/config.yaml @@ -43,12 +43,22 @@ clickhouse: password: "" max_total_conns: 0 # ceiling on open native connections across pools; 0 = none -# Each layer's implementation, chosen at boot. Only the in-process backend -# exists for each today, and it is the default. +# Each layer's implementation, chosen at boot. The in-process backend is +# each layer's default. mq: backend: embedded # NATS JetStream under /nats dedupe: - backend: pebble # Pebble under /pebble + backend: pebble # Pebble under /pebble; or dynamodb (below) + lease: 30s # how long a claimed id stays pending; at most 2m with the embedded mq + reserve_concurrency: 64 # parallel calls per request to a remote backend + # dynamodb: # read only when backend is dynamodb; credentials from the AWS SDK chain + # table: wavehouse-dedupe-prod + # region: "" # empty = AWS_REGION + # endpoint: "" # dynamodb-local only + # timeout: 250ms + # max_attempts: 3 + # retry_mode: standard # or adaptive + # create_table: false # dynamodb-local only; refused without endpoint coord: backend: local # reserved: nothing is elected yet diff --git a/docs/src/content/docs/api.md b/docs/src/content/docs/api.md index dc79ed6b9..fdaf46961 100644 --- a/docs/src/content/docs/api.md +++ b/docs/src/content/docs/api.md @@ -275,7 +275,7 @@ The body is a **flat JSON object** whose keys must match column names in the tar | 503 | `{"error":"schema not loaded yet"}` | The tenant's first schema discovery has not succeeded yet (its ClickHouse unreachable, or [no pool for it](/settings-directory#clickhouse)), so whether the table exists is not known; `Retry-After: 5`. Decided before the body is read | | 500 | `{"error":"publish failed"}` | Message queue error. With dedupe on, the record's id is given back, so a retry is published rather than reported as a duplicate. | | 503 | `{"error":"service unavailable"}` | The tenant's ingest queue is full (backpressure, for that tenant alone) or not open (see [Message Queue](/settings-directory#message-queue)). Response includes `Retry-After: 30` header. With dedupe on, the record's id is given back, so the retry is published rather than reported as a duplicate. | -| 503 | `{"error":"a request with the same dedupe id is in flight"}` | Dedupe is on and another request carrying the same id is still being published — usually a client's timeout-retry racing its own original. Its outcome decides whether this record is a duplicate, so retry after the `Retry-After` header (the dedupe lease, 30 seconds). | +| 503 | `{"error":"a request with the same dedupe id is in flight"}` | Dedupe is on and another request carrying the same id is still being published — usually a client's timeout-retry racing its own original. Its outcome decides whether this record is a duplicate, so retry after the `Retry-After` header (the dedupe lease, [`dedupe.lease`](/configuration#dedupe), 30 seconds by default). | | 503 | `{"error":"token verifier not ready: the tenant's JWKS has not been fetched yet"}` | A token was supplied, with no valid operator key, while the tenant's JWKS has not been fetched yet; refused before any policy runs, with a `Retry-After: 30` header — see [Authentication](#authentication) | **curl example:** @@ -387,7 +387,7 @@ A `200` is returned whenever the body was read and the records were processed | 415 | `{"error":"no Content-Type: ingest requires one of application/json, application/x-ndjson, …"}` (declared variant: `Content-Type "text/plain": ingest requires one of …` — see the note above on how declarations are echoed; conflicting variant: `conflicting Content-Type declarations "application/json", "application/x-ndjson": ingest reads one format per request, and requires one of …`) | The request declared no `Content-Type`, one whose media type is unsupported or does not parse, a comma-bearing value that does not parse as a single media type, or repeated header lines that disagree — different formats, or one supported and one not. Checked before the body is parsed | | 500 | `{"error":"publish failed"}` / `{"error":"dedupe failed"}` | Message-queue or dedup-backend failure mid-batch. After a publish failure the failing record's id is given back and the records before it keep theirs, so a whole-batch retry reports those as duplicates and publishes the rest | | 503 | `{"error":"service unavailable"}` | The tenant's ingest queue is full (backpressure) or not open, mid-batch; includes `Retry-After: 30`. As for `publish failed`, the failing record's id is given back and the records before it keep theirs | -| 503 | `{"error":"a request with the same dedupe id is in flight"}` | A record's dedupe id is held by another request still being published; includes `Retry-After` (the dedupe lease, 30 seconds). The records before it were published | +| 503 | `{"error":"a request with the same dedupe id is in flight"}` | A record's dedupe id is held by another request still being published; includes `Retry-After` (the dedupe lease, [`dedupe.lease`](/configuration#dedupe), 30 seconds by default). The records before it were published | | 503 | `{"error":"token verifier not ready: the tenant's JWKS has not been fetched yet"}` | A token was supplied, with no valid operator key, while the tenant's JWKS has not been fetched yet; refused before any policy runs, with a `Retry-After: 30` header — see [Authentication](#authentication) | :::caution[At-least-once on retry] diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 4b10348d6..0bb4030cf 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -90,7 +90,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability (after which each `Config.Warnings` line is logged at `WARN`), ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where each would be redelivered every ack wait for as long as the tenant is away and would hold that tenant's ack floor, so the sweeper could purge none of its queue past it. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultPolicy` reads it through the registry, where a flat directory's tenant `0` is always served. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) and then its table — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where each would be redelivered every ack wait for as long as the tenant is away and would hold that tenant's ack floor, so the sweeper could purge none of its queue past it. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultPolicy` reads it through the registry, where a flat directory's tenant `0` is always served. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) and then its table — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The `dynamodb` case builds the same `Stores` over `Dynamo.Tenant`, gated (`Factory.Gated`) on the table's check: boot runs `Dynamo.Check` (after `CreateTable`, when `dedupe.dynamodb.create_table` is on) whether or not any tenant has dedupe on, and a table that fails it follows the same rule as a Pebble instance that cannot open, with every reload checking again until it passes. It has no Pebble gauges. Both cases hand the ingest handler `dedupe.lease` (`IngestHandler.DedupeLease`). The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out @@ -126,10 +126,10 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key layout every backend stores: a version byte, the tenant, a NUL, the table, a NUL, the id ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)). No tenant id holds a NUL and settings validation refuses one in a table name, so neither tenants nor tables share ids; an id over 1,024 bytes (or one starting `0xFF`) is stored as `0xFF` plus its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. -- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (binary) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. +- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, selected by `dedupe.backend: dynamodb`: every tenant's ids in one shared table, `pk` (binary) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also refuses an unstorable key, reads a lease `<= 0` as `DefaultLease`, and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. -- **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. +- **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `Factory.Gated(ready)` wraps a factory so a store opens only once `ready` returns nil, and fails closed until then (the DynamoDB wiring's table check). `internal/app` drives it from the registry's `AfterAdopt` hook. ### `discovery/` — Schema Discovery & Validation @@ -232,7 +232,7 @@ Client POST /v1/ingest?table={table} setting it to null is published un-deduped + logged/counted, or rejected under require_id); once the record is encoded, reserve (tenant, table, id): a duplicate is skipped, an id another request holds → 503 + Retry-After - (the 30s lease) + (dedupe.lease, 30s by default) → Publish to NATS JetStream (ingest.{tenant}.{table}) → Commit the reserved id; on a failed publish, release it instead → 200 OK returned immediately diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 193a6c21b..3a7b8b12a 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -39,16 +39,39 @@ This page is boot config only — what the platform operator owns (wiring, lifec ### Backends -Each layer's implementation is chosen once, at boot. Today every layer has one backend, the in-process one, and it is the default, so a config that sets none of these keys runs as it always has. A value this build has no backend for refuses boot and names the valid ones. +Each layer's implementation is chosen once, at boot. The in-process backend is every layer's default, so a config that sets none of these keys runs as it always has. A value this build has no backend for refuses boot and names the valid ones. | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | | `mq.backend` | `WH_MQ_BACKEND` | `embedded` | The message queue. `embedded`: NATS JetStream inside this process, under `/nats`. It listens on no port, so no other process can reach its queue. | | `cache.backend` | `WH_CACHE_BACKEND` | `local` | The query-result cache. `local`: in this process, sized by `cache.l1_max_cost`. | -| `dedupe.backend` | `WH_DEDUPE_BACKEND` | `pebble` | Where ingest dedupe keeps the event ids it has seen. `pebble`: in this process, under `/pebble`, open while any tenant has dedupe on. | +| `dedupe.backend` | `WH_DEDUPE_BACKEND` | `pebble` | Where ingest dedupe keeps the event ids it has seen. `pebble`: in this process, under `/pebble`, open while any tenant has dedupe on; two processes do not share seen ids. `dynamodb`: one DynamoDB table that every tenant and every process shares, configured by [`dedupe.dynamodb`](#dynamodb-dedupe). | | `coord.backend` | `WH_COORD_BACKEND` | `local` | Reserved for the leases that will elect work only one process may do at a time, such as the sweeper. Nothing is elected yet: every process runs its own sweeper, and `local`, the only value, changes nothing. | -Settings for one backend will go in a sub-block named after it, `.`, read only when that backend is selected. No backend has settings yet, so today any such sub-block, `mq.embedded` included, is an unknown key and refuses boot. `mq` and `dedupe` also appear in the settings directory's `config.json`, with different keys (`mq.max_bytes_gb`, `dedupe.enabled`, …): those are per-tenant tunables and stay there, and one written in `config.yaml` refuses boot as an unknown key. +Settings for one backend go in a sub-block named after it, `.`, read only when that backend is selected. `dedupe.dynamodb` is the only one so far; any other, `mq.embedded` included, is an unknown key and refuses boot. `mq` and `dedupe` also appear in the settings directory's `config.json`, with different keys (`mq.max_bytes_gb`, `dedupe.enabled`, …): those are per-tenant tunables and stay there, and one written in `config.yaml` refuses boot as an unknown key. + +### Dedupe + +Whether a tenant dedupes, and on which field, are settings-directory keys ([Deduplication](/settings-directory#deduplication)). What is boot config is where the seen ids live and how a claim behaves. + +| YAML Key | Env Var | Default | Description | +| --- | --- | ------- | ----------- | +| `dedupe.lease` | `WH_DEDUPE_LEASE` | `30s` | How long a record's id stays claimed while the record is published. Another request carrying the same id meanwhile gets `503` with this as `Retry-After`, in whole seconds; a claim that is neither committed nor released, because its process died mid-publish, lapses after it. At most `2m` with `mq.backend: embedded`, the embedded queue's duplicate window: a longer lease refuses boot. A Go duration (`30s`, `1m`). | +| `dedupe.reserve_concurrency` | `WH_DEDUPE_RESERVE_CONCURRENCY` | `64` | The most calls one request makes to a remote dedupe backend at once. `pebble` ignores it. | + +#### DynamoDB dedupe + +Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (Binary) alone, and TTL off on `ex` is logged as a warning. A table that fails the check, or cannot be reached, refuses boot with a flat settings directory. With a nested one, the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and every reload checks again. The check runs whether or not any tenant has dedupe on. + +| YAML Key | Env Var | Default | Description | +| --- | --- | ------- | ----------- | +| `dedupe.dynamodb.table` | `WH_DEDUPE_DYNAMODB_TABLE` | *(none)* | The shared table. Required. | +| `dedupe.dynamodb.region` | `WH_DEDUPE_DYNAMODB_REGION` | *(empty)* | The table's region. Empty uses the SDK chain's (`AWS_REGION`). | +| `dedupe.dynamodb.endpoint` | `WH_DEDUPE_DYNAMODB_ENDPOINT` | *(empty)* | A custom endpoint, for [dynamodb-local](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/DynamoDBLocal.html) in development and tests. Leave it empty against AWS. | +| `dedupe.dynamodb.timeout` | `WH_DEDUPE_DYNAMODB_TIMEOUT` | `250ms` | Deadline for each DynamoDB call, the SDK's retries included. | +| `dedupe.dynamodb.max_attempts` | `WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS` | `3` | Attempts per call, the first included. | +| `dedupe.dynamodb.retry_mode` | `WH_DEDUPE_DYNAMODB_RETRY_MODE` | `standard` | `standard`, or `adaptive`, which also slows the client down after throttling. | +| `dedupe.dynamodb.create_table` | `WH_DEDUPE_DYNAMODB_CREATE_TABLE` | `false` | Development only: create the table at boot if it is missing, with TTL on `ex`. Refused unless `endpoint` is set, so it never creates a table in AWS; the production table belongs to your infrastructure code. | ### Server @@ -214,7 +237,17 @@ cache: l1_max_cost: 67108864 dedupe: - backend: pebble # in-process Pebble under /pebble + backend: pebble # in-process Pebble under /pebble; or dynamodb + lease: 30s # at most 2m with the embedded mq + reserve_concurrency: 64 + # dynamodb: # read only when backend is dynamodb + # table: wavehouse-dedupe-prod + # region: "" # empty = AWS_REGION + # endpoint: "" # dynamodb-local only + # timeout: 250ms + # max_attempts: 3 + # retry_mode: standard + # create_table: false # dynamodb-local only coord: backend: local # reserved: nothing is elected yet @@ -268,6 +301,16 @@ WH_MQ_BACKEND=embedded WH_CACHE_BACKEND=local WH_CACHE_L1_MAX_COST=67108864 WH_DEDUPE_BACKEND=pebble +WH_DEDUPE_LEASE=30s +WH_DEDUPE_RESERVE_CONCURRENCY=64 +# Read only with WH_DEDUPE_BACKEND=dynamodb: +# WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod +# WH_DEDUPE_DYNAMODB_REGION= +# WH_DEDUPE_DYNAMODB_ENDPOINT= +# WH_DEDUPE_DYNAMODB_TIMEOUT=250ms +# WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS=3 +# WH_DEDUPE_DYNAMODB_RETRY_MODE=standard +# WH_DEDUPE_DYNAMODB_CREATE_TABLE=false WH_COORD_BACKEND=local WH_AUTH_JWT_SECRET=change-me-in-production diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 9bea802e7..5d0cb1866 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -423,10 +423,6 @@ The dedupe key now carries the table as well as the tenant ([#222](https://githu ## A shared dedupe table on DynamoDB -:::note[Not selectable yet] -The DynamoDB dedupe backend is built and tested (`internal/dedupe/dynamodb.go`), but no boot key chooses it yet: every deployment still uses the embedded Pebble store. A boot key to select it lands with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot-config work. This section describes the table that backend expects, so the infrastructure can be ready first. -::: - Pebble is per process, so two pods on it do not share seen ids. The DynamoDB backend keeps every tenant's ids in **one shared table**, and a conditional write makes a claim atomic across every pod that uses the table. WaveHouse **never creates this table in production**: the table belongs to your infrastructure code. The backend refuses to create a table unless it is pointed at a custom endpoint, so table creation only works against [dynamodb-local](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/DynamoDBLocal.html). What the backend requires of the table: @@ -438,7 +434,7 @@ What the backend requires of the table: | `ex` | Number | Epoch seconds: the lease end while pending, the retention end once committed; absent = never expires. | | `tk` | Binary | The claim token that `Release` matches. | -Only `pk` is declared in the table definition. Turn TTL on for `ex`. Correctness never depends on TTL, because a claim whose `ex` has passed counts as absent whether or not DynamoDB has deleted it yet; TTL only reclaims the storage. **Today TTL removes only lapsed claims:** ingest commits every id with no retention, so a committed item carries no `ex` and is kept forever, and the table grows by one item (about 200 bytes) per distinct id. Per-tenant retention is [#220](https://github.com/Wave-RF/WaveHouse/issues/220). The backend's table check, which boot will run once the backend is selectable, refuses a table whose key schema does not match and logs a warning if TTL is off. +Only `pk` is declared in the table definition. Turn TTL on for `ex`. Correctness never depends on TTL, because a claim whose `ex` has passed counts as absent whether or not DynamoDB has deleted it yet; TTL only reclaims the storage. **Today TTL removes only lapsed claims:** ingest commits every id with no retention, so a committed item carries no `ex` and is kept forever, and the table grows by one item (about 200 bytes) per distinct id. Per-tenant retention is [#220](https://github.com/Wave-RF/WaveHouse/issues/220). Boot checks the table: it refuses one whose key schema does not match, and logs a warning if TTL is off. An example in Terraform. Its tags are the five that Wave RF's own deployments put on every AWS resource (`Name`, `Project`, `Environment`, `ManagedBy`, `CostCenter`, with lowercase-kebab values); use your own conventions in their place: @@ -487,6 +483,20 @@ data "aws_iam_policy_document" "wavehouse_dedupe" { } ``` +Select it in the boot config, on every pod that should share seen ids (all the keys are in the [Configuration Reference](/configuration#dynamodb-dedupe)): + +```yaml +dedupe: + backend: dynamodb + dynamodb: + table: wavehouse-dedupe-prod + region: us-east-1 # or leave empty for AWS_REGION +``` + +or `WH_DEDUPE_BACKEND=dynamodb`, `WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod`. A table that is missing, has the wrong key schema, or cannot be reached with the pod's credentials refuses boot over a flat settings directory; over a nested one the pod boots, every tenant with dedupe on fails its ingest closed, and each reload checks the table again. The check runs whether or not any tenant has `dedupe.enabled` on. The per-tenant switch stays in each tenant's `config.json`. + +For development against dynamodb-local, set `dedupe.dynamodb.endpoint` (for example `http://localhost:8000`) and `create_table: true`, and give the SDK any static credentials (`AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`) and a region. `create_table` without an `endpoint` refuses boot. + - **Credentials** come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; the environment or a profile locally), never from WaveHouse configuration. - **Point-in-time recovery** is not needed. The table records which ids have been seen, so losing it produces duplicate rows, not lost events. - **Cost:** every new event is two writes (the claim, then the commit), and a duplicate is one. On-demand, that is about $1.25 per million new events in us-east-1. Provisioned capacity with auto scaling is cheaper once traffic is steady. Storage is the other line: every distinct id stays in the table (see TTL above), at DynamoDB's per-GB-month rate. diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index df099659c..156421f72 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -183,10 +183,10 @@ What stays in boot config is only what cannot change under a running process — ## Deduplication -Every dedupe knob lives here — there are no boot-config keys for it. The switch and its fields are resolved per record from one snapshot (table override → global value): +Every per-tenant dedupe knob lives here. Where the seen ids are kept (`dedupe.backend`) and how long a claim is held (`dedupe.lease`) are [boot config](/configuration#dedupe), the same for every tenant. The switch and its fields are resolved per record from one snapshot (table override → global value): -- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store in the embedded Pebble instance at `/pebble`, so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. -- `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its 30-second lease, and a later retry of it is accepted again — counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero. +- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store (in the embedded Pebble instance at `/pebble`, or its share of the DynamoDB table under `dedupe.backend: dynamodb`), so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. +- `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its lease (`dedupe.lease`, 30 seconds by default), and a later retry of it is accepted again — counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. - `dedupe.tables.
.{id_field, require_id}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. diff --git a/internal/app/app.go b/internal/app/app.go index 51935e7d1..d8f18283a 100644 --- a/internal/app/app.go +++ b/internal/app/app.go @@ -174,7 +174,7 @@ func New(ctx context.Context, opts Options) (app *App, err error) { return nil, err } a.wireDiscovery(ctx) - if err := a.wireDedupe(); err != nil { + if err := a.wireDedupe(ctx); err != nil { return nil, err } if err := a.wireMQ(ctx); err != nil { diff --git a/internal/app/dedupe_dynamodb_test.go b/internal/app/dedupe_dynamodb_test.go new file mode 100644 index 000000000..ded483a9d --- /dev/null +++ b/internal/app/dedupe_dynamodb_test.go @@ -0,0 +1,164 @@ +package app + +import ( + "context" + "io" + "net/http" + "net/http/httptest" + "path/filepath" + "strings" + "sync" + "testing" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + + "github.com/Wave-RF/WaveHouse/internal/config" + "github.com/Wave-RF/WaveHouse/internal/dedupe" + "github.com/Wave-RF/WaveHouse/internal/dedupe/dedupetest" + "github.com/Wave-RF/WaveHouse/internal/tenant" +) + +// fakeDynamo answers the DynamoDB JSON protocol for one table, enough for +// boot's check, the dev create path, and a claim and its commit. Whether the +// table exists is the test's to switch. +type fakeDynamo struct { + mu sync.Mutex + exists bool + calls []string +} + +func (f *fakeDynamo) setExists(v bool) { + f.mu.Lock() + defer f.mu.Unlock() + f.exists = v +} + +func (f *fakeDynamo) called(op string) bool { + f.mu.Lock() + defer f.mu.Unlock() + for _, c := range f.calls { + if c == op { + return true + } + } + return false +} + +func (f *fakeDynamo) ServeHTTP(w http.ResponseWriter, r *http.Request) { + _, _ = io.Copy(io.Discard, r.Body) + _, op, _ := strings.Cut(r.Header.Get("X-Amz-Target"), ".") + f.mu.Lock() + f.calls = append(f.calls, op) + if op == "CreateTable" { + f.exists = true + } + exists := f.exists + f.mu.Unlock() + w.Header().Set("Content-Type", "application/x-amz-json-1.0") + if !exists { + w.WriteHeader(http.StatusBadRequest) + _, _ = io.WriteString(w, `{"__type":"com.amazonaws.dynamodb.v20120810#ResourceNotFoundException","message":"Requested resource not found"}`) + return + } + body := `{}` + switch op { + case "DescribeTable", "CreateTable": + body = `{"Table":{"TableName":"dedupe","TableStatus":"ACTIVE",` + + `"KeySchema":[{"AttributeName":"pk","KeyType":"HASH"}],` + + `"AttributeDefinitions":[{"AttributeName":"pk","AttributeType":"B"}]}}` + case "DescribeTimeToLive": + body = `{"TimeToLiveDescription":{"AttributeName":"ex","TimeToLiveStatus":"ENABLED"}}` + case "BatchWriteItem": + body = `{"UnprocessedItems":{}}` + } + _, _ = io.WriteString(w, body) +} + +// dynamoConfig points cfg's dedupe at a fake table, with credentials from the +// environment as the SDK's default chain reads them — and nothing from the +// developer's own AWS files. +func dynamoConfig(t *testing.T, cfg *config.Config, exists bool) *fakeDynamo { + t.Helper() + fake := &fakeDynamo{exists: exists} + srv := httptest.NewServer(fake) + t.Cleanup(srv.Close) + none := filepath.Join(t.TempDir(), "none") + for k, v := range map[string]string{ + "AWS_ACCESS_KEY_ID": "local", "AWS_SECRET_ACCESS_KEY": "local", "AWS_SESSION_TOKEN": "", + "AWS_PROFILE": "", "AWS_CONFIG_FILE": none, "AWS_SHARED_CREDENTIALS_FILE": none, + "AWS_EC2_METADATA_DISABLED": "true", + } { + t.Setenv(k, v) + } + cfg.Dedupe = config.Dedupe{Backend: config.DedupeDynamoDB, DynamoDB: config.DedupeDynamoDBConfig{ + Table: "dedupe", Region: "us-east-1", Endpoint: srv.URL, MaxAttempts: 1, + }} + return fake +} + +var dedupeOn = map[string]any{"dedupe": map[string]any{"enabled": true, "id_field": "event_id", "require_id": false, "tables": map[string]any{}}} + +func TestNew_DynamoDBDedupe(t *testing.T) { + cfg := testConfig(t, writeSettings(t, dedupeOn)) + fake := dynamoConfig(t, cfg, true) + a := newApp(t, cfg, Options{}) + + assert.True(t, fake.called("DescribeTable"), "boot checks the table") + assert.False(t, fake.called("CreateTable"), "and never creates it without create_table") + store := a.dedup.For(tenant.Default) + require.True(t, store.Open()) + dup, err := dedupetest.Mark(t.Context(), store, eventKey) + require.NoError(t, err) + assert.False(t, dup) + assert.True(t, fake.called("PutItem"), "the claim went to the table") + assert.True(t, fake.called("BatchWriteItem"), "and so did its commit") + assert.Nil(t, a.dedupeStats, "no Pebble instance, so no Pebble gauges") + assert.NoDirExists(t, filepath.Join(cfg.DataDir, "pebble")) +} + +func TestNew_DynamoDBDedupeCreatesTheTableOnlyWhenAsked(t *testing.T) { + cfg := testConfig(t, writeSettings(t, dedupeOn)) + fake := dynamoConfig(t, cfg, false) + cfg.Dedupe.DynamoDB.CreateTable = true + a := newApp(t, cfg, Options{}) + assert.True(t, fake.called("CreateTable")) + assert.True(t, fake.called("UpdateTimeToLive")) + assert.True(t, a.dedup.For(tenant.Default).Open()) +} + +// A table that fails the check follows the registry's rule for the shape, +// as a Pebble instance that cannot open does. +func TestNew_DynamoDBDedupeTableMissing(t *testing.T) { + t.Run("flat refuses boot", func(t *testing.T) { + for name, patch := range map[string]map[string]any{"dedupe on": dedupeOn, "dedupe off": nil} { + t.Run(name, func(t *testing.T) { + guardGlobals(t) + cfg := testConfig(t, writeSettings(t, patch)) + dynamoConfig(t, cfg, false) + _, err := New(t.Context(), Options{Config: cfg}) + require.ErrorContains(t, err, "dedupe open") + require.ErrorContains(t, err, "ResourceNotFoundException") + }) + } + }) + t.Run("nested fails closed until a reload passes the check", func(t *testing.T) { + root := writeNestedSettings(t, map[string]map[string]any{"acme": dedupeOn, "globex": nil}) + cfg := testConfig(t, root) + fake := dynamoConfig(t, cfg, false) + a := newApp(t, cfg, Options{}) + + acme := a.dedup.For("acme") + assert.False(t, acme.Open()) + _, err := dedupetest.Mark(t.Context(), acme, eventKey) + require.ErrorIs(t, err, dedupe.ErrUnavailable, "switched on, table missing: ingest fails closed") + _, err = dedupetest.Mark(t.Context(), a.dedup.For("globex"), eventKey) + require.ErrorIs(t, err, dedupe.ErrDisabled) + + fake.setExists(true) + a.tenants.Reload("test") + assert.True(t, acme.Open(), "the reload checked again and opened the store") + _, err = dedupetest.Mark(context.Background(), acme, eventKey) + require.NoError(t, err) + }) +} diff --git a/internal/app/wire.go b/internal/app/wire.go index 02164497e..426db8873 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -461,10 +461,12 @@ func (a *App) wireDiscovery(ctx context.Context) { // wireDedupe builds the dedupe stores — the one place the implementation is // chosen. -func (a *App) wireDedupe() error { +func (a *App) wireDedupe(ctx context.Context) error { switch b := a.cfg.Dedupe.Backend; b { case config.DedupePebble: return a.wirePebbleDedupe() + case config.DedupeDynamoDB: + return a.wireDynamoDedupe(ctx) default: return unreachableBackend("dedupe.backend", b) } @@ -529,6 +531,79 @@ func (a *App) wirePebbleDedupe() error { return nil } +// errDynamoUnchecked is a store's open before the first table check has run. +var errDynamoUnchecked = errors.New("dedupe: dynamodb table not checked yet") + +// wireDynamoDedupe builds the dedupe stores over one DynamoDB table that +// every tenant and every process shares (dedupe.Dynamo), so a tenant's store +// opens for free once the table has passed its check. Boot checks it (after +// creating it, with create_table on dynamodb-local) whether or not any tenant +// has dedupe on, and never creates it otherwise. A table that fails the check +// follows the registry's rule for the shape, as Pebble's instance does: a +// flat directory refuses boot; a nested one boots with every switched-on +// store closed, so its ingest fails closed, and each reload checks again. +func (a *App) wireDynamoDedupe(ctx context.Context) error { + c := a.cfg.Dedupe.DynamoDB + d, err := dedupe.NewDynamo(ctx, dedupe.DynamoConfig{ + Table: c.Table, Region: c.Region, Endpoint: c.Endpoint, + Timeout: c.Timeout, MaxAttempts: c.MaxAttempts, RetryMode: c.RetryMode, + ReserveConcurrency: a.cfg.Dedupe.ReserveConcurrency, + }) + if err != nil { + return err + } + var mu sync.Mutex + state := errDynamoUnchecked // nil once the table has passed + check := func(ctx context.Context) error { + mu.Lock() + defer mu.Unlock() + if state == nil { + return nil + } + if c.CreateTable { + if state = d.CreateTable(ctx); state != nil { + return state + } + } + state = d.Check(ctx) + return state + } + ready := func() error { + mu.Lock() + defer mu.Unlock() + return state + } + stores := dedupe.NewStores(dedupe.Factory(d.Tenant).Gated(ready)) + a.dedup = stores + a.add(component{name: "dedupe", close: withoutContext(stores.Close)}) + reconcile := func(ctx context.Context) error { + if err := stores.Retain(a.served); err != nil { + slog.Error("dedupe store close failed", "error", err) + } + checkErr := check(ctx) + if checkErr != nil { + slog.Error("dedupe: dynamodb table check failed; ingest with dedupe on fails closed until a reload passes it", + "table", c.Table, "error", checkErr) + } + for id, store := range a.tenants.All() { + m := stores.For(id) + enabled := store.DedupeEnabled() + wasOpen := m.Open() + // The one failure an open has is the check's, logged above. + _ = m.Apply(enabled) + if m.Open() != wasOpen { + slog.Info("dedupe store reconciled with settings", "tenant", id, "enabled", enabled) + } + } + return checkErr + } + a.tenants.AfterAdopt(func([]tenant.ID) { _ = reconcile(a.stopCtx) }) + if err := reconcile(ctx); err != nil && !a.tenants.Nested() { + return fmt.Errorf("dedupe open: %w", err) + } + return nil +} + // wireMQ starts the MQ — the one place the implementation is chosen; // everything after it sees mq.Broker. func (a *App) wireMQ(ctx context.Context) error { @@ -867,6 +942,7 @@ func (a *App) wireHTTP(authMW func(http.Handler) http.Handler) { ingestHandler.PolicySource = (*settings.Store).Policy ingestHandler.Dedup = func(s *settings.Store) dedupe.Deduplicator { return a.dedup.For(s.Tenant()) } ingestHandler.DedupeSettings = (*settings.Store).DedupeFor + ingestHandler.DedupeLease = a.cfg.Dedupe.Lease // Readiness pings every open pool at once and is ready at the first // answer: one tenant's ClickHouse outage is not the process's. diff --git a/internal/config/backends.go b/internal/config/backends.go index f2ab93308..9b0cb57a5 100644 --- a/internal/config/backends.go +++ b/internal/config/backends.go @@ -1,9 +1,11 @@ package config import ( + "errors" "fmt" "slices" "strings" + "time" ) // Each layer's implementation is chosen here, once, at boot: `.backend` @@ -55,20 +57,80 @@ func (c Cache) validate() error { // DedupeBackend names where ingest dedupe keeps the ids it has seen. type DedupeBackend string -// DedupePebble is the Pebble instance inside this process, under -// /pebble, opened while any tenant has dedupe on. -const DedupePebble DedupeBackend = "pebble" +const ( + // DedupePebble is the Pebble instance inside this process, under + // /pebble, opened while any tenant has dedupe on. Seen ids are + // per process. + DedupePebble DedupeBackend = "pebble" + // DedupeDynamoDB is one DynamoDB table every tenant and every process + // shares, configured by dedupe.dynamodb. + DedupeDynamoDB DedupeBackend = "dynamodb" +) -var dedupeBackends = []DedupeBackend{DedupePebble} +var dedupeBackends = []DedupeBackend{DedupePebble, DedupeDynamoDB} -// Dedupe selects the dedupe store. Whether a tenant dedupes, and on which -// field, are settings-directory keys, not this block's. +// Dedupe selects the dedupe store. Whether a tenant dedupes, on which field, +// and for how long are settings-directory keys, not this block's. type Dedupe struct { Backend DedupeBackend `yaml:"backend" env:"WH_DEDUPE_BACKEND" env-default:"pebble"` + // Lease is how long a claimed id stays pending while its record is + // published; a claim its request never settles lapses after it. + Lease time.Duration `yaml:"lease" env:"WH_DEDUPE_LEASE" env-default:"30s"` + // ReserveConcurrency bounds the parallel calls one request makes to a + // remote backend. Pebble ignores it. + ReserveConcurrency int `yaml:"reserve_concurrency" env:"WH_DEDUPE_RESERVE_CONCURRENCY" env-default:"64"` + DynamoDB DedupeDynamoDBConfig `yaml:"dynamodb"` +} + +// DedupeDynamoDBConfig is the dynamodb backend's block, read only when it is +// selected. Credentials are the AWS SDK's default chain (EKS Pod Identity, +// IRSA, AWS_* variables), never keys here. +type DedupeDynamoDBConfig struct { + // Table is the shared table; WaveHouse never creates it outside + // dynamodb-local. Required. + Table string `yaml:"table" env:"WH_DEDUPE_DYNAMODB_TABLE"` + // Region overrides the SDK chain's (AWS_REGION). + Region string `yaml:"region" env:"WH_DEDUPE_DYNAMODB_REGION"` + // Endpoint points the client at dynamodb-local. + Endpoint string `yaml:"endpoint" env:"WH_DEDUPE_DYNAMODB_ENDPOINT"` + Timeout time.Duration `yaml:"timeout" env:"WH_DEDUPE_DYNAMODB_TIMEOUT" env-default:"250ms"` + MaxAttempts int `yaml:"max_attempts" env:"WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS" env-default:"3"` + RetryMode string `yaml:"retry_mode" env:"WH_DEDUPE_DYNAMODB_RETRY_MODE" env-default:"standard"` + // CreateTable creates the table at boot if it is missing. Development + // only: refused unless Endpoint is set. + CreateTable bool `yaml:"create_table" env:"WH_DEDUPE_DYNAMODB_CREATE_TABLE" env-default:"false"` } func (d Dedupe) validate() error { - return checkBackend("dedupe.backend", "WH_DEDUPE_BACKEND", d.Backend, dedupeBackends) + if err := checkBackend("dedupe.backend", "WH_DEDUPE_BACKEND", d.Backend, dedupeBackends); err != nil { + return err + } + if d.Lease < 0 { + return fmt.Errorf("dedupe.lease (WH_DEDUPE_LEASE) must be >= 0, got %s", d.Lease) + } + if d.ReserveConcurrency < 0 { + return fmt.Errorf("dedupe.reserve_concurrency (WH_DEDUPE_RESERVE_CONCURRENCY) must be >= 0, got %d", d.ReserveConcurrency) + } + if d.Backend == DedupeDynamoDB { + return d.DynamoDB.validate() + } + return nil +} + +func (d DedupeDynamoDBConfig) validate() error { + switch { + case strings.TrimSpace(d.Table) == "": + return errors.New("dedupe.dynamodb.table (WH_DEDUPE_DYNAMODB_TABLE) is required when dedupe.backend is dynamodb") + case d.Timeout < 0: + return fmt.Errorf("dedupe.dynamodb.timeout (WH_DEDUPE_DYNAMODB_TIMEOUT) must be >= 0, got %s", d.Timeout) + case d.MaxAttempts < 0: + return fmt.Errorf("dedupe.dynamodb.max_attempts (WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS) must be >= 0, got %d", d.MaxAttempts) + case d.RetryMode != "" && d.RetryMode != "standard" && d.RetryMode != "adaptive": + return fmt.Errorf("dedupe.dynamodb.retry_mode (WH_DEDUPE_DYNAMODB_RETRY_MODE) %q: want standard or adaptive", d.RetryMode) + case d.CreateTable && d.Endpoint == "": + return errors.New("dedupe.dynamodb.create_table (WH_DEDUPE_DYNAMODB_CREATE_TABLE) is for dynamodb-local only: set dedupe.dynamodb.endpoint, or create the table with your infrastructure code") + } + return nil } // CoordBackend names where leases for singleton work (the sweeper) are held. @@ -104,13 +166,22 @@ func checkBackend[T ~string](key, env string, got T, valid []T) error { return fmt.Errorf("%s (%s) %q is not a backend this build has; valid: %s", key, env, got, strings.Join(names, ", ")) } -// validateBackends checks every layer's backend and its sub-block. +// embeddedDuplicateWindow mirrors mq.EmbeddedDuplicateWindow, the embedded +// ingest stream's duplicate window (#613 F2). A lease longer than it would let +// the republish of a publish whose outcome was unknown land twice. +const embeddedDuplicateWindow = 2 * time.Minute + +// validateBackends checks every layer's backend and its sub-block, then the +// rules that span two layers. func (c *Config) validateBackends() error { for _, check := range []func() error{c.MQ.validate, c.Cache.validate, c.Dedupe.validate, c.Coord.validate} { if err := check(); err != nil { return err } } + if c.MQ.Backend == MQEmbedded && c.Dedupe.Lease > embeddedDuplicateWindow { + return fmt.Errorf("dedupe.lease (WH_DEDUPE_LEASE) %s exceeds the embedded mq's %s duplicate window: a claim must lapse before the queue forgets the publish it guards", c.Dedupe.Lease, embeddedDuplicateWindow) + } return nil } diff --git a/internal/config/backends_test.go b/internal/config/backends_test.go index 0e70dc7a3..803b48d7a 100644 --- a/internal/config/backends_test.go +++ b/internal/config/backends_test.go @@ -4,6 +4,7 @@ import ( "os" "path/filepath" "testing" + "time" "github.com/stretchr/testify/assert" "github.com/stretchr/testify/require" @@ -28,6 +29,10 @@ func TestLoad_BackendDefaults(t *testing.T) { assert.Equal(t, MQEmbedded, cfg.MQ.Backend) assert.Equal(t, CacheLocal, cfg.Cache.Backend) assert.Equal(t, DedupePebble, cfg.Dedupe.Backend) + assert.Equal(t, Dedupe{ + Backend: DedupePebble, Lease: 30 * time.Second, ReserveConcurrency: 64, + DynamoDB: DedupeDynamoDBConfig{Timeout: 250 * time.Millisecond, MaxAttempts: 3, RetryMode: "standard"}, + }, cfg.Dedupe) assert.Equal(t, CoordLocal, cfg.Coord.Backend) assert.False(t, cfg.Distributed()) assert.True(t, cfg.NeedsDataDir()) @@ -112,7 +117,7 @@ func TestValidate_UnknownBackend(t *testing.T) { }{ {"mq", func(c *Config) { c.MQ.Backend = "kafka" }, `mq.backend (WH_MQ_BACKEND) "kafka" is not a backend this build has; valid: embedded`}, {"cache", func(c *Config) { c.Cache.Backend = "redis" }, `cache.backend (WH_CACHE_BACKEND) "redis" is not a backend this build has; valid: local`}, - {"dedupe", func(c *Config) { c.Dedupe.Backend = "dynamodb" }, `dedupe.backend (WH_DEDUPE_BACKEND) "dynamodb" is not a backend this build has; valid: pebble`}, + {"dedupe", func(c *Config) { c.Dedupe.Backend = "redis" }, `dedupe.backend (WH_DEDUPE_BACKEND) "redis" is not a backend this build has; valid: pebble, dynamodb`}, {"coord", func(c *Config) { c.Coord.Backend = "nats" }, `coord.backend (WH_COORD_BACKEND) "nats" is not a backend this build has; valid: local`}, // The zero value, which a Config built without Load carries. {"empty", func(c *Config) { c.MQ.Backend = "" }, `mq.backend (WH_MQ_BACKEND) "" is not a backend`}, @@ -159,3 +164,139 @@ func TestNeedsDataDir(t *testing.T) { cfg.MQ.Backend = MQEmbedded assert.True(t, cfg.NeedsDataDir(), "the embedded mq keeps state under data_dir") } + +func TestLoad_DedupeDynamoDBFromEnv(t *testing.T) { + for k, v := range map[string]string{ + "WH_DEDUPE_BACKEND": "dynamodb", + "WH_DEDUPE_LEASE": "45s", + "WH_DEDUPE_RESERVE_CONCURRENCY": "16", + "WH_DEDUPE_DYNAMODB_TABLE": "wavehouse-dedupe-dev", + "WH_DEDUPE_DYNAMODB_REGION": "us-east-2", + "WH_DEDUPE_DYNAMODB_ENDPOINT": "http://localhost:8000", + "WH_DEDUPE_DYNAMODB_TIMEOUT": "1s", + "WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS": "5", + "WH_DEDUPE_DYNAMODB_RETRY_MODE": "adaptive", + "WH_DEDUPE_DYNAMODB_CREATE_TABLE": "true", + } { + t.Setenv(k, v) + } + cfg, err := Load("nonexistent.yaml") + require.NoError(t, err) + assert.Equal(t, Dedupe{ + Backend: DedupeDynamoDB, Lease: 45 * time.Second, ReserveConcurrency: 16, + DynamoDB: DedupeDynamoDBConfig{ + Table: "wavehouse-dedupe-dev", Region: "us-east-2", Endpoint: "http://localhost:8000", + Timeout: time.Second, MaxAttempts: 5, RetryMode: "adaptive", CreateTable: true, + }, + }, cfg.Dedupe) + assert.True(t, cfg.NeedsDataDir(), "the embedded mq still keeps state under data_dir") +} + +func TestLoad_DedupeDynamoDBFromYAML(t *testing.T) { + t.Parallel() + path := filepath.Join(t.TempDir(), "config.yaml") + require.NoError(t, os.WriteFile(path, []byte(` +dedupe: + backend: dynamodb + lease: 20s + dynamodb: + table: wavehouse-dedupe-prod + timeout: 400ms +`), 0o600)) + cfg, err := Load(path) + require.NoError(t, err) + assert.Equal(t, DedupeDynamoDB, cfg.Dedupe.Backend) + assert.Equal(t, 20*time.Second, cfg.Dedupe.Lease) + assert.Equal(t, 64, cfg.Dedupe.ReserveConcurrency) + assert.Equal(t, DedupeDynamoDBConfig{ + Table: "wavehouse-dedupe-prod", Timeout: 400 * time.Millisecond, MaxAttempts: 3, RetryMode: "standard", + }, cfg.Dedupe.DynamoDB) +} + +func TestLoad_DedupeDynamoDBRefusesUnknownKeys(t *testing.T) { + t.Parallel() + path := filepath.Join(t.TempDir(), "config.yaml") + require.NoError(t, os.WriteFile(path, []byte(` +dedupe: + backend: dynamodb + dynamodb: + table: t + access_key_id: AKIA + redis: + addr: localhost:6379 +`), 0o600)) + _, err := Load(path) + require.Error(t, err) + assert.Contains(t, err.Error(), "dedupe.dynamodb.access_key_id, dedupe.redis") +} + +func TestUnboundEnv_KnowsTheDedupeVariables(t *testing.T) { + t.Parallel() + assert.Empty(t, unboundEnv([]string{ + "WH_DEDUPE_LEASE=30s", "WH_DEDUPE_RESERVE_CONCURRENCY=64", + "WH_DEDUPE_DYNAMODB_TABLE=t", "WH_DEDUPE_DYNAMODB_REGION=us-east-1", + "WH_DEDUPE_DYNAMODB_ENDPOINT=http://localhost:8000", "WH_DEDUPE_DYNAMODB_TIMEOUT=250ms", + "WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS=3", "WH_DEDUPE_DYNAMODB_RETRY_MODE=standard", + "WH_DEDUPE_DYNAMODB_CREATE_TABLE=false", + })) +} + +func TestValidate_Dedupe(t *testing.T) { + t.Parallel() + dynamo := func(c *Config) { + c.Dedupe.Backend = DedupeDynamoDB + c.Dedupe.DynamoDB = DedupeDynamoDBConfig{Table: "t", Timeout: time.Second, MaxAttempts: 3, RetryMode: "standard"} + } + cases := []struct { + name string + set func(*Config) + want string // "" = valid + }{ + {"dynamodb", dynamo, ""}, + {"zero values read as the defaults", func(c *Config) { + c.Dedupe.Backend = DedupeDynamoDB + c.Dedupe.DynamoDB = DedupeDynamoDBConfig{Table: "t"} + }, ""}, + {"create_table with an endpoint", func(c *Config) { + dynamo(c) + c.Dedupe.DynamoDB.Endpoint, c.Dedupe.DynamoDB.CreateTable = "http://localhost:8000", true + }, ""}, + {"the block is not read under pebble", func(c *Config) { c.Dedupe.DynamoDB.CreateTable = true }, ""}, + {"lease at the duplicate window", func(c *Config) { c.Dedupe.Lease = 2 * time.Minute }, ""}, + {"create_table without an endpoint", func(c *Config) { + dynamo(c) + c.Dedupe.DynamoDB.CreateTable = true + }, "dedupe.dynamodb.create_table (WH_DEDUPE_DYNAMODB_CREATE_TABLE) is for dynamodb-local only"}, + {"no table", func(c *Config) { dynamo(c); c.Dedupe.DynamoDB.Table = " " }, "dedupe.dynamodb.table (WH_DEDUPE_DYNAMODB_TABLE) is required"}, + {"retry mode", func(c *Config) { dynamo(c); c.Dedupe.DynamoDB.RetryMode = "legacy" }, `retry_mode (WH_DEDUPE_DYNAMODB_RETRY_MODE) "legacy"`}, + {"negative timeout", func(c *Config) { dynamo(c); c.Dedupe.DynamoDB.Timeout = -time.Second }, "dedupe.dynamodb.timeout"}, + {"negative attempts", func(c *Config) { dynamo(c); c.Dedupe.DynamoDB.MaxAttempts = -1 }, "dedupe.dynamodb.max_attempts"}, + {"negative lease", func(c *Config) { c.Dedupe.Lease = -time.Second }, "dedupe.lease (WH_DEDUPE_LEASE) must be >= 0"}, + {"negative concurrency", func(c *Config) { c.Dedupe.ReserveConcurrency = -1 }, "dedupe.reserve_concurrency"}, + {"lease past the duplicate window", func(c *Config) { c.Dedupe.Lease = 3 * time.Minute }, "exceeds the embedded mq's 2m0s duplicate window"}, + } + for _, tc := range cases { + t.Run(tc.name, func(t *testing.T) { + t.Parallel() + cfg := defaultBackends() + tc.set(&cfg) + err := cfg.Validate() + if tc.want == "" { + require.NoError(t, err) + return + } + require.Error(t, err) + assert.Contains(t, err.Error(), tc.want) + }) + } +} + +func TestNeedsDataDir_DynamoDBDedupe(t *testing.T) { + t.Parallel() + cfg := defaultBackends() + cfg.Dedupe.Backend = DedupeDynamoDB + assert.True(t, cfg.NeedsDataDir(), "the embedded mq keeps state under data_dir") + cfg.MQ.Backend = "shared" + assert.False(t, cfg.NeedsDataDir(), "neither a shared mq nor dynamodb dedupe keeps state under data_dir") + assert.Len(t, cfg.Warnings(), 1, "only the local cache warning: dynamodb dedupe is shared") +} diff --git a/internal/dedupe/stores.go b/internal/dedupe/stores.go index 614912fec..5bb393666 100644 --- a/internal/dedupe/stores.go +++ b/internal/dedupe/stores.go @@ -16,6 +16,25 @@ import ( // nothing that holds the Stores changes with it. type Factory func(id tenant.ID) *Managed +// Gated returns a Factory whose stores open only once ready returns nil, its +// error being the open's: a store switched on meanwhile stays closed and +// fails closed (ErrUnavailable) until an Apply finds the backend ready. For a +// backend whose tenant opens are free but whose shared resource (a remote +// table) is checked once. +func (f Factory) Gated(ready func() error) Factory { + return func(id tenant.ID) *Managed { + m := f(id) + open := m.open + m.open = func() (Deduplicator, error) { + if err := ready(); err != nil { + return nil, err + } + return open() + } + return m + } +} + // Stores is one Managed store per tenant (#583 story 7), each following its // own tenant's dedupe.enabled through Apply. A store is built on first use // and forgotten by Retain once its tenant is no longer served; its seen ids diff --git a/internal/dedupe/stores_test.go b/internal/dedupe/stores_test.go index ea2ae0654..03e6ecc66 100644 --- a/internal/dedupe/stores_test.go +++ b/internal/dedupe/stores_test.go @@ -2,6 +2,7 @@ package dedupe import ( "context" + "errors" "testing" "time" @@ -134,3 +135,21 @@ func TestStores_CloseClosesEveryStore(t *testing.T) { assert.False(t, e.Open(), "the instance closes with the last store") require.NoError(t, s.Close(), "closing again is a no-op") } + +func TestFactory_GatedOpensOnlyOnceReady(t *testing.T) { + t.Parallel() + notReady := errors.New("table missing") + ready := notReady + gated := NewStores(Factory(NewEmbedded(t.TempDir()).Tenant).Gated(func() error { return ready })) + t.Cleanup(func() { _ = gated.Close() }) + acme := gated.For("acme") + + require.ErrorIs(t, acme.Apply(true), notReady) + assert.False(t, acme.Open()) + _, err := mark(context.Background(), acme, "e1") + require.ErrorIs(t, err, ErrUnavailable, "switched on but not ready: fails closed, never open") + + ready = nil + require.NoError(t, acme.Apply(true), "the next apply finds it ready") + assert.True(t, acme.Open()) +} diff --git a/tests/integration/dedupe_dynamodb_app_test.go b/tests/integration/dedupe_dynamodb_app_test.go new file mode 100644 index 000000000..f94ccb4fd --- /dev/null +++ b/tests/integration/dedupe_dynamodb_app_test.go @@ -0,0 +1,169 @@ +//go:build integration + +package tests + +import ( + "context" + "encoding/json" + "fmt" + "io" + "net" + "net/http" + "net/url" + "path/filepath" + "strings" + "testing" + "time" + + "github.com/stretchr/testify/assert" + "github.com/stretchr/testify/require" + + "github.com/Wave-RF/WaveHouse/internal/app" + "github.com/Wave-RF/WaveHouse/internal/config" + "github.com/Wave-RF/WaveHouse/internal/dedupe" + "github.com/Wave-RF/WaveHouse/internal/settings" + "github.com/Wave-RF/WaveHouse/internal/tenant" +) + +// TestDynamoDBDedupe_TwoInstancesShareSeenIDs boots two apps the way two +// pods run — each its own data_dir, embedded queue and ingest worker — with +// dedupe.backend dynamodb over one table on dynamodb-local, and checks an id +// ingested through either is a duplicate through the other, that ClickHouse +// holds each id once, and that dedupe.lease reaches ingest as the in-flight +// answer's Retry-After. +func TestDynamoDBDedupe_TwoInstancesShareSeenIDs(t *testing.T) { + e := env(t) + ctx := context.Background() + // The SDK's default chain, as in production; never the developer's files. + none := filepath.Join(t.TempDir(), "none") + for k, v := range map[string]string{ + "AWS_ACCESS_KEY_ID": "local", "AWS_SECRET_ACCESS_KEY": "local", "AWS_SESSION_TOKEN": "", + "AWS_PROFILE": "", "AWS_CONFIG_FILE": none, "AWS_SHARED_CREDENTIALS_FILE": none, + "AWS_EC2_METADATA_DISABLED": "true", + } { + t.Setenv(k, v) + } + + chTable := createTable(t, "event_id String, n UInt32", "ORDER BY event_id") + ddbTable := newDynamoTable() + const lease = 7 * time.Second + + boot := func(name string) string { + t.Helper() + files, err := tenantSettings(e.ch, testCHDatabase) + require.NoError(t, err) + var doc map[string]json.RawMessage + require.NoError(t, json.Unmarshal(files[settings.FileConfig], &doc)) + doc["dedupe"] = json.RawMessage(`{"enabled": true, "id_field": "event_id", "require_id": true, "tables": {}}`) + files[settings.FileConfig], err = json.Marshal(doc) + require.NoError(t, err) + dir := filepath.Join(t.TempDir(), name) + require.NoError(t, writeSettingsFiles(dir, files)) + + var lc net.ListenConfig + ln, err := lc.Listen(ctx, "tcp", "127.0.0.1:0") + require.NoError(t, err) + cfg := &config.Config{ + DataDir: t.TempDir(), + Server: config.Server{ShutdownTimeout: 10}, + ClickHouse: config.ClickHouse{Password: testCHPassword}, + MQ: config.MQ{Backend: config.MQEmbedded}, + Cache: config.Cache{Backend: config.CacheLocal, L1MaxCost: 1 << 20}, + Dedupe: config.Dedupe{Backend: config.DedupeDynamoDB, Lease: lease, DynamoDB: config.DedupeDynamoDBConfig{ + Table: ddbTable, Region: "us-east-1", Endpoint: e.dynamoEndpoint, + // dynamodb-local under a parallel suite is slower than the real thing. + Timeout: 5 * time.Second, CreateTable: true, + }}, + Coord: config.Coord{Backend: config.CoordLocal}, + Settings: config.Settings{Dir: dir}, + } + a, err := app.New(ctx, app.Options{Config: cfg, Listener: ln}) + require.NoError(t, err) + runCtx, stop := context.WithCancel(ctx) + runDone := make(chan error, 1) + go func() { runDone <- a.Run(runCtx) }() + t.Cleanup(func() { + stop() + assert.NoError(t, <-runDone) + closeCtx, cancel := context.WithTimeout(context.Background(), 10*time.Second) + defer cancel() + assert.NoError(t, a.Close(closeCtx)) + }) + baseURL := "http://" + ln.Addr().String() + require.NoError(t, waitForLive(ctx, baseURL, 30*time.Second)) + return baseURL + } + // Both create the table: the second finds it and leaves it as it is. + podA, podB := boot("a"), boot("b") + + ingest := func(baseURL, id string, n int) (int, string, http.Header) { + t.Helper() + body := fmt.Sprintf(`{"event_id": %q, "n": %d}`, id, n) + req, err := http.NewRequestWithContext(ctx, http.MethodPost, baseURL+"/v1/ingest?table="+url.QueryEscape(chTable), strings.NewReader(body)) + require.NoError(t, err) + req.Header.Set("Content-Type", "application/json") + resp, err := http.DefaultClient.Do(req) + require.NoError(t, err) + defer func() { _ = resp.Body.Close() }() + b, err := io.ReadAll(resp.Body) + require.NoError(t, err) + return resp.StatusCode, strings.TrimSpace(string(b)), resp.Header + } + accepted := func(baseURL, id string, n int) { + t.Helper() + status, body, _ := ingest(baseURL, id, n) + require.Equal(t, http.StatusOK, status, body) + require.JSONEq(t, `{"ok": true}`, body) + } + duplicate := func(baseURL, id string, n int) { + t.Helper() + status, body, _ := ingest(baseURL, id, n) + require.Equal(t, http.StatusOK, status, body) + require.JSONEq(t, `{"duplicate": true}`, body) + } + + accepted(podA, "e1", 1) + duplicate(podB, "e1", 2) + accepted(podB, "e2", 3) + duplicate(podA, "e2", 4) + duplicate(podA, "e1", 5) + + // A claim another process holds is in flight on both pods, for as long + // as the configured lease says. + peer := dynamoClient(t, ddbTable, dedupe.DynamoConfig{}).Tenant(tenant.Default) + require.NoError(t, peer.Apply(true)) + claims, err := peer.Reserve(ctx, []dedupe.Key{{Table: chTable, ID: "e3"}}, time.Minute) + require.NoError(t, err) + require.Equal(t, dedupe.Claimed, claims[0].Status) + for _, pod := range []string{podA, podB} { + status, body, header := ingest(pod, "e3", 6) + require.Equal(t, http.StatusServiceUnavailable, status, body) + assert.Equal(t, "7", header.Get("Retry-After"), "dedupe.lease, in seconds") + } + require.NoError(t, peer.Release(ctx, claims)) + accepted(podB, "e3", 7) + duplicate(podA, "e3", 8) + + // Each pod's worker wrote only what its pod accepted: each id once. + type row struct { + ID string + N uint32 + } + want := []row{{"e1", 1}, {"e2", 3}, {"e3", 7}} + require.Eventually(t, func() bool { + rows, err := e.chConn.Query(ctx, fmt.Sprintf("SELECT event_id, n FROM %s ORDER BY event_id", chTable)) + if err != nil { + return false + } + defer func() { _ = rows.Close() }() + var got []row + for rows.Next() { + var r row + if rows.Scan(&r.ID, &r.N) != nil { + return false + } + got = append(got, r) + } + return assert.ObjectsAreEqual(want, got) + }, 30*time.Second, 500*time.Millisecond, "ClickHouse holds each id once, from the pod that accepted it") +} From a8e43bd2c41c285fa02d4b72e592e1b9f695ac82 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 04:13:48 -0400 Subject: [PATCH 08/38] fix(app): retry a failed DynamoDB table check; refuse no region A nested directory has no watcher, so a table check that failed at boot (a throttle, credentials not yet issued) left dedupe'd ingest failing closed until someone reloaded. The check now retries in the background, backing off 1s to 30s, and reconciles once it passes. NewDynamo refuses a config that resolves no region, a certain error caught at boot. Docs: reserve_concurrency has no effect while ingest sends one id per call; 0 = default; Pebble-only sentences scoped to dedupe.backend: pebble. Co-Authored-By: Claude Opus 5.5 (1M context) --- AGENTS.md | 2 +- CHANGELOG.md | 2 +- config.yaml | 2 +- docs/src/content/docs/architecture.md | 6 +-- docs/src/content/docs/configuration.mdx | 14 +++---- docs/src/content/docs/deployment.md | 6 +-- docs/src/content/docs/settings-directory.mdx | 2 +- internal/app/dedupe_dynamodb_test.go | 41 ++++++++++++++++++++ internal/app/wire.go | 27 +++++++++++-- internal/config/backends.go | 4 +- internal/dedupe/dynamodb.go | 3 ++ 11 files changed, 87 insertions(+), 22 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 790a728cd..4d062bbed 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -35,7 +35,7 @@ Eighteen internal packages under `internal/` (plus `internal/testutil/` for shar - **`chconn/`** — `Pools`, one `Manager` (a `driver.Conn`) per distinct `Identity{Addr, Database, Username, Password, TLS}` tuple among the served tenants, reconciled from the settings registry's `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): tenants naming one tuple share its pool, sized to their largest `max_open_conns`/`max_idle_conns`; a tenant whose tuple changed is repointed; a tuple no tenant names is released after the longest `query_timeout` among the tenants it had (never dials; a resize swaps the connection with the same grace). The boot config's `clickhouse.max_total_conns` bounds the open pools' `max_open_conns` together: boot refuses naming sum and ceiling; at a reload a resize above it keeps the pool's size, and a tuple that cannot be opened (the ceiling, an unreadable certificate, or options the driver refuses) leaves its tenants on the pool they had or on none — logged, retried by the next reload. Every consumer resolves its tenant's pool per call: `For` (nil for a tenant on no pool, a `503`), `Target` (the tenant's own HTTP wiring over its pool's TLS config), `SharingTables`, `Ping` (every pool at once, ready at the first answer). `HTTPClients` keeps one `http.Client` per TLS config - **`chsql/`** — dependency-free ClickHouse SQL helpers shared by `query`/`policy` (avoids an import cycle): `QuoteIdent` (backtick-quote every identifier) + `BindUnsafe` (reject names with a literal `?`) - **`config/`** — YAML + env var config loading (cleanenv); strict on both sides (undeclared YAML key, unbound `WH_*` variable) and probes `data_dir` writability when a selected backend keeps state there (`NeedsDataDir`); `backends.go` holds each layer's `.backend` (the in-process value by default; `dedupe.backend` also takes `dynamodb`, with its `dedupe.dynamodb` sub-block; `coord.backend` reserved) — boot is the validator, there is no dry run -- **`dedupe/`** — `Deduplicator` interface (two-phase `Reserve`/`Commit`/`Release` over `Key{Table, ID}`; every backend passes the `dedupetest` conformance suite) → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant and table, pending claims in memory, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges) or `Dynamo` (one shared DynamoDB table, conditional `PutItem` claims; conformance-tested against dynamodb-local, selected by `dedupe.backend: dynamodb`; boot checks the table and never creates it outside dynamodb-local), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) +- **`dedupe/`** — `Deduplicator` interface (two-phase `Reserve`/`Commit`/`Release` over `Key{Table, ID}`; every backend passes the `dedupetest` conformance suite) → `Embedded` (Pebble: every tenant's seen ids in one instance at `data_dir/pebble`, each key led by its tenant and table, pending claims in memory, open while any tenant's store is — the layout is the implementation's call, and the wiring hands it `data_dir` once; its `Stats` feed the system gauges) or `Dynamo` (one shared DynamoDB table, conditional `PutItem` claims; conformance-tested against dynamodb-local, selected by `dedupe.backend: dynamodb`; boot checks the table and never creates it outside dynamodb-local), wrapped by `Managed` whose open/closed state follows the hot-reloadable `dedupe.enabled` in the settings directory's `config.json`; `Stores` holds one `Managed` per tenant, built through a `Factory` (`func(tenant.ID) *Managed`, `Embedded.Tenant` or, gated on the table check (`Factory.Gated`), `Dynamo.Tenant` in production; `Managed` opens its store through a function, so every backend gets the same switch), and reconciled from the registry's `AfterAdopt` hook — open exactly when the tenant is served with its switch on, closed with its seen ids kept otherwise ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) stories 7 and 3) - **`discovery/`** — `SchemaRegistry`, one per served tenant over a `Source` read once per refresh — the tenant's pool's connection and the database that pool was opened for, one snapshot, so a refused move keeps discovering the database the tenant's queries still use (`internal/app`'s `discoveries` builds, runs and stops them from `AfterAdopt` and `App.Close`: `RetryRefresh` until the first success, then `StartAutoRefresh` with a random first tick; `Lookup` answers `ErrNotLoaded` before the first success — the handlers' `503` with `Retry-After` — and `ErrUnknownTable` after; a failed loop attempt counts in `wavehouse_schema_refresh_failures_total{tenant}`), that introspects ClickHouse `system.columns` (name/type/nullability plus `default_expression` and 1-based `position`) and `system.tables` (each table's `create_table_query`, kept in-process and never serialized — an external-engine table renders its wiring there unconditionally — endpoint, bucket/host, database, username, S3 access key id; ClickHouse masks the password as `[HIDDEN]` from ~23.9, so the exposure is the topology, not the secret), records the server version, + `Validate()` for ingest payloads + `CanonicalizeTimestamps()` rewriting top-level `DateTime`/`DateTime64` column values to the canonical RFC 3339 UTC wire form pre-publish (Key Design Decision #19) - **`ingest/`** — Ingest worker pipeline (`worker.go`: JetStream input → per-table batch INSERT with DLQ output). The pipeline is **insert-only**. The wire format `EventMessage` (`types.go`) carries `{table_name, scope, received_timestamp, format, columns, row}` and nothing else — `row` is one positional `JSONCompactEachRow` line and `columns` names its slots, the table's **insertable** columns (a `MATERIALIZED`/`ALIAS` column cannot be named in an `INSERT`); the worker batches per (tenant, table, column list), the tenant read off each message's `mq.Topic`, and inserts each batch into its tenant's own ClickHouse (`chconn.Pools.Target`); the worker accepts whatever table name the envelope carries (table existence was already checked by the HTTP ingest handler, which `404`s an unknown table before publish; the worker doesn't re-validate), then bulk-INSERTs. In the embedded-NATS deployment (the default), the server runs with `DontListen: true` (`internal/mq/embedded.go`), so the only Publishers reachable on the `ingest.>` subjects are in-process Go code — today, only the HTTP `/v1/ingest?table={table}` handler. Non-insert mutations (`DELETE`/`UPDATE`/`TRUNCATE`/…) must go through `POST /v1/ops/query` under the admin role (the same `RequireAdmin` gate as the rest of `/v1/ops/*`), so non-admin callers never reach the proxy. A request with no token (or an invalid one) resolves to the `default_role`, which in a production config is not the admin role (setting them equal is a loudly-warned dev-only setting), so it can't reach this endpoint. Plus `Sweeper` (Active Sweeper for NATS message lifecycle) + `EventMessage`/`BufferConsumerName` types (`types.go`) - **`mq/`** — the message-queue boundary: the **only** package that imports NATS/JetStream (Key Design Decision #20). and the only one that knows how the broker works. Everything else addresses events by `Topic{Tenant, Table, Scope}` (a validated tenant id and raw names — the tenant leads every subject, `ingest..
`, so one wildcard selects a tenant's traffic, and a topic without one is refused) and states intent through the interfaces — `Publisher` (`ErrQueueFull` is the backpressure signal), `Subscriber`, `ConsumerManager`/`Consumer`/`ConsumerConfig` (the ingest worker's durable consumer), `DeadLetterer` and `DeadLetterStats` (park a message, count what is parked), `Purger` (drop what is both acked and older than a cutoff — the sweeper), `Replayer` (SSE gap-fill) — composed into `Broker`, which adds each tenant's byte budget (`SetMaxBytes`/`MaxBytes`: the `mq.max_bytes_gb` reload, which opens a tenant's queue the first time) and `Stats` (the system gauges' source). Every interface speaks per tenant, never per stream: the embedded implementation gives each tenant a queue of its own (a stream pair, `INGEST_`/`DLQ_`), and nothing outside the package may assume that layout — an external implementation may keep one shared stream. Subjects, prefixes, wildcards, token encoding, stream names, sequences, and ack floors are private to the one implementation, `EmbeddedNATS` (`embedded.go`, `subject.go`, `purge.go`); `internal/app` constructs it and hands everything else a `mq.Broker` diff --git a/CHANGELOG.md b/CHANGELOG.md index f74c6cb50..8f6e31a49 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/backends.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/stores.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `2m`, the embedded queue's duplicate window), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` Binary alone; TTL off on `ex` is a warning) whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until a reload's check passes. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. +- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/backends.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `2m`, the embedded queue's duplicate window), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`; no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` Binary alone; TTL off on `ex` is a warning) whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and on every reload. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. - **A DynamoDB dedupe backend** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. `Reserve` is a conditional `PutItem` that is atomic across pods. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. `dedupe.backend: dynamodb` selects it (entry above). New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block; no backend has settings yet, so every such sub-block is an unknown key for now and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name the `mq`, `cache` and `dedupe` backends: the zero value is not the default, and `app.New` refuses it. `coord.backend` is reserved: nothing reads it until the lease layer lands, and the sweeper still runs in every process. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. diff --git a/config.yaml b/config.yaml index 7d4cbd450..93f66539f 100644 --- a/config.yaml +++ b/config.yaml @@ -50,7 +50,7 @@ mq: dedupe: backend: pebble # Pebble under /pebble; or dynamodb (below) lease: 30s # how long a claimed id stays pending; at most 2m with the embedded mq - reserve_concurrency: 64 # parallel calls per request to a remote backend + reserve_concurrency: 64 # parallel calls per Reserve/Commit/Release to a remote backend; no effect yet (ingest sends one id per call) # dynamodb: # read only when backend is dynamodb; credentials from the AWS SDK chain # table: wavehouse-dedupe-prod # region: "" # empty = AWS_REGION diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 0bb4030cf..a5ca5ca37 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -90,7 +90,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability (after which each `Config.Warnings` line is logged at `WARN`), ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where each would be redelivered every ack wait for as long as the tenant is away and would hold that tenant's ack floor, so the sweeper could purge none of its queue past it. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultPolicy` reads it through the registry, where a flat directory's tenant `0` is always served. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) and then its table — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The `dynamodb` case builds the same `Stores` over `Dynamo.Tenant`, gated (`Factory.Gated`) on the table's check: boot runs `Dynamo.Check` (after `CreateTable`, when `dedupe.dynamodb.create_table` is on) whether or not any tenant has dedupe on, and a table that fails it follows the same rule as a Pebble instance that cannot open, with every reload checking again until it passes. It has no Pebble gauges. Both cases hand the ingest handler `dedupe.lease` (`IngestHandler.DedupeLease`). The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where each would be redelivered every ack wait for as long as the tenant is away and would hold that tenant's ack floor, so the sweeper could purge none of its queue past it. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultPolicy` reads it through the registry, where a flat directory's tenant `0` is always served. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own. The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) and then its table — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The `dynamodb` case builds the same `Stores` over `Dynamo.Tenant`, gated (`Factory.Gated`) on the table's check: boot runs `Dynamo.Check` (after `CreateTable`, when `dedupe.dynamodb.create_table` is on) whether or not any tenant has dedupe on, and a table that fails it follows the same rule as a Pebble instance that cannot open, with the check retried until it passes: by a background component that backs off from one second to thirty (a nested directory has no watcher), and by every reload. It has no Pebble gauges. `wireHTTP` hands the ingest handler `dedupe.lease` (`IngestHandler.DedupeLease`) whichever backend is chosen. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out @@ -117,7 +117,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — the implementation each layer runs on, resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). - **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load` when `Config.NeedsDataDir` says a selected backend keeps state there, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. -- **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` at the end of `Validate`), which refuses a value not on the list and names the ones that are. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns the valid combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), which `app.New` logs at `WARN`. +- **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` at the end of `Validate`), which refuses a value not on the list and names the ones that are. One rule spans two layers: `dedupe.lease` may not exceed the embedded MQ's 2m duplicate window (`embeddedDuplicateWindow`) while `mq.backend` is `embedded`. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns the valid combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), which `app.New` logs at `WARN`. - **strict.go** — `rejectUnknownKeys`, the YAML half: re-reads the file as a generic tree and walks it against the struct's `yaml` tags, listing every key the struct doesn't declare. cleanenv itself is lenient by design, which is exactly wrong for boot config once keys have moved to the settings directory. - **persistence.go** — `WarnIfFreshDataDir` logs the startup `WARN` when `data_dir` is missing or empty (on a redeploy, the sign that the volume didn't persist); `LogStorageInitError` attaches the UID-65532 `permissionHint` to a NATS or Pebble open failure that looks like a permission denial — the same hint string `CheckDataDir` uses. @@ -126,7 +126,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key layout every backend stores: a version byte, the tenant, a NUL, the table, a NUL, the id ([#222](https://github.com/Wave-RF/WaveHouse/issues/222)). No tenant id holds a NUL and settings validation refuses one in a table name, so neither tenants nor tables share ids; an id over 1,024 bytes (or one starting `0xFF`) is stored as `0xFF` plus its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. -- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, selected by `dedupe.backend: dynamodb`: every tenant's ids in one shared table, `pk` (binary) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. +- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, selected by `dedupe.backend: dynamodb`: every tenant's ids in one shared table, `pk` (binary) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local, and refuses a config that resolves no region. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also refuses an unstorable key, reads a lease `<= 0` as `DefaultLease`, and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `Factory.Gated(ready)` wraps a factory so a store opens only once `ready` returns nil, and fails closed until then (the DynamoDB wiring's table check). `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 3a7b8b12a..d6f071a04 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -35,7 +35,7 @@ This page is boot config only — what the platform operator owns (wiring, lifec | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | -| `data_dir` | `WH_DATA_DIR` | `./data` | Root directory for embedded state. NATS JetStream lives at `/nats`; Pebble, holding every tenant's dedupe store while any tenant has dedupe enabled, at `/pebble`. Subdirectory names are conventions, not config — one knob, one mount. **In a container this MUST resolve to a host-backed volume**; the relative default is for local binary use. WaveHouse logs a startup `WARN` when the directory is missing or empty (no prior state). See [Persistent Storage](/deployment#persistent-storage-required-for-containers). | +| `data_dir` | `WH_DATA_DIR` | `./data` | Root directory for embedded state. NATS JetStream lives at `/nats`; Pebble (with `dedupe.backend: pebble`), holding every tenant's dedupe store while any tenant has dedupe enabled, at `/pebble`. Subdirectory names are conventions, not config — one knob, one mount. **In a container this MUST resolve to a host-backed volume**; the relative default is for local binary use. WaveHouse logs a startup `WARN` when the directory is missing or empty (no prior state). See [Persistent Storage](/deployment#persistent-storage-required-for-containers). | ### Backends @@ -56,20 +56,20 @@ Whether a tenant dedupes, and on which field, are settings-directory keys ([Dedu | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | -| `dedupe.lease` | `WH_DEDUPE_LEASE` | `30s` | How long a record's id stays claimed while the record is published. Another request carrying the same id meanwhile gets `503` with this as `Retry-After`, in whole seconds; a claim that is neither committed nor released, because its process died mid-publish, lapses after it. At most `2m` with `mq.backend: embedded`, the embedded queue's duplicate window: a longer lease refuses boot. A Go duration (`30s`, `1m`). | -| `dedupe.reserve_concurrency` | `WH_DEDUPE_RESERVE_CONCURRENCY` | `64` | The most calls one request makes to a remote dedupe backend at once. `pebble` ignores it. | +| `dedupe.lease` | `WH_DEDUPE_LEASE` | `30s` | How long a record's id stays claimed while the record is published. Another request carrying the same id meanwhile gets `503` with this as `Retry-After`, in whole seconds; a claim that is neither committed nor released, because its process died mid-publish, lapses after it. At most `2m` with `mq.backend: embedded`, the embedded queue's duplicate window: a longer lease refuses boot. A Go duration (`30s`, `1m`); `0` = the default. | +| `dedupe.reserve_concurrency` | `WH_DEDUPE_RESERVE_CONCURRENCY` | `64` | The most parallel calls one Reserve, Commit or Release makes to a remote dedupe backend. Ingest sends one id per call today, so it has no effect yet; `pebble` ignores it. `0` = the default. | #### DynamoDB dedupe -Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (Binary) alone, and TTL off on `ex` is logged as a warning. A table that fails the check, or cannot be reached, refuses boot with a flat settings directory. With a nested one, the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and every reload checks again. The check runs whether or not any tenant has dedupe on. +Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (Binary) alone, and TTL off on `ex` is logged as a warning. A table that fails the check, or cannot be reached, refuses boot with a flat settings directory. With a nested one, the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and the check is retried in the background (backing off from one second to thirty) and on every reload, so a table that comes good is picked up without a restart. The check runs whether or not any tenant has dedupe on. | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | | `dedupe.dynamodb.table` | `WH_DEDUPE_DYNAMODB_TABLE` | *(none)* | The shared table. Required. | -| `dedupe.dynamodb.region` | `WH_DEDUPE_DYNAMODB_REGION` | *(empty)* | The table's region. Empty uses the SDK chain's (`AWS_REGION`). | +| `dedupe.dynamodb.region` | `WH_DEDUPE_DYNAMODB_REGION` | *(empty)* | The table's region. Empty uses the SDK chain's (`AWS_REGION`); no region from either refuses boot. | | `dedupe.dynamodb.endpoint` | `WH_DEDUPE_DYNAMODB_ENDPOINT` | *(empty)* | A custom endpoint, for [dynamodb-local](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/DynamoDBLocal.html) in development and tests. Leave it empty against AWS. | -| `dedupe.dynamodb.timeout` | `WH_DEDUPE_DYNAMODB_TIMEOUT` | `250ms` | Deadline for each DynamoDB call, the SDK's retries included. | -| `dedupe.dynamodb.max_attempts` | `WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS` | `3` | Attempts per call, the first included. | +| `dedupe.dynamodb.timeout` | `WH_DEDUPE_DYNAMODB_TIMEOUT` | `250ms` | Deadline for each DynamoDB call, the SDK's retries included. `0` = the default. | +| `dedupe.dynamodb.max_attempts` | `WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS` | `3` | Attempts per call, the first included. `0` = the default. | | `dedupe.dynamodb.retry_mode` | `WH_DEDUPE_DYNAMODB_RETRY_MODE` | `standard` | `standard`, or `adaptive`, which also slows the client down after throttling. | | `dedupe.dynamodb.create_table` | `WH_DEDUPE_DYNAMODB_CREATE_TABLE` | `false` | Development only: create the table at boot if it is missing, with TTL on `ex`. Refused unless `endpoint` is set, so it never creates a table in AWS; the production table belongs to your infrastructure code. | diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 5d0cb1866..3919bec49 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -169,7 +169,7 @@ WH_SETTINGS_DIR=/etc/wavehouse/settings WaveHouse keeps all embedded state under a single configurable root, `WH_DATA_DIR` (yaml: `data_dir`). Subdirectories are convention, not config: - `/nats` — embedded NATS JetStream. Holds in-flight events between an ingest POST and the ingest worker → ClickHouse flush, plus the `stream.gap_window_minutes` window (settings directory) of history that powers SSE gap-fill across restarts. -- `/pebble` — the Pebble dedup KV: one instance shared by every tenant, each key led by its tenant and table. Only used while some tenant's `dedupe.enabled` is `true` in its `config.json` (opened and closed on reload). +- `/pebble` — the Pebble dedup KV (with `dedupe.backend: pebble`, the default): one instance shared by every tenant, each key led by its tenant and table. Only used while some tenant's `dedupe.enabled` is `true` in its `config.json` (opened and closed on reload). In a Docker / Podman / Kubernetes deployment, **`data_dir` must resolve to a host-backed volume**. The reference compose file `deployments/compose/standalone.yaml` sets `WH_DATA_DIR=/app/data` and binds a `wavehouse-data:/app/data` volume — copy that pattern. The bundled Dockerfiles pre-create `/app/data` and `/app/settings` owned by the nonroot user (UID 65532); the binary creates the `nats/` and `pebble/` subdirectories under `/app/data` itself on first run. @@ -177,7 +177,7 @@ If `data_dir` resolves into the container's writable overlay layer instead, **Je Beyond persistence, the *speed* of that volume matters: JetStream `fsync`s every event to `/nats` before the ingest endpoint returns `200`, so the volume's `fsync` latency is your ingest latency floor. Managed cloud block storage handles this without thinking; commodity or virtualized substrates (ZFS without a SLOG, qcow2-on-`ext4`, spinning disks) can stall ingest with multi-second `fsync` tails. See [Durability & Storage](/durability) to measure yours before going live. -WaveHouse runs a simple existence check on startup and logs a `WARN` if `/nats` (or `/pebble`, when dedupe is on) is missing or empty: +WaveHouse runs a simple existence check on startup and logs a `WARN` if `/nats` (or `/pebble`, when dedupe is on with the `pebble` backend) is missing or empty: ```text wrap=false WARN data directory does not exist — starting with no prior state. @@ -493,7 +493,7 @@ dedupe: region: us-east-1 # or leave empty for AWS_REGION ``` -or `WH_DEDUPE_BACKEND=dynamodb`, `WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod`. A table that is missing, has the wrong key schema, or cannot be reached with the pod's credentials refuses boot over a flat settings directory; over a nested one the pod boots, every tenant with dedupe on fails its ingest closed, and each reload checks the table again. The check runs whether or not any tenant has `dedupe.enabled` on. The per-tenant switch stays in each tenant's `config.json`. +or `WH_DEDUPE_BACKEND=dynamodb`, `WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod`. A table that is missing, has the wrong key schema, or cannot be reached with the pod's credentials refuses boot over a flat settings directory; over a nested one the pod boots, every tenant with dedupe on fails its ingest closed, and the check is retried in the background (backing off from one second to thirty) and on every reload. No region at all (neither `region` nor `AWS_REGION`) refuses boot in both shapes. The check runs whether or not any tenant has `dedupe.enabled` on. The per-tenant switch stays in each tenant's `config.json`. For development against dynamodb-local, set `dedupe.dynamodb.endpoint` (for example `http://localhost:8000`) and `create_table: true`, and give the SDK any static credentials (`AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`) and a region. `create_table` without an `endpoint` refuses boot. diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index 156421f72..c7377620c 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -185,7 +185,7 @@ What stays in boot config is only what cannot change under a running process — Every per-tenant dedupe knob lives here. Where the seen ids are kept (`dedupe.backend`) and how long a claim is held (`dedupe.lease`) are [boot config](/configuration#dedupe), the same for every tenant. The switch and its fields are resolved per record from one snapshot (table override → global value): -- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store (in the embedded Pebble instance at `/pebble`, or its share of the DynamoDB table under `dedupe.backend: dynamodb`), so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. +- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store (in the embedded Pebble instance at `/pebble`, or its share of the DynamoDB table under `dedupe.backend: dynamodb`), so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) with `dedupe.backend: pebble`, every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. Under `dedupe.backend: dynamodb` the table check plays the instance's part: a table that fails it fails every tenant with dedupe on closed, whatever the switches say, until the check, retried in the background and on every reload, passes ([Configuration](/configuration#dynamodb-dedupe)). - `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its lease (`dedupe.lease`, 30 seconds by default), and a later retry of it is accepted again — counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. - `dedupe.tables.
.{id_field, require_id}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. diff --git a/internal/app/dedupe_dynamodb_test.go b/internal/app/dedupe_dynamodb_test.go index ded483a9d..c86d2294b 100644 --- a/internal/app/dedupe_dynamodb_test.go +++ b/internal/app/dedupe_dynamodb_test.go @@ -3,12 +3,14 @@ package app import ( "context" "io" + "net" "net/http" "net/http/httptest" "path/filepath" "strings" "sync" "testing" + "time" "github.com/stretchr/testify/assert" "github.com/stretchr/testify/require" @@ -162,3 +164,42 @@ func TestNew_DynamoDBDedupeTableMissing(t *testing.T) { require.NoError(t, err) }) } + +// A nested directory has no watcher, so a table that comes good is picked up +// by the background retry, not only by a reload someone has to send. +func TestRun_DynamoDBDedupeRetriesTheTableCheck(t *testing.T) { + root := writeNestedSettings(t, map[string]map[string]any{"acme": dedupeOn}) + cfg := testConfig(t, root) + fake := dynamoConfig(t, cfg, false) + var lc net.ListenConfig + ln, err := lc.Listen(t.Context(), "tcp", "127.0.0.1:0") + require.NoError(t, err) + a := newApp(t, cfg, Options{Listener: ln}) + acme := a.dedup.For("acme") + require.False(t, acme.Open()) + + _, stop := runApp(t, a, ln) + fake.setExists(true) + require.Eventually(t, acme.Open, 10*time.Second, 50*time.Millisecond, "the retry opened the store without a reload") + require.NoError(t, stop()) +} + +// No region anywhere is a certain config error: refused at boot in either +// shape rather than failing every check afterwards. +func TestNew_DynamoDBDedupeRefusesNoRegion(t *testing.T) { + for name, dir := range map[string]func(*testing.T) string{ + "flat": func(t *testing.T) string { return writeSettings(t, dedupeOn) }, + "nested": func(t *testing.T) string { return writeNestedSettings(t, map[string]map[string]any{"acme": dedupeOn}) }, + } { + t.Run(name, func(t *testing.T) { + guardGlobals(t) + cfg := testConfig(t, dir(t)) + dynamoConfig(t, cfg, true) + cfg.Dedupe.DynamoDB.Region = "" + t.Setenv("AWS_REGION", "") + t.Setenv("AWS_DEFAULT_REGION", "") + _, err := New(t.Context(), Options{Config: cfg}) + require.ErrorContains(t, err, "dynamodb region is not set") + }) + } +} diff --git a/internal/app/wire.go b/internal/app/wire.go index 426db8873..a2e23b7c2 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -541,7 +541,10 @@ var errDynamoUnchecked = errors.New("dedupe: dynamodb table not checked yet") // has dedupe on, and never creates it otherwise. A table that fails the check // follows the registry's rule for the shape, as Pebble's instance does: a // flat directory refuses boot; a nested one boots with every switched-on -// store closed, so its ingest fails closed, and each reload checks again. +// store closed, so its ingest fails closed. Unlike a local disk, a remote +// table's failure is usually brief (a throttle, credentials not yet issued +// mid-rollout), and a nested directory has no watcher to reload it, so the +// check is also retried in the background, with backoff, until it passes. func (a *App) wireDynamoDedupe(ctx context.Context) error { c := a.cfg.Dedupe.DynamoDB d, err := dedupe.NewDynamo(ctx, dedupe.DynamoConfig{ @@ -576,7 +579,10 @@ func (a *App) wireDynamoDedupe(ctx context.Context) error { stores := dedupe.NewStores(dedupe.Factory(d.Tenant).Gated(ready)) a.dedup = stores a.add(component{name: "dedupe", close: withoutContext(stores.Close)}) + var reconciling sync.Mutex // the hook and the retry loop both reconcile reconcile := func(ctx context.Context) error { + reconciling.Lock() + defer reconciling.Unlock() if err := stores.Retain(a.served); err != nil { slog.Error("dedupe store close failed", "error", err) } @@ -598,8 +604,23 @@ func (a *App) wireDynamoDedupe(ctx context.Context) error { return checkErr } a.tenants.AfterAdopt(func([]tenant.ID) { _ = reconcile(a.stopCtx) }) - if err := reconcile(ctx); err != nil && !a.tenants.Nested() { - return fmt.Errorf("dedupe open: %w", err) + if err := reconcile(ctx); err != nil { + if !a.tenants.Nested() { + return fmt.Errorf("dedupe open: %w", err) + } + a.add(component{name: "dedupe table check", run: func(ctx context.Context) error { + for wait := time.Second; ready() != nil; wait = min(2*wait, 30*time.Second) { + select { + case <-ctx.Done(): + return nil + case <-time.After(wait): + } + if reconcile(ctx) == nil { + slog.Info("dedupe: dynamodb table check passed", "table", c.Table) + } + } + return nil + }}) } return nil } diff --git a/internal/config/backends.go b/internal/config/backends.go index 9b0cb57a5..0b586d57e 100644 --- a/internal/config/backends.go +++ b/internal/config/backends.go @@ -76,8 +76,8 @@ type Dedupe struct { // Lease is how long a claimed id stays pending while its record is // published; a claim its request never settles lapses after it. Lease time.Duration `yaml:"lease" env:"WH_DEDUPE_LEASE" env-default:"30s"` - // ReserveConcurrency bounds the parallel calls one request makes to a - // remote backend. Pebble ignores it. + // ReserveConcurrency bounds the parallel calls one Reserve, Commit or + // Release makes to a remote backend. Pebble ignores it. ReserveConcurrency int `yaml:"reserve_concurrency" env:"WH_DEDUPE_RESERVE_CONCURRENCY" env-default:"64"` DynamoDB DedupeDynamoDBConfig `yaml:"dynamodb"` } diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index eb1c83838..f07904af2 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -139,6 +139,9 @@ func NewDynamo(ctx context.Context, cfg DynamoConfig, extra ...func(*config.Load if err != nil { return nil, fmt.Errorf("dedupe: aws config: %w", err) } + if awsCfg.Region == "" { + return nil, errors.New("dedupe: dynamodb region is not set: set dedupe.dynamodb.region or AWS_REGION") + } client := dynamodb.NewFromConfig(awsCfg, func(o *dynamodb.Options) { if cfg.Endpoint != "" { o.BaseEndpoint = aws.String(cfg.Endpoint) From 0b604510236fa0eb19dea64a2fcde189f220452b Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 04:20:25 -0400 Subject: [PATCH 09/38] docs(dedupe): name every region source; untangle the dynamodb check clause Co-Authored-By: Claude Opus 5.5 (1M context) --- docs/src/content/docs/deployment.md | 2 +- docs/src/content/docs/settings-directory.mdx | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 3919bec49..b889008a4 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -493,7 +493,7 @@ dedupe: region: us-east-1 # or leave empty for AWS_REGION ``` -or `WH_DEDUPE_BACKEND=dynamodb`, `WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod`. A table that is missing, has the wrong key schema, or cannot be reached with the pod's credentials refuses boot over a flat settings directory; over a nested one the pod boots, every tenant with dedupe on fails its ingest closed, and the check is retried in the background (backing off from one second to thirty) and on every reload. No region at all (neither `region` nor `AWS_REGION`) refuses boot in both shapes. The check runs whether or not any tenant has `dedupe.enabled` on. The per-tenant switch stays in each tenant's `config.json`. +or `WH_DEDUPE_BACKEND=dynamodb`, `WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod`. A table that is missing, has the wrong key schema, or cannot be reached with the pod's credentials refuses boot over a flat settings directory; over a nested one the pod boots, every tenant with dedupe on fails its ingest closed, and the check is retried in the background (backing off from one second to thirty) and on every reload. No region at all (neither `region` nor one from the SDK chain: `AWS_REGION`, `AWS_DEFAULT_REGION` or a profile) refuses boot in both shapes. The check runs whether or not any tenant has `dedupe.enabled` on. The per-tenant switch stays in each tenant's `config.json`. For development against dynamodb-local, set `dedupe.dynamodb.endpoint` (for example `http://localhost:8000`) and `create_table: true`, and give the SDK any static credentials (`AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`) and a region. `create_table` without an `endpoint` refuses boot. diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index c7377620c..524c39dcb 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -185,7 +185,7 @@ What stays in boot config is only what cannot change under a running process — Every per-tenant dedupe knob lives here. Where the seen ids are kept (`dedupe.backend`) and how long a claim is held (`dedupe.lease`) are [boot config](/configuration#dedupe), the same for every tenant. The switch and its fields are resolved per record from one snapshot (table override → global value): -- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store (in the embedded Pebble instance at `/pebble`, or its share of the DynamoDB table under `dedupe.backend: dynamodb`), so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) with `dedupe.backend: pebble`, every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. Under `dedupe.backend: dynamodb` the table check plays the instance's part: a table that fails it fails every tenant with dedupe on closed, whatever the switches say, until the check, retried in the background and on every reload, passes ([Configuration](/configuration#dynamodb-dedupe)). +- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store (in the embedded Pebble instance at `/pebble`, or its share of the DynamoDB table under `dedupe.backend: dynamodb`), so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) with `dedupe.backend: pebble`, every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. Under `dedupe.backend: dynamodb` the table check plays the instance's part: the table is checked whether or not any tenant's switch is on, and a table that fails it fails every tenant with dedupe on closed until the check, retried in the background and on every reload, passes ([Configuration](/configuration#dynamodb-dedupe)). - `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its lease (`dedupe.lease`, 30 seconds by default), and a later retry of it is accepted again — counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. - `dedupe.tables.
.{id_field, require_id}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. From 02d7056bf76c42aa712758aaffa718a4785bfed5 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 11:25:11 -0400 Subject: [PATCH 10/38] docs(deployment): use neutral tags in the DynamoDB Terraform example Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- docs/src/content/docs/deployment.md | 9 ++++----- 1 file changed, 4 insertions(+), 5 deletions(-) diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 9bea802e7..c7e88c2cc 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -440,7 +440,7 @@ What the backend requires of the table: Only `pk` is declared in the table definition. Turn TTL on for `ex`. Correctness never depends on TTL, because a claim whose `ex` has passed counts as absent whether or not DynamoDB has deleted it yet; TTL only reclaims the storage. **Today TTL removes only lapsed claims:** ingest commits every id with no retention, so a committed item carries no `ex` and is kept forever, and the table grows by one item (about 200 bytes) per distinct id. Per-tenant retention is [#220](https://github.com/Wave-RF/WaveHouse/issues/220). The backend's table check, which boot will run once the backend is selectable, refuses a table whose key schema does not match and logs a warning if TTL is off. -An example in Terraform. Its tags are the five that Wave RF's own deployments put on every AWS resource (`Name`, `Project`, `Environment`, `ManagedBy`, `CostCenter`, with lowercase-kebab values); use your own conventions in their place: +An example in Terraform. Replace the tags with your own conventions: ```hcl resource "aws_dynamodb_table" "wavehouse_dedupe" { @@ -465,10 +465,9 @@ resource "aws_dynamodb_table" "wavehouse_dedupe" { tags = { Name = "wavehouse-dedupe-${var.environment}" - Project = "wavehouse-cloud" - Environment = var.environment # prod | dev | ci | demo | benchmark - ManagedBy = "wavehouse-cloud/infra/stacks/prod-platform" - CostCenter = "data-plane" + Project = "wavehouse" + Environment = var.environment + ManagedBy = "terraform" } } From cfdc06cd8fae71fc8167f25d8cff63195bb47c52 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 13:13:12 -0400 Subject: [PATCH 11/38] feat(dedupe): key the DynamoDB table by a String pk The dedupe key is now ASCII text, so the table's partition key is a String: it reads as-is in the console and in get-item output, with the same 2,048-byte limit. Check() now requires pk to be S, CreateTable declares it S, and the Deployment table and Terraform example follow. Part of #613. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/deployment.md | 4 ++-- internal/dedupe/dynamodb.go | 20 ++++++++++---------- internal/dedupe/dynamodb_test.go | 16 ++++++++-------- tests/integration/dedupe_dynamodb_test.go | 12 ++++++------ 6 files changed, 28 insertions(+), 28 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index a0062f44b..e79d07b98 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. `Reserve` is a conditional `PutItem` that is atomic across pods. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. +- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt%2D123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 6a63f13d8..8995a4bda 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -126,7 +126,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key every backend stores, as text: `/
/` (for example `acme/clicks/evt%2D123`, [#222](https://github.com/Wave-RF/WaveHouse/issues/222)). The table and id are escaped by `internal/keyenc` — the escaping NATS subject tokens use — which never writes `/`, and a tenant id cannot hold one, so a table name may hold any byte, NUL included, and neither tenants nor tables share ids; a key is ASCII, so it reads as-is in a console and is a valid DynamoDB String. An id whose escaped form is over 1,024 bytes is stored as `#` plus its SHA-256 in hex (`#` is never written by the escaping), counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. -- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (binary) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. +- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also reads a lease `<= 0` as `DefaultLease` and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index c7e88c2cc..a612e3ad5 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -433,7 +433,7 @@ What the backend requires of the table: | Attribute | Type | Role | |---|---|---| -| `pk` | Binary | Partition key, and the only key: tenant, table and id. No sort key. | +| `pk` | String | Partition key, and the only key: tenant, table and id as readable text, for example `acme/clicks/evt%2D123` (the table and id escaped the way NATS subject tokens are). No sort key. | | `st` | Number | `1` = pending claim, `2` = committed. | | `ex` | Number | Epoch seconds: the lease end while pending, the retention end once committed; absent = never expires. | | `tk` | Binary | The claim token that `Release` matches. | @@ -451,7 +451,7 @@ resource "aws_dynamodb_table" "wavehouse_dedupe" { attribute { name = "pk" - type = "B" + type = "S" } ttl { diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index eb1c83838..bb39c060b 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -187,7 +187,7 @@ func (d *Dynamo) Tenant(id tenant.ID) *Managed { } // Check verifies the table exists with the key schema this backend writes: -// pk, binary, as the only key. TTL not enabled on ex is logged, not refused: +// pk, a string, as the only key. TTL not enabled on ex is logged, not refused: // expiry never depends on it, only storage does. func (d *Dynamo) Check(ctx context.Context) error { ctx, cancel := context.WithTimeout(ctx, 10*d.cfg.Timeout) @@ -201,8 +201,8 @@ func (d *Dynamo) Check(ctx context.Context) error { return fmt.Errorf("dedupe: dynamodb table %s: key schema must be %s (HASH) alone", d.cfg.Table, attrKey) } for _, a := range t.AttributeDefinitions { - if aws.ToString(a.AttributeName) == attrKey && a.AttributeType != types.ScalarAttributeTypeB { - return fmt.Errorf("dedupe: dynamodb table %s: %s must be binary (B), is %s", d.cfg.Table, attrKey, a.AttributeType) + if aws.ToString(a.AttributeName) == attrKey && a.AttributeType != types.ScalarAttributeTypeS { + return fmt.Errorf("dedupe: dynamodb table %s: %s must be a string (S), is %s", d.cfg.Table, attrKey, a.AttributeType) } } ttl, err := d.api.DescribeTimeToLive(ctx, &dynamodb.DescribeTimeToLiveInput{TableName: &d.cfg.Table}) @@ -229,7 +229,7 @@ func (d *Dynamo) CreateTable(ctx context.Context) error { _, err := d.api.CreateTable(ctx, &dynamodb.CreateTableInput{ TableName: &d.cfg.Table, BillingMode: types.BillingModePayPerRequest, - AttributeDefinitions: []types.AttributeDefinition{{AttributeName: aws.String(attrKey), AttributeType: types.ScalarAttributeTypeB}}, + AttributeDefinitions: []types.AttributeDefinition{{AttributeName: aws.String(attrKey), AttributeType: types.ScalarAttributeTypeS}}, KeySchema: []types.KeySchemaElement{{AttributeName: aws.String(attrKey), KeyType: types.KeyTypeHash}}, }) var inUse *types.ResourceInUseException @@ -332,7 +332,7 @@ func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Durati func (s *dynamoStore) reserve(ctx context.Context, k Key, token, nowSec string, exp int64) (Status, error) { item := map[string]types.AttributeValue{ - attrKey: &types.AttributeValueMemberB{Value: AppendKey(nil, s.prefix, k)}, + attrKey: &types.AttributeValueMemberS{Value: string(AppendKey(nil, s.prefix, k))}, attrState: &types.AttributeValueMemberN{Value: statePending}, attrExpiry: &types.AttributeValueMemberN{Value: strconv.FormatInt(exp, 10)}, attrToken: &types.AttributeValueMemberB{Value: []byte(token)}, @@ -375,13 +375,13 @@ func (s *dynamoStore) Commit(ctx context.Context, claims []Claim, retention time seen := make(map[string]bool, len(claims)) writes := make([]types.WriteRequest, 0, len(claims)) for _, c := range claims { - pk := AppendKey(nil, s.prefix, c.Key) - if seen[string(pk)] { + pk := string(AppendKey(nil, s.prefix, c.Key)) + if seen[pk] { continue } - seen[string(pk)] = true + seen[pk] = true item := map[string]types.AttributeValue{ - attrKey: &types.AttributeValueMemberB{Value: pk}, + attrKey: &types.AttributeValueMemberS{Value: pk}, attrState: &types.AttributeValueMemberN{Value: stateCommitted}, attrToken: &types.AttributeValueMemberB{Value: []byte(c.Token)}, } @@ -442,7 +442,7 @@ func (s *dynamoStore) Release(ctx context.Context, claims []Claim) error { err := s.d.call(ctx, "delete_item", func(ctx context.Context) error { _, err := s.d.api.DeleteItem(ctx, &dynamodb.DeleteItemInput{ TableName: &s.d.cfg.Table, - Key: map[string]types.AttributeValue{attrKey: &types.AttributeValueMemberB{Value: AppendKey(nil, s.prefix, c.Key)}}, + Key: map[string]types.AttributeValue{attrKey: &types.AttributeValueMemberS{Value: string(AppendKey(nil, s.prefix, c.Key))}}, ConditionExpression: aws.String(condRelease), ExpressionAttributeValues: map[string]types.AttributeValue{ ":tk": &types.AttributeValueMemberB{Value: []byte(c.Token)}, diff --git a/internal/dedupe/dynamodb_test.go b/internal/dedupe/dynamodb_test.go index dc74d9d1d..d2f5f740b 100644 --- a/internal/dedupe/dynamodb_test.go +++ b/internal/dedupe/dynamodb_test.go @@ -137,7 +137,7 @@ func TestClassify(t *testing.T) { func TestDynamo_ReserveReadsTheHeldItem(t *testing.T) { t.Parallel() _, m := openFake(t, &fakeDynamo{put: func(_ context.Context, in *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { - id := string(in.Item[attrKey].(*types.AttributeValueMemberB).Value) + id := in.Item[attrKey].(*types.AttributeValueMemberS).Value switch id[len(id)-1] { case 'd': return nil, &types.ConditionalCheckFailedException{Item: map[string]types.AttributeValue{attrState: &types.AttributeValueMemberN{Value: stateCommitted}}} @@ -162,7 +162,7 @@ func TestDynamo_FailedReserveReleasesEveryPutThatMayHaveLanded(t *testing.T) { var released []string _, m := openFake(t, &fakeDynamo{ put: func(_ context.Context, in *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { - id := string(in.Item[attrKey].(*types.AttributeValueMemberB).Value) + id := in.Item[attrKey].(*types.AttributeValueMemberS).Value mu.Lock() putTokens[id] = string(in.Item[attrToken].(*types.AttributeValueMemberB).Value) mu.Unlock() @@ -175,7 +175,7 @@ func TestDynamo_FailedReserveReleasesEveryPutThatMayHaveLanded(t *testing.T) { return &dynamodb.PutItemOutput{}, nil }, del: func(_ context.Context, in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { - id := string(in.Key[attrKey].(*types.AttributeValueMemberB).Value) + id := in.Key[attrKey].(*types.AttributeValueMemberS).Value mu.Lock() defer mu.Unlock() assert.Equal(t, putTokens[id], string(in.ExpressionAttributeValues[":tk"].(*types.AttributeValueMemberB).Value), "released by the token it was put with") @@ -210,7 +210,7 @@ func TestDynamo_CommitRetriesUnprocessedItems(t *testing.T) { defer mu.Unlock() var left []types.WriteRequest for i, r := range reqs { - pk := string(r.PutRequest.Item[attrKey].(*types.AttributeValueMemberB).Value) + pk := r.PutRequest.Item[attrKey].(*types.AttributeValueMemberS).Value assert.Equal(t, stateCommitted, r.PutRequest.Item[attrState].(*types.AttributeValueMemberN).Value) assert.Contains(t, r.PutRequest.Item, attrExpiry) if i == len(reqs)-1 && !heldBack[pk] && len(reqs) > 1 { @@ -252,7 +252,7 @@ func TestDynamo_ReleaseTreatsAFailedConditionAsDone(t *testing.T) { t.Parallel() _, m := openFake(t, &fakeDynamo{del: func(_ context.Context, in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { assert.Equal(t, condRelease, aws.ToString(in.ConditionExpression)) - id := string(in.Key[attrKey].(*types.AttributeValueMemberB).Value) + id := in.Key[attrKey].(*types.AttributeValueMemberS).Value if id[len(id)-1] == 'x' { return nil, &types.ResourceNotFoundException{} } @@ -320,7 +320,7 @@ func TestDynamo_Check(t *testing.T) { t.Parallel() good := &dynamodb.DescribeTableOutput{Table: &types.TableDescription{ KeySchema: []types.KeySchemaElement{{AttributeName: aws.String("pk"), KeyType: types.KeyTypeHash}}, - AttributeDefinitions: []types.AttributeDefinition{{AttributeName: aws.String("pk"), AttributeType: types.ScalarAttributeTypeB}}, + AttributeDefinitions: []types.AttributeDefinition{{AttributeName: aws.String("pk"), AttributeType: types.ScalarAttributeTypeS}}, }} ttlOn := &dynamodb.DescribeTimeToLiveOutput{TimeToLiveDescription: &types.TimeToLiveDescription{ AttributeName: aws.String("ex"), TimeToLiveStatus: types.TimeToLiveStatusEnabled, @@ -375,8 +375,8 @@ func TestExpiresAt(t *testing.T) { } func idOf(av types.AttributeValue) string { - b := av.(*types.AttributeValueMemberB).Value - return string(b[len(b)-2:]) + s := av.(*types.AttributeValueMemberS).Value + return s[len(s)-2:] } func TestDynamo_CommitAttemptsEveryChunk(t *testing.T) { diff --git a/tests/integration/dedupe_dynamodb_test.go b/tests/integration/dedupe_dynamodb_test.go index 192734a97..79bea18ab 100644 --- a/tests/integration/dedupe_dynamodb_test.go +++ b/tests/integration/dedupe_dynamodb_test.go @@ -222,11 +222,11 @@ func TestDedupeDynamo_Check(t *testing.T) { _, err := raw.CreateTable(t.Context(), &dynamodb.CreateTableInput{ TableName: aws.String(table), BillingMode: types.BillingModePayPerRequest, - AttributeDefinitions: []types.AttributeDefinition{{AttributeName: aws.String("pk"), AttributeType: types.ScalarAttributeTypeS}}, + AttributeDefinitions: []types.AttributeDefinition{{AttributeName: aws.String("pk"), AttributeType: types.ScalarAttributeTypeB}}, KeySchema: []types.KeySchemaElement{{AttributeName: aws.String("pk"), KeyType: types.KeyTypeHash}}, }) require.NoError(t, err) - assert.ErrorContains(t, dynamoClient(t, table, dedupe.DynamoConfig{}).Check(t.Context()), "must be binary") + assert.ErrorContains(t, dynamoClient(t, table, dedupe.DynamoConfig{}).Check(t.Context()), "must be a string") fresh := newDynamoTable() d := dynamoClient(t, fresh, dedupe.DynamoConfig{}) @@ -248,13 +248,13 @@ func TestDedupeDynamo_Expiry(t *testing.T) { require.NoError(t, d.CreateTable(t.Context())) m := d.Tenant("acme") require.NoError(t, m.Apply(true)) - pk := func(id string) []byte { - return dedupe.AppendKey(nil, dedupe.KeyPrefix("acme"), dedupe.Key{Table: "events", ID: id}) + pk := func(id string) string { + return string(dedupe.AppendKey(nil, dedupe.KeyPrefix("acme"), dedupe.Key{Table: "events", ID: id})) } item := func(id string) map[string]types.AttributeValue { out, err := raw.GetItem(t.Context(), &dynamodb.GetItemInput{ TableName: aws.String(table), ConsistentRead: aws.Bool(true), - Key: map[string]types.AttributeValue{"pk": &types.AttributeValueMemberB{Value: pk(id)}}, + Key: map[string]types.AttributeValue{"pk": &types.AttributeValueMemberS{Value: pk(id)}}, }) require.NoError(t, err) return out.Item @@ -280,7 +280,7 @@ func TestDedupeDynamo_Expiry(t *testing.T) { // TTL deletes lazily; an item whose ex has passed is absent all the same. for _, st := range []string{"1", "2"} { _, err = raw.PutItem(t.Context(), &dynamodb.PutItemInput{TableName: aws.String(table), Item: map[string]types.AttributeValue{ - "pk": &types.AttributeValueMemberB{Value: pk("stale-" + st)}, + "pk": &types.AttributeValueMemberS{Value: pk("stale-" + st)}, "st": &types.AttributeValueMemberN{Value: st}, "ex": &types.AttributeValueMemberN{Value: strconv.FormatInt(time.Now().Add(-time.Minute).Unix(), 10)}, "tk": &types.AttributeValueMemberB{Value: []byte("old")}, From 7082907a576486b2948aad21b2c029d08a17e92d Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 13:44:33 -0400 Subject: [PATCH 12/38] fix(dedupe): let a sent DynamoDB put answer before its release Reserve sent every put on the errgroup's context, so the first failure cancelled siblings already on the wire. The client gave up on them but the table could still apply one after the undo's conditional delete had found nothing, holding the key for the whole lease: the conformance case "a failed reserve leaves nothing claimed" failed about one run in ten against dynamodb-local. A sent put now runs on the caller's context; the group's context only skips the puts not yet sent. 30 consecutive conformance runs against dynamodb-local pass. Part of #613. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- internal/dedupe/dynamodb.go | 9 +++++--- internal/dedupe/dynamodb_test.go | 35 +++++++++++++++++++------------- 2 files changed, 27 insertions(+), 17 deletions(-) diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index bb39c060b..5b444d5ec 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -292,8 +292,11 @@ func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Durati claims := make([]Claim, len(keys)) tried := make([]Claim, len(keys)) sent := make([]bool, len(keys)) - // The first failure cancels the puts not yet sent: the Reserve fails - // either way, and a throttled table should not take the rest. + // The first failure skips the puts not yet sent: the Reserve fails + // either way, and a throttled table should not take the rest. A put + // already sent runs on ctx, not gctx, so it finishes and its outcome is + // known before the undo below; cancelled mid-flight, it could land after + // its release and hold the key for the lease. g, gctx := errgroup.WithContext(ctx) g.SetLimit(s.d.cfg.ReserveConcurrency) for i, k := range keys { @@ -304,7 +307,7 @@ func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Durati return err } sent[i] = true - status, err := s.reserve(gctx, k, token, nowSec, exp) + status, err := s.reserve(ctx, k, token, nowSec, exp) if err != nil { return err } diff --git a/internal/dedupe/dynamodb_test.go b/internal/dedupe/dynamodb_test.go index d2f5f740b..685cb7f78 100644 --- a/internal/dedupe/dynamodb_test.go +++ b/internal/dedupe/dynamodb_test.go @@ -419,12 +419,14 @@ func TestDynamo_ReleaseAttemptsEveryClaim(t *testing.T) { assert.Equal(t, int64(4), deletes.Load()) } -// One throttled put in a multi-key Reserve: the unsent puts are neither sent -// nor released, and the cancelled siblings do not reset the breaker. +// One throttled put in a multi-key Reserve: the unsent puts are never sent, +// and a sibling already sent runs to its answer before the undo releases it, +// so a put cannot land after its own release. func TestDynamo_FailedMultiKeyReserve(t *testing.T) { t.Parallel() var mu sync.Mutex var put, released []string + landed := map[string]bool{} _, m := openFakeWith(t, &fakeDynamo{ put: func(ctx context.Context, in *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { id := idOf(in.Item[attrKey]) @@ -434,26 +436,31 @@ func TestDynamo_FailedMultiKeyReserve(t *testing.T) { if id == "k0" { return nil, &types.ProvisionedThroughputExceededException{} } - <-ctx.Done() - return nil, ctx.Err() + time.Sleep(20 * time.Millisecond) + if err := ctx.Err(); err != nil { + return nil, err + } + mu.Lock() + landed[id] = true + mu.Unlock() + return &dynamodb.PutItemOutput{}, nil }, del: func(_ context.Context, in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { + id := idOf(in.Key[attrKey]) mu.Lock() - released = append(released, idOf(in.Key[attrKey])) - mu.Unlock() - return nil, &types.ConditionalCheckFailedException{} + defer mu.Unlock() + released = append(released, id) + if id != "k0" && !landed[id] { + t.Errorf("%s released before its put answered", id) + } + return &dynamodb.DeleteItemOutput{}, nil }, }, DynamoConfig{Table: "dedupe", ReserveConcurrency: 2}) ks := keys("k0", "k1", "k2", "k3", "k4", "k5", "k6", "k7") - for range breakerTrips { - _, err := m.Reserve(t.Context(), ks, time.Minute) - require.ErrorIs(t, err, ErrUnavailable) - require.NotErrorIs(t, err, errBreakerOpen) - } _, err := m.Reserve(t.Context(), ks, time.Minute) - require.ErrorIs(t, err, errBreakerOpen, "the cancelled siblings did not reset the count") + require.ErrorIs(t, err, ErrUnavailable) mu.Lock() defer mu.Unlock() assert.ElementsMatch(t, put, released, "exactly the sent puts are released") - assert.Less(t, len(put), breakerTrips*len(ks), "unsent puts were never sent") + assert.Less(t, len(put), len(ks), "unsent puts were never sent") } From 484c72e056a876aca9118f1a804d343715c36ab9 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 13:45:08 -0400 Subject: [PATCH 13/38] docs(changelog): the boot-backends entry predates dedupe.dynamodb Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- CHANGELOG.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 65f18ee91..3e5571400 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -12,7 +12,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), - **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/backends.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `2m`, the embedded queue's duplicate window), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`; no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and on every reload. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. - **A DynamoDB dedupe backend** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt%2D123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. `dedupe.backend: dynamodb` selects it (entry above). New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. -- **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer has only its in-process backend so far, and it is the default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings will go in a `.` sub-block; no backend has settings yet, so every such sub-block is an unknown key for now and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name the `mq`, `cache` and `dedupe` backends: the zero value is not the default, and `app.New` refuses it. `coord.backend` is reserved: nothing reads it until the lease layer lands, and the sweeper still runs in every process. +- **Each layer's implementation is chosen at boot** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `cmd/wavehouse/main.go`, `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,architecture.md}`): the first step of running WaveHouse as more than one process ([#613](https://github.com/Wave-RF/WaveHouse/issues/613)). The boot config gains `mq.backend` (`WH_MQ_BACKEND`, default `embedded`), `cache.backend` (`WH_CACHE_BACKEND`, `local`), `dedupe.backend` (`WH_DEDUPE_BACKEND`, `pebble`) and `coord.backend` (`WH_COORD_BACKEND`, `local`). Each layer's in-process backend is its default, so nothing changes for a config that sets none of them; a value with no backend refuses boot and names the valid ones. A backend's own settings go in a `.` sub-block (`dedupe.dynamodb` is the first, entry above); any other sub-block is an unknown key and refuses boot. `internal/app` picks each implementation in one `switch` per layer (`wireMQ`, `wireCache`, `wireDedupe`), `data_dir` is probed only when a selected backend keeps state there (`Config.NeedsDataDir`), and boot logs at `WARN` each line of `Config.Warnings`, the combinations that are correct for one replica only once a shared queue exists. A `config.Config` built without `config.Load` must now name the `mq`, `cache` and `dedupe` backends: the zero value is not the default, and `app.New` refuses it. `coord.backend` is reserved: nothing reads it until the lease layer lands, and the sweeper still runs in every process. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. From 4c2c2a407f52b4386c748d0d2d50a6a4b8dd788f Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Fri, 25 Sep 2026 13:49:45 -0400 Subject: [PATCH 14/38] test(dedupe): a caller-cancelled put does not reset the breaker The guard's comment named sibling cancellation, which no longer happens; it now names the caller going away, and a test pins it (removing the guard fails it). Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01EJr5tY4WQUy2sc4MbW67vL --- internal/dedupe/dynamodb.go | 5 +++-- internal/dedupe/dynamodb_test.go | 30 ++++++++++++++++++++++++++++++ 2 files changed, 33 insertions(+), 2 deletions(-) diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index 5b444d5ec..d298a8c70 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -260,8 +260,9 @@ func (d *Dynamo) call(ctx context.Context, op string, do func(context.Context) e start := time.Now() err := classify(op, do(ctx)) d.metrics.record(ctx, op, time.Since(start), err) - // A request cancelled because a sibling failed says nothing about the - // table, and must not reset the breaker's count. + // A request cancelled because its caller went away (a client + // disconnecting mid-Reserve) says nothing about the table, and must not + // reset the breaker's count. if op == opReserve && !errors.Is(err, context.Canceled) { d.breaker.record(err) } diff --git a/internal/dedupe/dynamodb_test.go b/internal/dedupe/dynamodb_test.go index 685cb7f78..88bc3f768 100644 --- a/internal/dedupe/dynamodb_test.go +++ b/internal/dedupe/dynamodb_test.go @@ -401,6 +401,36 @@ func TestDynamo_CommitAttemptsEveryChunk(t *testing.T) { assert.Equal(t, 2*batchWriteMax, written, "a failed chunk does not cancel the others: their records are published") } +// A put cancelled by its caller is not an answer from the table: it does +// not reset the breaker's count of throttled puts. +func TestDynamo_CallerCancelDoesNotResetBreaker(t *testing.T) { + t.Parallel() + var hang atomic.Bool + started := make(chan struct{}, 1) + _, m := openFake(t, &fakeDynamo{put: func(ctx context.Context, _ *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { + if hang.Load() { + started <- struct{}{} + <-ctx.Done() + return nil, ctx.Err() + } + return nil, &types.ProvisionedThroughputExceededException{} + }}) + for range breakerTrips - 1 { + _, err := m.Reserve(t.Context(), keys("a"), time.Minute) + require.ErrorIs(t, err, ErrUnavailable) + } + hang.Store(true) + ctx, cancel := context.WithCancel(t.Context()) + go func() { <-started; cancel() }() + _, err := m.Reserve(ctx, keys("a"), time.Minute) + require.ErrorIs(t, err, context.Canceled) + hang.Store(false) + _, err = m.Reserve(t.Context(), keys("a"), time.Minute) + require.ErrorIs(t, err, ErrUnavailable) + _, err = m.Reserve(t.Context(), keys("a"), time.Minute) + require.ErrorIs(t, err, errBreakerOpen, "the cancelled put did not reset the count") +} + func TestDynamo_ReleaseAttemptsEveryClaim(t *testing.T) { t.Parallel() var deletes atomic.Int64 From 3248e7eadaf786bf9144dc1b40bf6570564911b0 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 04:47:57 -0400 Subject: [PATCH 15/38] fix(dedupe): a retried DynamoDB put keeps its own claim When the SDK retries a conditional PutItem whose first attempt was applied (a 500 or a reset connection after the write), the retry fails its condition on the caller's own item. Reserve read that as another request's live claim and answered InFlight: nobody would commit or release the item, the id stayed locked for the lease, and the producer got a spurious 503. A pending item carrying the put's own token now answers Claimed, so the caller commits it, or a failed Reserve's undo releases it. Pinned by a unit test (the fake hands back the put's own item) and an integration test that forwards the first PutItem to dynamodb-local and answers it 500. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- internal/dedupe/dynamodb.go | 18 +++++++-- internal/dedupe/dynamodb_test.go | 48 +++++++++++++++++++++++ tests/integration/dedupe_dynamodb_test.go | 46 ++++++++++++++++++++++ 5 files changed, 111 insertions(+), 5 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index a81bf89a6..792cb65dd 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt%2D123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. +- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt%2D123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 7482a7b61..9fb7dd8f8 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -126,7 +126,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key every backend stores, as text: `/
/` (for example `acme/clicks/evt%2D123`, [#222](https://github.com/Wave-RF/WaveHouse/issues/222)). The table and id are escaped by `internal/keyenc` — the escaping NATS subject tokens use — which never writes `/`, and a tenant id cannot hold one, so a table name may hold any byte, NUL included, and neither tenants nor tables share ids; a key is ASCII, so it reads as-is in a console and is a valid DynamoDB String. An id whose escaped form is over 1,024 bytes is stored as `#` plus its SHA-256 in hex (`#` is never written by the escaping), counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. -- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. +- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read, or `Claimed` when it is the put's own pending item (same token): an SDK retry of an attempt DynamoDB applied but whose answer was lost. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also reads a lease `<= 0` as `DefaultLease` and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index d298a8c70..5e71db3fc 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -277,8 +277,9 @@ type dynamoStore struct { // Reserve puts every key's pending item in parallel, each conditional on no // live item holding the key. A failed condition hands back the live item, -// whose state says Duplicate or InFlight without a read. On any error every -// put that may have landed is released by its token. +// whose state says Duplicate or InFlight without a read, or Claimed when its +// token is the put's own. On any error every put that may have landed is +// released by its token. func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Duration) ([]Claim, error) { if len(keys) == 0 { return []Claim{}, nil @@ -358,12 +359,23 @@ func (s *dynamoStore) reserve(ctx context.Context, k Key, token, nowSec string, } return Claimed, nil } - if st, ok := held.Item[attrState].(*types.AttributeValueMemberN); ok && st.Value == stateCommitted { + st, _ := held.Item[attrState].(*types.AttributeValueMemberN) + switch { + case st != nil && st.Value == stateCommitted: return Duplicate, nil + case st != nil && st.Value == statePending && heldBy(held.Item, token): + // This put's own item: an SDK retry of an attempt that was applied + // but whose answer was lost (a 500, a connection reset). + return Claimed, nil } return InFlight, nil } +func heldBy(item map[string]types.AttributeValue, token string) bool { + tk, ok := item[attrToken].(*types.AttributeValueMemberB) + return ok && string(tk.Value) == token +} + // Commit overwrites every claim's item as committed, unconditionally, 25 to a // BatchWriteItem, retrying the items DynamoDB leaves unprocessed. func (s *dynamoStore) Commit(ctx context.Context, claims []Claim, retention time.Duration) error { diff --git a/internal/dedupe/dynamodb_test.go b/internal/dedupe/dynamodb_test.go index 88bc3f768..65397b271 100644 --- a/internal/dedupe/dynamodb_test.go +++ b/internal/dedupe/dynamodb_test.go @@ -4,6 +4,7 @@ import ( "context" "errors" "fmt" + "maps" "net" "sync" "sync/atomic" @@ -155,6 +156,53 @@ func TestDynamo_ReserveReadsTheHeldItem(t *testing.T) { assert.Empty(t, claims[1].Token) } +// An SDK retry of a put whose first attempt was applied fails its condition +// on the put's own item: that is the caller's claim, not another request's. +func TestDynamo_RetriedPutKeepsItsOwnClaim(t *testing.T) { + t.Parallel() + var mu sync.Mutex + putTokens := map[string]string{} + var released []string + fake := &fakeDynamo{ + put: func(_ context.Context, in *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { + id := idOf(in.Item[attrKey]) + mu.Lock() + putTokens[id] = string(in.Item[attrToken].(*types.AttributeValueMemberB).Value) + mu.Unlock() + switch id { + case "k0": + return nil, &types.ConditionalCheckFailedException{Item: in.Item} + case "k1": + theirs := maps.Clone(in.Item) + theirs[attrToken] = &types.AttributeValueMemberB{Value: []byte("theirs")} + return nil, &types.ConditionalCheckFailedException{Item: theirs} + } + return nil, &types.InternalServerError{} + }, + del: func(_ context.Context, in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { + id := idOf(in.Key[attrKey]) + mu.Lock() + defer mu.Unlock() + assert.Equal(t, putTokens[id], string(in.ExpressionAttributeValues[":tk"].(*types.AttributeValueMemberB).Value)) + released = append(released, id) + return &dynamodb.DeleteItemOutput{}, nil + }, + } + _, m := openFakeWith(t, fake, DynamoConfig{Table: "dedupe", ReserveConcurrency: 1}) + + claims, err := m.Reserve(t.Context(), keys("k0", "k1"), time.Minute) + require.NoError(t, err) + assert.Equal(t, []Status{Claimed, InFlight}, []Status{claims[0].Status, claims[1].Status}, "a pending item is InFlight only under another token") + assert.Equal(t, putTokens["k0"], claims[0].Token) + + // k0 is sent and answered before k2 fails, so the undo owns it. + _, err = m.Reserve(t.Context(), keys("k0", "k2"), time.Minute) + require.ErrorIs(t, err, ErrUnavailable) + mu.Lock() + defer mu.Unlock() + assert.Contains(t, released, "k0", "the failed Reserve's undo releases the retried put's own item") +} + func TestDynamo_FailedReserveReleasesEveryPutThatMayHaveLanded(t *testing.T) { t.Parallel() var mu sync.Mutex diff --git a/tests/integration/dedupe_dynamodb_test.go b/tests/integration/dedupe_dynamodb_test.go index 79bea18ab..69be6b16e 100644 --- a/tests/integration/dedupe_dynamodb_test.go +++ b/tests/integration/dedupe_dynamodb_test.go @@ -86,6 +86,23 @@ func awsError(status int, code string) *http.Response { const putItem = "DynamoDB_20120810.PutItem" +// landThenFail sends the first PutItem, then answers it 500 as if the +// response were lost: the write is applied and the SDK retries it. +type landThenFail struct { + next *http.Client + failed atomic.Bool +} + +func (l *landThenFail) Do(r *http.Request) (*http.Response, error) { + resp, err := l.next.Do(r) + if err != nil || r.Header.Get("X-Amz-Target") != putItem || !l.failed.CompareAndSwap(false, true) { + return resp, err + } + _, _ = io.Copy(io.Discard, resp.Body) + _ = resp.Body.Close() + return awsError(http.StatusInternalServerError, "InternalServerError"), nil +} + func TestDedupeDynamo_Conformance(t *testing.T) { t.Parallel() dedupetest.Run(t, func(t *testing.T) dedupetest.Harness { @@ -183,6 +200,35 @@ func TestDedupeDynamo_Throttled(t *testing.T) { } } +// The SDK's retry of an applied put fails its condition on the put's own +// item, which is still the caller's claim: without that, the id would be +// held InFlight for the lease by a claim nobody commits or releases. +func TestDedupeDynamo_RetriedPutKeepsItsClaim(t *testing.T) { + t.Parallel() + table := newDynamoTable() + lossy := &landThenFail{next: http.DefaultClient} + d := dynamoClient(t, table, dedupe.DynamoConfig{}, config.WithHTTPClient(lossy)) + require.NoError(t, d.CreateTable(t.Context())) + m := d.Tenant("acme") + require.NoError(t, m.Apply(true)) + peer := dynamoClient(t, table, dedupe.DynamoConfig{}).Tenant("acme") + require.NoError(t, peer.Apply(true)) + k := []dedupe.Key{{Table: "events", ID: "e1"}} + + claims, err := m.Reserve(t.Context(), k, time.Minute) + require.NoError(t, err) + require.True(t, lossy.failed.Load(), "the applied attempt was answered 500") + require.Equal(t, dedupe.Claimed, claims[0].Status) + other, err := peer.Reserve(t.Context(), k, time.Minute) + require.NoError(t, err) + assert.Equal(t, dedupe.InFlight, other[0].Status) + + require.NoError(t, m.Release(t.Context(), claims)) + other, err = peer.Reserve(t.Context(), k, time.Minute) + require.NoError(t, err) + assert.Equal(t, dedupe.Claimed, other[0].Status, "the claim's token was the applied put's, so Release freed the id") +} + func TestDedupeDynamo_Unreachable(t *testing.T) { t.Parallel() d, err := dedupe.NewDynamo(t.Context(), dedupe.DynamoConfig{ From 88958638bfa4cdcf4194a75dd6062845a8eb3b5a Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 04:51:20 -0400 Subject: [PATCH 16/38] fix(dedupe): jittered DynamoDB retries that fit the call timeout With the retryer's MaxBackoff at 200 ms, the SDK's legacy backoff waited a fixed 200 ms before every retry, with no jitter: under the 250 ms default per-call timeout the third attempt was never sent, a throttled call ended as DeadlineExceeded with the throttle cause dropped, every put throttled together retried in lockstep, and a throttled Reserve took about twice the timeout once its undo hit the same schedule. The retryer now uses an explicit full-jitter backoff, uniform over [0, min(25 ms * 2^attempt, ceiling)], with the ceiling derived from the config as Timeout / (2 * (MaxAttempts - 1)). A call's retries therefore wait at most half its Timeout in all, whatever Timeout and MaxAttempts are set to, and a throttled call ends on its last attempt's answer: a max-attempts error carrying the throttle, mapped to ErrUnavailable. Commit's wait between rounds of unprocessed items is jittered the same way (25 ms doubling to 200 ms). Tests: the backoff's worst case fits half the Timeout across several configs and is jittered; a put throttled through the real SDK stack at the defaults ends as a MaxAttemptsError wrapping the throttle, inside the Timeout, after exactly MaxAttempts sends. Both failed on the old schedule. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- internal/dedupe/dynamodb.go | 37 +++++++- internal/dedupe/dynamodb_test.go | 121 ++++++++++++++++++++++++++ 4 files changed, 156 insertions(+), 6 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 792cb65dd..babb8f4e9 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt%2D123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. +- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt%2D123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 9fb7dd8f8..aeaed95aa 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -126,7 +126,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key every backend stores, as text: `/
/` (for example `acme/clicks/evt%2D123`, [#222](https://github.com/Wave-RF/WaveHouse/issues/222)). The table and id are escaped by `internal/keyenc` — the escaping NATS subject tokens use — which never writes `/`, and a tenant id cannot hold one, so a table name may hold any byte, NUL included, and neither tenants nor tables share ids; a key is ASCII, so it reads as-is in a console and is a valid DynamoDB String. An id whose escaped form is over 1,024 bytes is stored as `#` plus its SHA-256 in hex (`#` is never written by the escaping), counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. -- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read, or `Claimed` when it is the put's own pending item (same token): an SDK retry of an attempt DynamoDB applied but whose answer was lost. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. +- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read, or `Claimed` when it is the put's own pending item (same token): an SDK retry of an attempt DynamoDB applied but whose answer was lost. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with jittered backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Each call has a `Timeout` (250 ms) covering the SDK's retries (`MaxAttempts`, 3), which back off with full jitter under a ceiling capped at `Timeout/(2·(MaxAttempts−1))`, so a call's retries wait at most half its timeout and a throttled call fails on its last attempt's answer rather than on the deadline. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also reads a lease `<= 0` as `DefaultLease` and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index 5e71db3fc..bd9a74858 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -6,6 +6,7 @@ import ( "errors" "fmt" "log/slog" + mathrand "math/rand/v2" "strconv" "sync" "time" @@ -45,7 +46,12 @@ const ( // commitRounds bounds the BatchWriteItem rounds one chunk gets before // its still-unprocessed items fail the Commit. commitRounds = 8 - tokenBytes = 16 + // commitBase and commitCeiling bound the jittered wait between Commit + // rounds; retryBase is the one the SDK retryer's backoff doubles from. + commitBase = 25 * time.Millisecond + commitCeiling = 200 * time.Millisecond + retryBase = 25 * time.Millisecond + tokenBytes = 16 // opReserve is the operation the breaker watches: Release and Commit // answers say nothing about whether a new Reserve would get through. @@ -64,7 +70,8 @@ type DynamoConfig struct { // only; it is also what unlocks CreateTable. Endpoint string // Timeout bounds each DynamoDB call, its SDK retries included. - // 0 = 250ms. + // 0 = 250ms. The retries' jittered backoff is capped so that together + // it waits at most half of Timeout (retryBackoff). Timeout time.Duration // MaxAttempts is the SDK retryer's attempts per call. 0 = 3. MaxAttempts int @@ -150,7 +157,7 @@ func NewDynamo(ctx context.Context, cfg DynamoConfig, extra ...func(*config.Load func newRetryer(cfg DynamoConfig) (func() aws.Retryer, error) { standard := func(o *retry.StandardOptions) { o.MaxAttempts = cfg.MaxAttempts - o.MaxBackoff = 200 * time.Millisecond + o.Backoff = retryBackoff(cfg) } switch cfg.RetryMode { case "standard": @@ -165,6 +172,28 @@ func newRetryer(cfg DynamoConfig) (func() aws.Retryer, error) { return nil, fmt.Errorf("dedupe: dynamodb retry_mode %q: want standard or adaptive", cfg.RetryMode) } +// retryBackoff is the SDK retryer's wait before a retry: full jitter, so +// puts throttled together do not retry in lockstep, under a ceiling that +// doubles from retryBase up to Timeout/(2·(MaxAttempts-1)). A call's retries +// then wait at most half its Timeout in all, so a throttled call ends on its +// last attempt's answer (ErrUnavailable, the throttle as its cause) unless +// the attempts themselves take the other half. +func retryBackoff(cfg DynamoConfig) retry.BackoffDelayerFunc { + ceiling := cfg.Timeout / time.Duration(2*max(cfg.MaxAttempts-1, 1)) + return func(attempt int, _ error) (time.Duration, error) { + return fullJitter(retryBase, ceiling, attempt), nil + } +} + +// fullJitter is uniform over [0, min(base·2^attempt, ceiling)]. +func fullJitter(base, ceiling time.Duration, attempt int) time.Duration { + d := min(base< Date: Sat, 26 Sep 2026 04:52:51 -0400 Subject: [PATCH 17/38] fix(dedupe): size the DynamoDB client's idle pool to its fan-out Reserve, Commit and Release fan out to ReserveConcurrency (64) calls at once, but the SDK's default transport keeps only 10 idle connections per host, so most of a wide Reserve's puts dialed a new connection every time. Against dynamodb-local, a warm 64-key Reserve opened 46 to 54 new connections; with the pool sized it reuses all 64. NewDynamo now passes a BuildableClient whose transport keeps ReserveConcurrency idle connections per host (and at least as many in all). It goes ahead of the caller's extra options, so a test's fault-injecting HTTP client still replaces it. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- internal/dedupe/dynamodb.go | 19 ++++++++++++++++--- internal/dedupe/dynamodb_test.go | 17 +++++++++++++++++ 4 files changed, 35 insertions(+), 5 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index babb8f4e9..118c45215 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt%2D123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. +- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt%2D123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index aeaed95aa..8ee1ec6ad 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -126,7 +126,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key every backend stores, as text: `/
/` (for example `acme/clicks/evt%2D123`, [#222](https://github.com/Wave-RF/WaveHouse/issues/222)). The table and id are escaped by `internal/keyenc` — the escaping NATS subject tokens use — which never writes `/`, and a tenant id cannot hold one, so a table name may hold any byte, NUL included, and neither tenants nor tables share ids; a key is ASCII, so it reads as-is in a console and is a valid DynamoDB String. An id whose escaped form is over 1,024 bytes is stored as `#` plus its SHA-256 in hex (`#` is never written by the escaping), counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. -- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency`. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read, or `Claimed` when it is the put's own pending item (same token): an SDK retry of an attempt DynamoDB applied but whose answer was lost. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with jittered backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Each call has a `Timeout` (250 ms) covering the SDK's retries (`MaxAttempts`, 3), which back off with full jitter under a ceiling capped at `Timeout/(2·(MaxAttempts−1))`, so a call's retries wait at most half its timeout and a throttled call fails on its last attempt's answer rather than on the deadline. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. +- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency` (64), with as many idle connections kept per host so a wide `Reserve` reuses them rather than dial. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read, or `Claimed` when it is the put's own pending item (same token): an SDK retry of an attempt DynamoDB applied but whose answer was lost. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with jittered backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Each call has a `Timeout` (250 ms) covering the SDK's retries (`MaxAttempts`, 3), which back off with full jitter under a ceiling capped at `Timeout/(2·(MaxAttempts−1))`, so a call's retries wait at most half its timeout and a throttled call fails on its last attempt's answer rather than on the deadline. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also reads a lease `<= 0` as `DefaultLease` and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index bd9a74858..bd15f84db 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -7,12 +7,14 @@ import ( "fmt" "log/slog" mathrand "math/rand/v2" + "net/http" "strconv" "sync" "time" "github.com/aws/aws-sdk-go-v2/aws" "github.com/aws/aws-sdk-go-v2/aws/retry" + awshttp "github.com/aws/aws-sdk-go-v2/aws/transport/http" "github.com/aws/aws-sdk-go-v2/config" "github.com/aws/aws-sdk-go-v2/service/dynamodb" "github.com/aws/aws-sdk-go-v2/service/dynamodb/types" @@ -79,7 +81,8 @@ type DynamoConfig struct { // the client after throttles. RetryMode string // ReserveConcurrency bounds the parallel calls one Reserve, Commit or - // Release makes. 0 = 64. + // Release makes, and sizes the client's idle connection pool to match. + // 0 = 64. ReserveConcurrency int } @@ -128,7 +131,7 @@ type Dynamo struct { // NewDynamo builds the backend over a client from the SDK's default config // chain. extra is appended to the chain's options (a test's static -// credentials, say). It dials nothing: Check does. +// credentials or HTTP client, say). It dials nothing: Check does. func NewDynamo(ctx context.Context, cfg DynamoConfig, extra ...func(*config.LoadOptions) error) (*Dynamo, error) { if cfg.Table == "" { return nil, errors.New("dedupe: dynamodb table is required") @@ -138,7 +141,7 @@ func NewDynamo(ctx context.Context, cfg DynamoConfig, extra ...func(*config.Load if err != nil { return nil, err } - opts := []func(*config.LoadOptions) error{config.WithRetryer(retryer)} + opts := []func(*config.LoadOptions) error{config.WithRetryer(retryer), config.WithHTTPClient(newHTTPClient(cfg))} if cfg.Region != "" { opts = append(opts, config.WithRegion(cfg.Region)) } @@ -154,6 +157,16 @@ func NewDynamo(ctx context.Context, cfg DynamoConfig, extra ...func(*config.Load return newDynamo(client, cfg), nil } +// newHTTPClient keeps an idle connection for every call one Reserve can have +// in flight: with the SDK's default of 10 per host, a wide Reserve would dial +// most of its puts afresh. +func newHTTPClient(cfg DynamoConfig) *awshttp.BuildableClient { + return awshttp.NewBuildableClient().WithTransportOptions(func(tr *http.Transport) { + tr.MaxIdleConnsPerHost = cfg.ReserveConcurrency + tr.MaxIdleConns = max(tr.MaxIdleConns, cfg.ReserveConcurrency) + }) +} + func newRetryer(cfg DynamoConfig) (func() aws.Retryer, error) { standard := func(o *retry.StandardOptions) { o.MaxAttempts = cfg.MaxAttempts diff --git a/internal/dedupe/dynamodb_test.go b/internal/dedupe/dynamodb_test.go index d09af7b95..ddffeb480 100644 --- a/internal/dedupe/dynamodb_test.go +++ b/internal/dedupe/dynamodb_test.go @@ -16,6 +16,7 @@ import ( "github.com/aws/aws-sdk-go-v2/aws" "github.com/aws/aws-sdk-go-v2/aws/retry" + awshttp "github.com/aws/aws-sdk-go-v2/aws/transport/http" "github.com/aws/aws-sdk-go-v2/config" "github.com/aws/aws-sdk-go-v2/credentials" "github.com/aws/aws-sdk-go-v2/service/dynamodb" @@ -535,6 +536,22 @@ func TestDynamo_ThrottledCallEndsOnItsLastAttempt(t *testing.T) { assert.Equal(t, int64(calls*d.cfg.MaxAttempts), h.puts.Load()) } +// The idle pool holds a connection for every call one Reserve can have in +// flight, so a wide Reserve reuses them rather than dial; an HTTP client in +// extra (TestDynamo_ThrottledCallEndsOnItsLastAttempt's) replaces it. +func TestNewDynamo_SizesTheIdlePool(t *testing.T) { + t.Parallel() + for _, n := range []int{0, 8, 200} { + d, err := NewDynamo(t.Context(), DynamoConfig{Table: "dedupe", Region: "us-east-1", ReserveConcurrency: n}) + require.NoError(t, err) + client, ok := d.api.(*dynamodb.Client).Options().HTTPClient.(*awshttp.BuildableClient) + require.True(t, ok) + tr := client.GetTransport() + assert.GreaterOrEqual(t, tr.MaxIdleConnsPerHost, d.cfg.ReserveConcurrency, "ReserveConcurrency %d", n) + assert.GreaterOrEqual(t, tr.MaxIdleConns, d.cfg.ReserveConcurrency, "ReserveConcurrency %d", n) + } +} + func TestExpiresAt(t *testing.T) { t.Parallel() base := time.Unix(100, 0) From eb5dea9035b8daf76ecbdb13f0e9cef70e1befb3 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 04:55:03 -0400 Subject: [PATCH 18/38] fix(dedupe): retry a throttled DynamoDB commit like unprocessed items commitChunk returned on the first ErrUnavailable from BatchWriteItem. DynamoDB raises the throttle exception only when it processed none of the batch, so a total throttle got the SDK's retries inside one call while a partial throttle got eight backoff rounds: Commit gave up sooner under heavier throttling. These records are already published, and a lost commit lets a retry publish them again. A transient (ErrUnavailable-class) error from the call now counts as a round in which every item came back unprocessed: the same jittered backoff, the same eight-round cap, and the last round's error (with its cause) if the cap is reached. The puts are idempotent, so resending a batch whose fate a timeout left unknown is safe. Non-transient errors stay final. The unprocessed-items metric still counts only what DynamoDB reported unprocessed. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- internal/dedupe/dynamodb.go | 23 +++++++++++--- internal/dedupe/dynamodb_test.go | 44 +++++++++++++++++++++++++++ 4 files changed, 64 insertions(+), 7 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 118c45215..671b8bd35 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt%2D123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. +- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt%2D123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, retrying for up to eight jittered rounds both the items DynamoDB leaves unprocessed and a batch it throttled whole, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **A tenant removed or rejected at runtime has its open streams ended** (`internal/stream/{hub,subscriber,bucket}.go` (+ tests), `internal/api/stream.go` (+ tests), `internal/api/router.go`, `internal/settings/{registry,tree}.go` (+ tests), `internal/ingest/worker.go` (+ tests), `internal/app/wire.go` (+ tests), `clients/ts/src/stream/sse.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/settings-directory.mdx`, `AGENTS.md`): story 3 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). A `GET /v1/stream` used to outlive its tenant: the hub read no policy for it and withheld every row while the keepalive wheel held the connection open, so the client could not tell it from a quiet table. After every reload the hub now evicts the subscribers of each tenant no longer served, removed or rejected alike (`Hub.Prune`, through the close-once `Subscriber.Evict`), and the handler ends the stream: a gap-fill in progress included, and a stream `TenantMW` admitted just before the reload but registered just after it, which the handler checks for as it registers. The client's reconnect gets `404` (removed) or `503` (rejected); the SDK stops on the first and retries the second, resuming from `Last-Event-ID` once the folder is back — except in a browser going cross-origin where tenant `0` is not served or its CORS list does not admit the page, which cannot read either refusal (both are decorated from tenant `0`'s list) and re-dials as after a dropped connection. A flat directory never stops serving tenant `0`, so nothing changes there. Removing a tenant is deleting its folder, then reloading the whole directory, the last folder included: a server started with tenant folders reads the emptied directory as no folder left, where it read as a change of shape and the reload was rejected whole, leaving that tenant served; at boot an empty directory still reads as the four files, missing. Reloading the deleted folder by name leaves its tenant rejected. `GET /v1/health` keeps resolving a tenant, deliberately: its `404` tells a caller with no token no more than every tenant route's does, since they all answer before authenticating, and resolving answers a served tenant's ping from that tenant's own CORS list. A batch queued for a tenant with no ClickHouse connection — one no longer served, or one no pool could be opened for (such as by the connection ceiling) — skips the row-by-row retry, which no row of it could pass, and meets its DLQ switch once, whole, logged once per batch rather than twice per row; a tenant no longer served reads the switch as on, so its queued rows are parked under its own subject rather than dropped, inserted into another tenant's ClickHouse, or left unacked, redelivered for as long as the tenant is away and holding its queue's ack floor, which the purge waits on. Over a nested directory `/livez` no longer keeps naming a tenant that stopped being served before any tenant completed a first discovery: the diagnostic goes back to `no tenant has completed a first discovery yet`. - **One ClickHouse pool per tuple and one schema registry per tenant** (`internal/chconn/chconn.go` (+ tests), `internal/discovery/discovery.go` (+ tests), `internal/app/discoveries.go` (new), `internal/app/{app,wire}.go` (+ tests), `internal/api/{schema,ingest,structured_query,pipes,query,health,errors,clickhouse_exec}.go` (+ tests), `internal/stream/hub.go`, `internal/ingest/worker.go`, `internal/cache/{cache,local,version_manager}.go` (+ tests), `internal/testutil/{testutil,mocks}.go`, `tests/integration/{setup,tenants,boot_resilience,query_limits}_test.go`, `clients/ts/src/{schema,table,sql,client,types}.ts` (+ tests), `tests/e2e/sdk/admin.test.ts`, `docs/src/content/docs/{api,deployment,architecture,ingest-pipeline}.md`, `docs/src/content/docs/{settings-directory,configuration,access-control,reverse-proxy}.mdx`, `docs/src/content/docs/sdk/{admin,reference,queries}.md`, `AGENTS.md`): the second slice of story 6 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)), with no behavior change for a settings directory that holds the four files beyond the three noted at the end. The process opens one native pool per distinct `clickhouse.addr` / `database` / `username` / password / `tls` tuple among the served tenants (`chconn.Identity`, `chconn.Pools`), shared by the tenants naming it and sized to their largest `max_open_conns` and `max_idle_conns`; `http_port`, `http_scheme`, `headers` and `query_timeout` stay each tenant's own. Every reload reconciles the pools: a new tuple opens (never dials), a tenant whose tuple changed is repointed, a tuple no tenant names closes after the longest `query_timeout` among the tenants it had, and a pool whose largest ask changed is resized with the same grace. The boot config's `clickhouse.max_total_conns` now bounds the open pools together: boot is refused naming the sum and the ceiling; at a reload a resize above it is refused with the pool kept at its size, and a tuple that cannot be opened — the ceiling, a certificate file that cannot be read, or options the driver refuses — leaves its tenants on the pool they had (the keep-previous-wiring rule of the connection ceiling) or on none when they had none; both are logged and the next reload retries. A tenant on no pool fails closed: `503` with `Retry-After: 30` on `POST /v1/query`, `GET/POST /v1/pipes/{name}`, `POST /v1/ops/query` and `POST /v1/ops/schema/refresh`, ahead of the cache. Each served tenant gets a `discovery.SchemaRegistry` of its own over its pool, kept fresh by its own loop — the boot retry until the first success, then `schema.refresh_interval` with the first refresh at a random point within the interval so tenants adopted together do not refresh together — created and stopped from the settings reload and stopped under `App.Close`; `wavehouse_schema_refresh_failures_total{tenant}` counts a loop's failed attempts. `GET /v1/ops/schema`, `POST /v1/ops/schema/refresh` and `POST /v1/ops/query` take the strict `?tenant=` the pipe reads take (absent is tenant `0`, `400` malformed, `404` unknown, `503` rejected); the SDK sends it as the `tenant` option of `wh.schema.list()`, `wh.schema.refresh()`, `wh.from(t).schema()` and `wh.sql()`. Over a nested directory `/livez` (and `/v1/health`) is `503` with the latest discovery failure, naming its tenant, while no tenant has completed a first discovery, then `200` for the rest of the process lifetime; `/readyz` pings every open pool at once, is ready at the first answer, and names every pool that did not answer when none does — one tenant's ClickHouse outage is its log line and counter, never a probe failure. The ingest worker inserts each batch into its own tenant's ClickHouse — the tenant the message's topic names, through that tenant's HTTP wiring (`chconn.Pools.Target`); a tenant on no pool takes the failure path an unreachable ClickHouse takes — and its cache invalidation fans out to the tenants on the same address and database as the batch's tenant, whatever their user or `tls` block (they read the same tables), rather than to every known tenant, and a tenant adopted after an absence — rejected or removed, so out of that fan-out — or moved to another address or database has its cached structured-query results orphaned in one step (`Cache.InvalidateTenant`, a tenant generation in every version key), so a repaired folder never serves query rows cached before the inserts it missed (a pipe result names no table, so neither this nor any insert invalidates it: it stays until its TTL expires). Three changes reach the single-tenant directory too: a table lookup before the first discovery is a `503` with `Retry-After: 5` (`schema not loaded yet`) rather than a `404`, on `POST /v1/ingest`, `POST /v1/query` and `GET /v1/ops/schema` (the list included, where `[]` would read as no tables); the first periodic refresh fires at a random point within the interval rather than a full interval after boot; and a reload that moves `clickhouse.addr` or `clickhouse.database` now orphans the structured-query results cached before it, where they were served until their TTL. Until [#529](https://github.com/Wave-RF/WaveHouse/issues/529) every tenant's user authenticates with `WH_CH_PASSWORD`, so the tuple is in effect the address, database, username and `tls` block. - **One token verifier per tenant, built off the boot and reload paths** (`internal/auth/auth.go` (+ tests), `internal/api/router.go` (+ tests), `internal/app/{app,wire}.go` (+ tests), `go.mod`, `docs/src/content/docs/sdk/{reference,streaming}.md`, `docs/src/content/docs/{architecture,deployment,api}.md`, `docs/src/content/docs/{settings-directory,configuration}.mdx`, `SECURITY.md`): story 9 of the multi-tenant epic ([#583](https://github.com/Wave-RF/WaveHouse/issues/583)). Over a nested settings directory each tenant's folder now wires that tenant's verifier (`auth.jwks_url`, `auth.role_claim`), so a JWKS-issued token verifies only under the tenants whose `jwks_url` names its identity provider's key set — under any other tenant's header it is refused as invalid — where before one verifier, tenant `0`'s, accepted a token under any header. `auth.Config` is now the boot-config half alone (the HMAC secret and the operator key, shared by every tenant) and the new `auth.Wiring` a tenant's half; `NewAuthenticator` builds no verifier, `Reconfigure(id, wiring)` gives a tenant one — swapped atomically when its wiring changed, kept when it did not — `Prune` drops the verifiers of the tenants a reload stopped serving, rejected or removed alike (no work runs for a tenant that is not served; a folder adopted again gets a fresh verifier), and `Close` stops every JWKS refresh as a component of `App.Close`. This holds on every reload because the registry's `AfterAdopt` hooks run after every reload it applied, empty list included, so a per-tenant reload that rejects a folder drops its verifier then rather than at the next adoption. The middleware reads the request tenant's verifier through an injected `auth.TenantSource` — the store `api.TenantMW` resolved names its tenant with `settings.Store.Tenant()`, one read of the context — a tenant-exempt route (the ops tree) verifies as tenant `0`, and a tenant with no verifier fails closed. The operator key stamps the request tenant's `admin_role`, read from that tenant's policy, rather than tenant `0`'s. A JWKS key set is fetched on its own goroutine: the verifier is in place at once and *pending* until a set has been stored — the first fetch retried with backoff from one second to a minute, a stored set then kept fresh by the library hourly and, rate-limited, on an unknown key id — so neither boot nor a reload (which holds the registry's lock) waits on the endpoint, and **an unreachable JWKS no longer refuses boot** — it logs (`jwks refresh failed; no token validates until it succeeds`) and that tenant alone is affected. A token checked against a pending verifier is a new outcome, `auth.ErrVerifierPending`: every `/v1` route answers it `503 {"error": "token verifier not ready: …"}` with `Retry-After: 30` (`api.refuseUnverifiable`) rather than evaluate the request under the `default_role`, which could accept its data under a lesser role while another pod holding the keys would have served it as its own; a request without a token, and the operator key, are unaffected. Every fetch goes through one client that caps the response at 1 MiB, refusing a larger one as unreachable with `jwks response exceeds 1048576 bytes` as the logged cause. Nothing changes for a flat directory beyond that boot rule: its one tenant gets exactly the verifier it had. The boot warning for the secretless posture (no `auth.jwt_secret` and no `jwks_url`, whose token refusal landed in [#607](https://github.com/Wave-RF/WaveHouse/pull/607)) is now per tenant — each served tenant with no `jwks_url` while the boot secret is unset — where one line that any JWKS tenant silenced used to stand for the whole directory. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 8ee1ec6ad..b8ab38ccb 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -126,7 +126,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key every backend stores, as text: `/
/` (for example `acme/clicks/evt%2D123`, [#222](https://github.com/Wave-RF/WaveHouse/issues/222)). The table and id are escaped by `internal/keyenc` — the escaping NATS subject tokens use — which never writes `/`, and a tenant id cannot hold one, so a table name may hold any byte, NUL included, and neither tenants nor tables share ids; a key is ASCII, so it reads as-is in a console and is a valid DynamoDB String. An id whose escaped form is over 1,024 bytes is stored as `#` plus its SHA-256 in hex (`#` is never written by the escaping), counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. -- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency` (64), with as many idle connections kept per host so a wide `Reserve` reuses them rather than dial. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read, or `Claimed` when it is the put's own pending item (same token): an SDK retry of an attempt DynamoDB applied but whose answer was lost. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying unprocessed items with jittered backoff. `Release` is a `DeleteItem` conditional on the token and the pending state. Each call has a `Timeout` (250 ms) covering the SDK's retries (`MaxAttempts`, 3), which back off with full jitter under a ceiling capped at `Timeout/(2·(MaxAttempts−1))`, so a call's retries wait at most half its timeout and a throttled call fails on its last attempt's answer rather than on the deadline. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. +- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency` (64), with as many idle connections kept per host so a wide `Reserve` reuses them rather than dial. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read, or `Claimed` when it is the put's own pending item (same token): an SDK retry of an attempt DynamoDB applied but whose answer was lost. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying with jittered backoff, for up to eight rounds, both the items DynamoDB leaves unprocessed and a batch that failed transiently (a throttle means it processed none of it); the records are already published, and a table that throttles every round delays the ingest response by at most about 3 s at the defaults (eight 250 ms calls and the waits between them) before the commit is given up. `Release` is a `DeleteItem` conditional on the token and the pending state. Each call has a `Timeout` (250 ms) covering the SDK's retries (`MaxAttempts`, 3), which back off with full jitter under a ceiling capped at `Timeout/(2·(MaxAttempts−1))`, so a call's retries wait at most half its timeout and a throttled call fails on its last attempt's answer rather than on the deadline. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also reads a lease `<= 0` as `DefaultLease` and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index bd15f84db..f5c22b59a 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -46,7 +46,7 @@ const ( // batchWriteMax is BatchWriteItem's per-call item limit. batchWriteMax = 25 // commitRounds bounds the BatchWriteItem rounds one chunk gets before - // its still-unprocessed items fail the Commit. + // its still-unprocessed (or still-throttled) items fail the Commit. commitRounds = 8 // commitBase and commitCeiling bound the jittered wait between Commit // rounds; retryBase is the one the SDK retryer's backoff doubles from. @@ -125,7 +125,7 @@ type Dynamo struct { breaker *breaker metrics dynamoMetrics // commitBackoff is the wait before retrying the attempt'th round of - // unprocessed items. + // unprocessed or throttled items. commitBackoff func(attempt int) time.Duration } @@ -419,7 +419,8 @@ func heldBy(item map[string]types.AttributeValue, token string) bool { } // Commit overwrites every claim's item as committed, unconditionally, 25 to a -// BatchWriteItem, retrying the items DynamoDB leaves unprocessed. +// BatchWriteItem, retrying the items DynamoDB leaves unprocessed and a batch +// that failed transiently. func (s *dynamoStore) Commit(ctx context.Context, claims []Claim, retention time.Duration) error { if len(claims) == 0 { return nil @@ -468,16 +469,28 @@ func (s *dynamoStore) commitChunk(ctx context.Context, writes []types.WriteReque } return err }) - if err != nil { + switch { + case err == nil: + case errors.Is(err, ErrUnavailable): + // DynamoDB throttles a batch whole only when it processed none of + // it, and a timeout leaves its fate unknown: retry it whole, as a + // round that left every item unprocessed. The puts are idempotent. + unprocessed = writes + default: return err } if len(unprocessed) == 0 { return nil } if attempt+1 >= commitRounds { + if err != nil { + return err + } return fmt.Errorf("%w: dynamodb batch_write_item: %d items still unprocessed", ErrUnavailable, len(unprocessed)) } - s.d.metrics.unprocessed.Add(ctx, int64(len(unprocessed))) + if err == nil { + s.d.metrics.unprocessed.Add(ctx, int64(len(unprocessed))) + } writes = unprocessed select { case <-ctx.Done(): diff --git a/internal/dedupe/dynamodb_test.go b/internal/dedupe/dynamodb_test.go index ddffeb480..31f54f85e 100644 --- a/internal/dedupe/dynamodb_test.go +++ b/internal/dedupe/dynamodb_test.go @@ -303,6 +303,50 @@ func TestDynamo_CommitGivesUpOnItemsThatStayUnprocessed(t *testing.T) { require.ErrorIs(t, err, ErrUnavailable) } +// DynamoDB answers a BatchWriteItem it processed none of with a throttle, +// not with every item unprocessed: the chunk gets the same rounds either way. +func TestDynamo_CommitRetriesAFailedBatch(t *testing.T) { + t.Parallel() + throttle := &types.ProvisionedThroughputExceededException{} + for _, tc := range []struct { + name string + fails int + err error + calls int64 + committed bool + unavailable bool + }{ + {"throttled, then through", 3, throttle, 4, true, false}, + {"throttled every round", 1 << 10, throttle, commitRounds, false, true}, + {"timed out, then through", 1, fmt.Errorf("op: %w", context.DeadlineExceeded), 2, true, false}, + {"a configuration bug is final", 1 << 10, &types.ResourceNotFoundException{}, 1, false, false}, + } { + t.Run(tc.name, func(t *testing.T) { + t.Parallel() + var calls atomic.Int64 + _, m := openFake(t, &fakeDynamo{batch: func(_ context.Context, in *dynamodb.BatchWriteItemInput) (*dynamodb.BatchWriteItemOutput, error) { + assert.Len(t, in.RequestItems["dedupe"], 2, "the whole batch, every round") + if calls.Add(1) <= int64(tc.fails) { + return nil, tc.err + } + return &dynamodb.BatchWriteItemOutput{}, nil + }}) + var claims []Claim + for _, k := range keys("a", "b") { + claims = append(claims, Claim{Key: k, Status: Claimed, Token: "t"}) + } + err := m.Commit(t.Context(), claims, 0) + assert.Equal(t, tc.calls, calls.Load()) + if tc.committed { + require.NoError(t, err) + return + } + require.ErrorIs(t, err, tc.err, "the last round's cause is kept") + assert.Equal(t, tc.unavailable, errors.Is(err, ErrUnavailable)) + }) + } +} + func TestDynamo_ReleaseTreatsAFailedConditionAsDone(t *testing.T) { t.Parallel() _, m := openFake(t, &fakeDynamo{del: func(_ context.Context, in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { From f40c76aa4acba1ebea48439f86a5584732353a17 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 04:55:18 -0400 Subject: [PATCH 19/38] fix(config): dedupe defaults in defaults(); refuse an explicit zero The dedupe keys carried cleanenv env-default tags, which main no longer allows: cleanenv applies them after the YAML decode to any field still zero, so an explicit zero in config.yaml silently became the default. dedupe.lease, reserve_concurrency and the dynamodb block's timeout, max_attempts and retry_mode now default in defaults(), and a zero lease, concurrency, timeout or attempt count, or an empty retry_mode, refuses boot (the dynamodb block's only while dynamodb is selected). Each is a refusedZeros case; the docs-defaults test gains a duration parser. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- CHANGELOG.md | 2 +- docs/src/content/docs/configuration.mdx | 12 ++--- internal/config/backends.go | 30 ++++++------- internal/config/backends_test.go | 17 +++---- internal/config/config.go | 8 +++- internal/config/defaults_test.go | 59 +++++++++++++++++-------- 6 files changed, 77 insertions(+), 51 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 2d5b586b3..db2333b61 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/backends.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `2m`, the embedded queue's duplicate window), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`; no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and on every reload. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. +- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `2m`, the embedded queue's duplicate window), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`; no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and on every reload. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. - **A DynamoDB dedupe backend** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. `dedupe.backend: dynamodb` selects it (entry above). New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index e98b69aed..7dbd31710 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -56,8 +56,8 @@ Whether a tenant dedupes, and on which field, are settings-directory keys ([Dedu | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | -| `dedupe.lease` | `WH_DEDUPE_LEASE` | `30s` | How long a record's id stays claimed while the record is published. Another request carrying the same id meanwhile gets `503` with this as `Retry-After`, in whole seconds; a claim that is neither committed nor released, because its process died mid-publish, lapses after it. At most `2m` with `mq.backend: embedded`, the embedded queue's duplicate window: a longer lease refuses boot. A Go duration (`30s`, `1m`); `0` = the default. | -| `dedupe.reserve_concurrency` | `WH_DEDUPE_RESERVE_CONCURRENCY` | `64` | The most parallel calls one Reserve, Commit or Release makes to a remote dedupe backend. Ingest sends one id per call today, so it has no effect yet; `pebble` ignores it. `0` = the default. | +| `dedupe.lease` | `WH_DEDUPE_LEASE` | `30s` | How long a record's id stays claimed while the record is published. Another request carrying the same id meanwhile gets `503` with this as `Retry-After`, in whole seconds; a claim that is neither committed nor released, because its process died mid-publish, lapses after it. At most `2m` with `mq.backend: embedded`, the embedded queue's duplicate window: a longer lease refuses boot. A Go duration (`30s`, `1m`); `0` refuses boot. | +| `dedupe.reserve_concurrency` | `WH_DEDUPE_RESERVE_CONCURRENCY` | `64` | The most parallel calls one Reserve, Commit or Release makes to a remote dedupe backend. Ingest sends one id per call today, so it has no effect yet; `pebble` ignores it. `0` refuses boot. | #### DynamoDB dedupe @@ -65,12 +65,12 @@ Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | -| `dedupe.dynamodb.table` | `WH_DEDUPE_DYNAMODB_TABLE` | *(none)* | The shared table. Required. | +| `dedupe.dynamodb.table` | `WH_DEDUPE_DYNAMODB_TABLE` | *(required)* | The shared table. | | `dedupe.dynamodb.region` | `WH_DEDUPE_DYNAMODB_REGION` | *(empty)* | The table's region. Empty uses the SDK chain's (`AWS_REGION`); no region from either refuses boot. | | `dedupe.dynamodb.endpoint` | `WH_DEDUPE_DYNAMODB_ENDPOINT` | *(empty)* | A custom endpoint, for [dynamodb-local](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/DynamoDBLocal.html) in development and tests. Leave it empty against AWS. | -| `dedupe.dynamodb.timeout` | `WH_DEDUPE_DYNAMODB_TIMEOUT` | `250ms` | Deadline for each DynamoDB call, the SDK's retries included. `0` = the default. | -| `dedupe.dynamodb.max_attempts` | `WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS` | `3` | Attempts per call, the first included. `0` = the default. | -| `dedupe.dynamodb.retry_mode` | `WH_DEDUPE_DYNAMODB_RETRY_MODE` | `standard` | `standard`, or `adaptive`, which also slows the client down after throttling. | +| `dedupe.dynamodb.timeout` | `WH_DEDUPE_DYNAMODB_TIMEOUT` | `250ms` | Deadline for each DynamoDB call, the SDK's retries included. `0` refuses boot. | +| `dedupe.dynamodb.max_attempts` | `WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS` | `3` | Attempts per call, the first included. `0` refuses boot. | +| `dedupe.dynamodb.retry_mode` | `WH_DEDUPE_DYNAMODB_RETRY_MODE` | `standard` | `standard`, or `adaptive`, which also slows the client down after throttling. Anything else, empty included, refuses boot. | | `dedupe.dynamodb.create_table` | `WH_DEDUPE_DYNAMODB_CREATE_TABLE` | `false` | Development only: create the table at boot if it is missing, with TTL on `ex`. Refused unless `endpoint` is set, so it never creates a table in AWS; the production table belongs to your infrastructure code. | ### Process roles diff --git a/internal/config/backends.go b/internal/config/backends.go index 6286956ba..493e9488c 100644 --- a/internal/config/backends.go +++ b/internal/config/backends.go @@ -75,10 +75,10 @@ type Dedupe struct { Backend DedupeBackend `yaml:"backend" env:"WH_DEDUPE_BACKEND"` // Lease is how long a claimed id stays pending while its record is // published; a claim its request never settles lapses after it. - Lease time.Duration `yaml:"lease" env:"WH_DEDUPE_LEASE" env-default:"30s"` + Lease time.Duration `yaml:"lease" env:"WH_DEDUPE_LEASE"` // ReserveConcurrency bounds the parallel calls one Reserve, Commit or // Release makes to a remote backend. Pebble ignores it. - ReserveConcurrency int `yaml:"reserve_concurrency" env:"WH_DEDUPE_RESERVE_CONCURRENCY" env-default:"64"` + ReserveConcurrency int `yaml:"reserve_concurrency" env:"WH_DEDUPE_RESERVE_CONCURRENCY"` DynamoDB DedupeDynamoDBConfig `yaml:"dynamodb"` } @@ -93,23 +93,23 @@ type DedupeDynamoDBConfig struct { Region string `yaml:"region" env:"WH_DEDUPE_DYNAMODB_REGION"` // Endpoint points the client at dynamodb-local. Endpoint string `yaml:"endpoint" env:"WH_DEDUPE_DYNAMODB_ENDPOINT"` - Timeout time.Duration `yaml:"timeout" env:"WH_DEDUPE_DYNAMODB_TIMEOUT" env-default:"250ms"` - MaxAttempts int `yaml:"max_attempts" env:"WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS" env-default:"3"` - RetryMode string `yaml:"retry_mode" env:"WH_DEDUPE_DYNAMODB_RETRY_MODE" env-default:"standard"` + Timeout time.Duration `yaml:"timeout" env:"WH_DEDUPE_DYNAMODB_TIMEOUT"` + MaxAttempts int `yaml:"max_attempts" env:"WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS"` + RetryMode string `yaml:"retry_mode" env:"WH_DEDUPE_DYNAMODB_RETRY_MODE"` // CreateTable creates the table at boot if it is missing. Development // only: refused unless Endpoint is set. - CreateTable bool `yaml:"create_table" env:"WH_DEDUPE_DYNAMODB_CREATE_TABLE" env-default:"false"` + CreateTable bool `yaml:"create_table" env:"WH_DEDUPE_DYNAMODB_CREATE_TABLE"` } func (d Dedupe) validate() error { if err := checkBackend("dedupe.backend", "WH_DEDUPE_BACKEND", d.Backend, dedupeBackends); err != nil { return err } - if d.Lease < 0 { - return fmt.Errorf("dedupe.lease (WH_DEDUPE_LEASE) must be >= 0, got %s", d.Lease) + if d.Lease <= 0 { + return fmt.Errorf("dedupe.lease (WH_DEDUPE_LEASE) must be > 0, got %s", d.Lease) } - if d.ReserveConcurrency < 0 { - return fmt.Errorf("dedupe.reserve_concurrency (WH_DEDUPE_RESERVE_CONCURRENCY) must be >= 0, got %d", d.ReserveConcurrency) + if d.ReserveConcurrency <= 0 { + return fmt.Errorf("dedupe.reserve_concurrency (WH_DEDUPE_RESERVE_CONCURRENCY) must be > 0, got %d", d.ReserveConcurrency) } if d.Backend == DedupeDynamoDB { return d.DynamoDB.validate() @@ -121,11 +121,11 @@ func (d DedupeDynamoDBConfig) validate() error { switch { case strings.TrimSpace(d.Table) == "": return errors.New("dedupe.dynamodb.table (WH_DEDUPE_DYNAMODB_TABLE) is required when dedupe.backend is dynamodb") - case d.Timeout < 0: - return fmt.Errorf("dedupe.dynamodb.timeout (WH_DEDUPE_DYNAMODB_TIMEOUT) must be >= 0, got %s", d.Timeout) - case d.MaxAttempts < 0: - return fmt.Errorf("dedupe.dynamodb.max_attempts (WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS) must be >= 0, got %d", d.MaxAttempts) - case d.RetryMode != "" && d.RetryMode != "standard" && d.RetryMode != "adaptive": + case d.Timeout <= 0: + return fmt.Errorf("dedupe.dynamodb.timeout (WH_DEDUPE_DYNAMODB_TIMEOUT) must be > 0, got %s", d.Timeout) + case d.MaxAttempts <= 0: + return fmt.Errorf("dedupe.dynamodb.max_attempts (WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS) must be > 0, got %d", d.MaxAttempts) + case d.RetryMode != "standard" && d.RetryMode != "adaptive": return fmt.Errorf("dedupe.dynamodb.retry_mode (WH_DEDUPE_DYNAMODB_RETRY_MODE) %q: want standard or adaptive", d.RetryMode) case d.CreateTable && d.Endpoint == "": return errors.New("dedupe.dynamodb.create_table (WH_DEDUPE_DYNAMODB_CREATE_TABLE) is for dynamodb-local only: set dedupe.dynamodb.endpoint, or create the table with your infrastructure code") diff --git a/internal/config/backends_test.go b/internal/config/backends_test.go index 26144915d..56f0d083e 100644 --- a/internal/config/backends_test.go +++ b/internal/config/backends_test.go @@ -11,11 +11,11 @@ import ( ) // withDefaultBackends sets what defaults() would: a literal Config -// names no backend and no role, and Validate refuses that. +// names no backend, no role and no dedupe lease, and Validate refuses that. func withDefaultBackends(c Config) *Config { c.Roles = AllRoles() c.MQ.Backend, c.Cache.Backend = MQEmbedded, CacheLocal - c.Dedupe.Backend, c.Coord.Backend = DedupePebble, CoordLocal + c.Dedupe, c.Coord.Backend = defaults().Dedupe, CoordLocal return &c } @@ -254,15 +254,11 @@ func TestValidate_Dedupe(t *testing.T) { want string // "" = valid }{ {"dynamodb", dynamo, ""}, - {"zero values read as the defaults", func(c *Config) { - c.Dedupe.Backend = DedupeDynamoDB - c.Dedupe.DynamoDB = DedupeDynamoDBConfig{Table: "t"} - }, ""}, {"create_table with an endpoint", func(c *Config) { dynamo(c) c.Dedupe.DynamoDB.Endpoint, c.Dedupe.DynamoDB.CreateTable = "http://localhost:8000", true }, ""}, - {"the block is not read under pebble", func(c *Config) { c.Dedupe.DynamoDB.CreateTable = true }, ""}, + {"the block is not read under pebble", func(c *Config) { c.Dedupe.DynamoDB = DedupeDynamoDBConfig{CreateTable: true} }, ""}, {"lease at the duplicate window", func(c *Config) { c.Dedupe.Lease = 2 * time.Minute }, ""}, {"create_table without an endpoint", func(c *Config) { dynamo(c) @@ -270,9 +266,14 @@ func TestValidate_Dedupe(t *testing.T) { }, "dedupe.dynamodb.create_table (WH_DEDUPE_DYNAMODB_CREATE_TABLE) is for dynamodb-local only"}, {"no table", func(c *Config) { dynamo(c); c.Dedupe.DynamoDB.Table = " " }, "dedupe.dynamodb.table (WH_DEDUPE_DYNAMODB_TABLE) is required"}, {"retry mode", func(c *Config) { dynamo(c); c.Dedupe.DynamoDB.RetryMode = "legacy" }, `retry_mode (WH_DEDUPE_DYNAMODB_RETRY_MODE) "legacy"`}, + {"zero timeout", func(c *Config) { dynamo(c); c.Dedupe.DynamoDB.Timeout = 0 }, "dedupe.dynamodb.timeout (WH_DEDUPE_DYNAMODB_TIMEOUT) must be > 0"}, {"negative timeout", func(c *Config) { dynamo(c); c.Dedupe.DynamoDB.Timeout = -time.Second }, "dedupe.dynamodb.timeout"}, + {"zero attempts", func(c *Config) { dynamo(c); c.Dedupe.DynamoDB.MaxAttempts = 0 }, "dedupe.dynamodb.max_attempts (WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS) must be > 0"}, {"negative attempts", func(c *Config) { dynamo(c); c.Dedupe.DynamoDB.MaxAttempts = -1 }, "dedupe.dynamodb.max_attempts"}, - {"negative lease", func(c *Config) { c.Dedupe.Lease = -time.Second }, "dedupe.lease (WH_DEDUPE_LEASE) must be >= 0"}, + {"no retry mode", func(c *Config) { dynamo(c); c.Dedupe.DynamoDB.RetryMode = "" }, `retry_mode (WH_DEDUPE_DYNAMODB_RETRY_MODE) ""`}, + {"zero lease", func(c *Config) { c.Dedupe.Lease = 0 }, "dedupe.lease (WH_DEDUPE_LEASE) must be > 0, got 0s"}, + {"negative lease", func(c *Config) { c.Dedupe.Lease = -time.Second }, "dedupe.lease (WH_DEDUPE_LEASE) must be > 0"}, + {"zero concurrency", func(c *Config) { c.Dedupe.ReserveConcurrency = 0 }, "dedupe.reserve_concurrency (WH_DEDUPE_RESERVE_CONCURRENCY) must be > 0"}, {"negative concurrency", func(c *Config) { c.Dedupe.ReserveConcurrency = -1 }, "dedupe.reserve_concurrency"}, {"lease past the duplicate window", func(c *Config) { c.Dedupe.Lease = 3 * time.Minute }, "exceeds the embedded mq's 2m0s duplicate window"}, } diff --git a/internal/config/config.go b/internal/config/config.go index 95bfacb15..fcd0cd731 100644 --- a/internal/config/config.go +++ b/internal/config/config.go @@ -7,6 +7,7 @@ import ( "os" "slices" "strings" + "time" "github.com/ilyakaznacheev/cleanenv" ) @@ -256,8 +257,11 @@ func defaults() Config { Server: Server{Port: 8080, ShutdownTimeout: 10}, MQ: MQ{Backend: MQEmbedded}, Cache: Cache{Backend: CacheLocal, L1MaxCost: 64 << 20}, - Dedupe: Dedupe{Backend: DedupePebble}, - Coord: Coord{Backend: CoordLocal}, + Dedupe: Dedupe{ + Backend: DedupePebble, Lease: 30 * time.Second, ReserveConcurrency: 64, + DynamoDB: DedupeDynamoDBConfig{Timeout: 250 * time.Millisecond, MaxAttempts: 3, RetryMode: "standard"}, + }, + Coord: Coord{Backend: CoordLocal}, OTel: OTel{ Traces: OTelTraces{Enabled: true, SampleRate: 1.0}, Metrics: OTelMetrics{Enabled: true}, diff --git a/internal/config/defaults_test.go b/internal/config/defaults_test.go index 6877be4bb..cbbfdef8a 100644 --- a/internal/config/defaults_test.go +++ b/internal/config/defaults_test.go @@ -9,6 +9,7 @@ import ( "strconv" "strings" "testing" + "time" "github.com/stretchr/testify/assert" "github.com/stretchr/testify/require" @@ -41,35 +42,53 @@ var zeroCases = []zeroCase{ // refusedZeros are the non-zero defaults whose zero Validate refuses: written // in the file, the zero must reach Validate rather than become the default. +// also holds the keys a sub-block's zero needs to be read at all. var refusedZeros = []struct { key string zero any err string + also map[string]any }{ - {"server.port", 0, "server.port 0 out of range"}, - {"mq.backend", "", `mq.backend (WH_MQ_BACKEND) ""`}, - {"cache.backend", "", `cache.backend (WH_CACHE_BACKEND) ""`}, - {"dedupe.backend", "", `dedupe.backend (WH_DEDUPE_BACKEND) ""`}, - {"coord.backend", "", `coord.backend (WH_COORD_BACKEND) ""`}, - {"roles", []string{}, "roles (WH_ROLES) is empty"}, + {"server.port", 0, "server.port 0 out of range", nil}, + {"mq.backend", "", `mq.backend (WH_MQ_BACKEND) ""`, nil}, + {"cache.backend", "", `cache.backend (WH_CACHE_BACKEND) ""`, nil}, + {"dedupe.backend", "", `dedupe.backend (WH_DEDUPE_BACKEND) ""`, nil}, + {"coord.backend", "", `coord.backend (WH_COORD_BACKEND) ""`, nil}, + {"roles", []string{}, "roles (WH_ROLES) is empty", nil}, + {"dedupe.lease", "0s", "dedupe.lease (WH_DEDUPE_LEASE) must be > 0", nil}, + {"dedupe.reserve_concurrency", 0, "dedupe.reserve_concurrency (WH_DEDUPE_RESERVE_CONCURRENCY) must be > 0", nil}, + {"dedupe.dynamodb.timeout", "0s", "dedupe.dynamodb.timeout (WH_DEDUPE_DYNAMODB_TIMEOUT) must be > 0", dynamoSelected}, + {"dedupe.dynamodb.max_attempts", 0, "dedupe.dynamodb.max_attempts (WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS) must be > 0", dynamoSelected}, + {"dedupe.dynamodb.retry_mode", "", `dedupe.dynamodb.retry_mode (WH_DEDUPE_DYNAMODB_RETRY_MODE) ""`, dynamoSelected}, } -// yamlAt renders a file setting key to value, plus otel.enabled: true so -// the test can tell the file was read. -func yamlAt(t *testing.T, key string, value any) string { +// dynamoSelected is what the dedupe.dynamodb block needs to be read. +var dynamoSelected = map[string]any{"dedupe.backend": "dynamodb", "dedupe.dynamodb.table": "t"} + +// yamlAt renders a file setting key to value, and each dotted key of also to +// its value, plus otel.enabled: true so the test can tell the file was read. +func yamlAt(t *testing.T, key string, value any, also ...map[string]any) string { t.Helper() tree := map[string]any{"otel": map[string]any{"enabled": true}} - node := tree - parts := strings.Split(key, ".") - for _, p := range parts[:len(parts)-1] { - sub, ok := node[p].(map[string]any) - if !ok { - sub = map[string]any{} - node[p] = sub + set := func(key string, value any) { + node := tree + parts := strings.Split(key, ".") + for _, p := range parts[:len(parts)-1] { + sub, ok := node[p].(map[string]any) + if !ok { + sub = map[string]any{} + node[p] = sub + } + node = sub + } + node[parts[len(parts)-1]] = value + } + for _, m := range also { + for k, v := range m { + set(k, v) } - node = sub } - node[parts[len(parts)-1]] = value + set(key, value) out, err := yaml.Marshal(tree) require.NoError(t, err) return string(out) @@ -136,7 +155,7 @@ func TestLoad_YAMLZeroIsRefused(t *testing.T) { for _, tc := range refusedZeros { t.Run(tc.key, func(t *testing.T) { t.Parallel() - _, err := Load(writeYAML(t, yamlAt(t, tc.key, tc.zero))) + _, err := Load(writeYAML(t, yamlAt(t, tc.key, tc.zero, tc.also))) require.ErrorContains(t, err, tc.err, "the zero reaches Validate instead of becoming the default") }) } @@ -296,6 +315,8 @@ func parseDocDefault(t *testing.T, key, cell string, like any) any { v, err = strconv.ParseInt(cell, 10, 64) case float64: v, err = strconv.ParseFloat(cell, 64) + case time.Duration: + v, err = time.ParseDuration(cell) default: rt := reflect.TypeOf(like) switch { From c56103e90dda639cf182bca41f84b3d8b3741998 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 04:56:05 -0400 Subject: [PATCH 20/38] docs(dedupe): scope the DynamoDB Reserve undo to what it guarantees The comments on Reserve's undo, and the architecture page, said every put that may have landed is released. That holds for a sibling's failure, which never cuts off a put already sent, but not for a put cut off by the caller's cancellation or its own call deadline: DynamoDB can still apply it after the undo's conditional delete, and it then holds its key InFlight until the lease ends. A release that fails leaves a claim the same way. Both are bounded by the lease, as a crashed request's claims are, which the Deduplicator contract allows; the comments now say so instead of promising more. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- docs/src/content/docs/architecture.md | 2 +- internal/dedupe/dynamodb.go | 17 ++++++++++------- internal/dedupe/dynamodb_test.go | 2 +- 3 files changed, 12 insertions(+), 9 deletions(-) diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index b8ab38ccb..3c7330a69 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -126,7 +126,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key every backend stores, as text: `/
/` (for example `acme/clicks/evt%2D123`, [#222](https://github.com/Wave-RF/WaveHouse/issues/222)). The table and id are escaped by `internal/keyenc` — the escaping NATS subject tokens use — which never writes `/`, and a tenant id cannot hold one, so a table name may hold any byte, NUL included, and neither tenants nor tables share ids; a key is ASCII, so it reads as-is in a console and is a valid DynamoDB String. An id whose escaped form is over 1,024 bytes is stored as `#` plus its SHA-256 in hex (`#` is never written by the escaping), counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. -- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency` (64), with as many idle connections kept per host so a wide `Reserve` reuses them rather than dial. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read, or `Claimed` when it is the put's own pending item (same token): an SDK retry of an attempt DynamoDB applied but whose answer was lost. If any put errors, every put that may have landed is released by its token. `Commit` is `BatchWriteItem`, 25 at a time, retrying with jittered backoff, for up to eight rounds, both the items DynamoDB leaves unprocessed and a batch that failed transiently (a throttle means it processed none of it); the records are already published, and a table that throttles every round delays the ingest response by at most about 3 s at the defaults (eight 250 ms calls and the waits between them) before the commit is given up. `Release` is a `DeleteItem` conditional on the token and the pending state. Each call has a `Timeout` (250 ms) covering the SDK's retries (`MaxAttempts`, 3), which back off with full jitter under a ceiling capped at `Timeout/(2·(MaxAttempts−1))`, so a call's retries wait at most half its timeout and a throttled call fails on its last attempt's answer rather than on the deadline. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. +- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency` (64), with as many idle connections kept per host so a wide `Reserve` reuses them rather than dial. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read, or `Claimed` when it is the put's own pending item (same token): an SDK retry of an attempt DynamoDB applied but whose answer was lost. If any put errors, every put that may have landed is released by its token. A put already sent answers before that undo when a sibling fails, but one cut off by the caller's cancellation or its call deadline can still be applied after its release; it then holds its key `InFlight` until the lease ends, as a crashed request's claim does. `Commit` is `BatchWriteItem`, 25 at a time, retrying with jittered backoff, for up to eight rounds, both the items DynamoDB leaves unprocessed and a batch that failed transiently (a throttle means it processed none of it); the records are already published, and a table that throttles every round delays the ingest response by at most about 3 s at the defaults (eight 250 ms calls and the waits between them) before the commit is given up. `Release` is a `DeleteItem` conditional on the token and the pending state. Each call has a `Timeout` (250 ms) covering the SDK's retries (`MaxAttempts`, 3), which back off with full jitter under a ceiling capped at `Timeout/(2·(MaxAttempts−1))`, so a call's retries wait at most half its timeout and a throttled call fails on its last attempt's answer rather than on the deadline. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also reads a lease `<= 0` as `DefaultLease` and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index f5c22b59a..2181a1f55 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -320,8 +320,9 @@ type dynamoStore struct { // Reserve puts every key's pending item in parallel, each conditional on no // live item holding the key. A failed condition hands back the live item, // whose state says Duplicate or InFlight without a read, or Claimed when its -// token is the put's own. On any error every put that may have landed is -// released by its token. +// token is the put's own. On any error it releases, by token, every put that +// may have landed; what that undo misses (below) holds its key InFlight +// until the lease ends, as a crashed request's claim does. func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Duration) ([]Claim, error) { if len(keys) == 0 { return []Claim{}, nil @@ -338,9 +339,10 @@ func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Durati sent := make([]bool, len(keys)) // The first failure skips the puts not yet sent: the Reserve fails // either way, and a throttled table should not take the rest. A put - // already sent runs on ctx, not gctx, so it finishes and its outcome is - // known before the undo below; cancelled mid-flight, it could land after - // its release and hold the key for the lease. + // already sent runs on ctx, not gctx, so a sibling's failure never cuts + // it off: it answers before the undo below. The caller's cancellation or + // the put's own deadline can, and DynamoDB may then apply it after its + // release. g, gctx := errgroup.WithContext(ctx) g.SetLimit(s.d.cfg.ReserveConcurrency) for i, k := range keys { @@ -363,8 +365,9 @@ func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Durati }) } if err := g.Wait(); err != nil { - // A put that was sent and errored may still have landed; its token - // is known, and releasing a key it does not hold is a no-op. + // A sent put that errored may have landed anyway; releasing a key + // its token does not hold is a no-op. Best effort: a put applied + // after this, or a release that fails, lapses with the lease. var undo []Claim for i, c := range claims { if sent[i] && (c.Status == Claimed || c.Status == 0) { diff --git a/internal/dedupe/dynamodb_test.go b/internal/dedupe/dynamodb_test.go index 31f54f85e..1a30791be 100644 --- a/internal/dedupe/dynamodb_test.go +++ b/internal/dedupe/dynamodb_test.go @@ -681,7 +681,7 @@ func TestDynamo_ReleaseAttemptsEveryClaim(t *testing.T) { // One throttled put in a multi-key Reserve: the unsent puts are never sent, // and a sibling already sent runs to its answer before the undo releases it, -// so a put cannot land after its own release. +// so a sibling's failure never lets a put land after its own release. func TestDynamo_FailedMultiKeyReserve(t *testing.T) { t.Parallel() var mu sync.Mutex From 8b616b4a309d7c7e96f964817d8f57c30d8cf5e1 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 04:57:19 -0400 Subject: [PATCH 21/38] fix(config): cap dedupe.lease at 59.5s over the embedded queue The cap sat at the duplicate window itself (2m), which is the wrong boundary. The in-flight 503 answers with the whole lease as Retry-After, so a client that obeys it after a publish whose outcome it never learned republishes up to twice the lease after the claim, and DynamoDB rounds a claim's expiry up to the second. With mq.backend=embedded a lease is now refused unless twice it plus one second fits the 2m window. The 30s default is unchanged; 59s and 59.5s pass, 60s is refused. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- CHANGELOG.md | 2 +- config.yaml | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/configuration.mdx | 4 ++-- docs/src/content/docs/sdk/reference.md | 2 +- internal/config/backends.go | 15 ++++++++++----- internal/config/backends_test.go | 7 +++++-- 7 files changed, 21 insertions(+), 13 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index db2333b61..2abb1c9b6 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `2m`, the embedded queue's duplicate window), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`; no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and on every reload. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. +- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59.5s` with the embedded queue, so that twice the lease plus a second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish republishes up to twice the lease after the claim, and DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`; no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and on every reload. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. - **A DynamoDB dedupe backend** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. `dedupe.backend: dynamodb` selects it (entry above). New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/config.yaml b/config.yaml index 7107dd38c..0ca27bd6b 100644 --- a/config.yaml +++ b/config.yaml @@ -57,7 +57,7 @@ mq: backend: embedded # NATS JetStream under /nats dedupe: backend: pebble # Pebble under /pebble; or dynamodb (below) - lease: 30s # how long a claimed id stays pending; at most 2m with the embedded mq + lease: 30s # how long a claimed id stays pending; at most 59.5s with the embedded mq (2*lease + 1s within its 2m duplicate window) reserve_concurrency: 64 # parallel calls per Reserve/Commit/Release to a remote backend; no effect yet (ingest sends one id per call) # dynamodb: # read only when backend is dynamodb; credentials from the AWS SDK chain # table: wavehouse-dedupe-prod diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 1484f79db..934ee671e 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -120,7 +120,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — the implementation each layer runs on, the process's `roles`, resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). - **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load` when `Config.NeedsDataDir` says a selected backend keeps state there, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. -- **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` in `Validate`, between `validateRoles` and `validateTopology`), which refuses a value not on the list and names the ones that are. One rule spans two layers: `dedupe.lease` may not exceed the embedded MQ's 2m duplicate window (`embeddedDuplicateWindow`) while `mq.backend` is `embedded`. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns the valid combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), which `app.New` logs at `WARN`. +- **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` in `Validate`, between `validateRoles` and `validateTopology`), which refuses a value not on the list and names the ones that are. One rule spans two layers: while `mq.backend` is `embedded`, twice `dedupe.lease` plus one second must fit the embedded MQ's 2m duplicate window (`embeddedDuplicateWindow`), a cap of 59.5s (`maxEmbeddedLease`), because a client obeying the in-flight `503`'s `Retry-After` republishes up to twice the lease after the claim and DynamoDB rounds a claim's expiry up to the second. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns the valid combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), which `app.New` logs at `WARN`. - **config.go**, roles — `roles` (`[]Role`: `api`, `ingest`, `sweeper`; `AllRoles` by default; `Has(Role)`) picks which components `internal/app` wires, and `instance_id` names the process (`-<8 hex>` when empty, resolved in `Load`; today only logged at boot, and a distributed coordinator will record it as a lease's holder). `validateRoles` refuses an empty list, an empty entry, an unknown or a repeated role; `validateTopology` refuses a role set the backends cannot serve: any split over the embedded MQ, and a process with exactly one of `api` and `ingest` over a local cache. `NeedsDataDir` counts Pebble only for a process running `api`, and `Warnings` is empty without `api`, since only that role opens a cache it reads or a dedupe store. - **strict.go** — `rejectUnknownKeys`, the YAML half: re-reads the file as a generic tree and walks it against the struct's `yaml` tags, listing every key the struct doesn't declare. cleanenv itself is lenient by design, which is exactly wrong for boot config once keys have moved to the settings directory. - **persistence.go** — `WarnIfFreshDataDir` logs the startup `WARN` when `data_dir` is missing or empty (on a redeploy, the sign that the volume didn't persist); `LogStorageInitError` attaches the UID-65532 `permissionHint` to a NATS or Pebble open failure that looks like a permission denial — the same hint string `CheckDataDir` uses. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 7dbd31710..12c0eeb0f 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -56,7 +56,7 @@ Whether a tenant dedupes, and on which field, are settings-directory keys ([Dedu | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | -| `dedupe.lease` | `WH_DEDUPE_LEASE` | `30s` | How long a record's id stays claimed while the record is published. Another request carrying the same id meanwhile gets `503` with this as `Retry-After`, in whole seconds; a claim that is neither committed nor released, because its process died mid-publish, lapses after it. At most `2m` with `mq.backend: embedded`, the embedded queue's duplicate window: a longer lease refuses boot. A Go duration (`30s`, `1m`); `0` refuses boot. | +| `dedupe.lease` | `WH_DEDUPE_LEASE` | `30s` | How long a record's id stays claimed while the record is published. Another request carrying the same id meanwhile gets `503` with this as `Retry-After`, in whole seconds; a claim that is neither committed nor released, because its process died mid-publish, lapses after it. With `mq.backend: embedded`, twice the lease plus one second must fit the embedded queue's 2-minute duplicate window, so the lease is at most `59.5s`: a client that obeys `Retry-After` after a publish whose outcome it never learned republishes up to twice the lease after the claim, and DynamoDB rounds a claim's expiry up to the second. A longer lease refuses boot. A Go duration (`30s`, `45s`); `0` refuses boot. | | `dedupe.reserve_concurrency` | `WH_DEDUPE_RESERVE_CONCURRENCY` | `64` | The most parallel calls one Reserve, Commit or Release makes to a remote dedupe backend. Ingest sends one id per call today, so it has no effect yet; `pebble` ignores it. `0` refuses boot. | #### DynamoDB dedupe @@ -265,7 +265,7 @@ cache: dedupe: backend: pebble # in-process Pebble under /pebble; or dynamodb - lease: 30s # at most 2m with the embedded mq + lease: 30s # at most 59.5s with the embedded mq reserve_concurrency: 64 # dynamodb: # read only when backend is dynamodb # table: wavehouse-dedupe-prod diff --git a/docs/src/content/docs/sdk/reference.md b/docs/src/content/docs/sdk/reference.md index 5679d7852..390b0411a 100644 --- a/docs/src/content/docs/sdk/reference.md +++ b/docs/src/content/docs/sdk/reference.md @@ -40,7 +40,7 @@ The SDK **never throws** for anything the server returns — all API errors come | 502 | `clickhouse.misconfigured` | No | ClickHouse refused WaveHouse's own credentials or database, or the route to it is wrong (a redirect, or a `4xx` other than `408`/`413`/`429`, with no exception code) — an operator fix | | 502 | `clickhouse.response_too_large` | No | A raw-SQL (`wh.sql`) response over the 64 MiB cap | | 503 | `clickhouse.unavailable` | Yes | ClickHouse is down, unreachable or overloaded; `Retry-After: 5`, honored between attempts | -| 503 | `HTTP_503` | Yes | Service unavailable, a tenant whose settings folder was rejected, a schema not discovered yet, a tenant on no ClickHouse pool, a token sent while that tenant's JWKS has not been fetched yet (`token verifier not ready`, `Retry-After: 30`), or a record whose dedupe id another request is still publishing (`a request with the same dedupe id is in flight`, `Retry-After`: the 30 s dedupe lease). REST calls auto-retry, honoring `Retry-After` when the response carries one — so each attempt on those last two causes waits the 30 s; a stream re-dials on its own jittered backoff instead | +| 503 | `HTTP_503` | Yes | Service unavailable, a tenant whose settings folder was rejected, a schema not discovered yet, a tenant on no ClickHouse pool, a token sent while that tenant's JWKS has not been fetched yet (`token verifier not ready`, `Retry-After: 30`), or a record whose dedupe id another request is still publishing (`a request with the same dedupe id is in flight`, `Retry-After`: the server's dedupe lease, 30 s by default). REST calls auto-retry, honoring `Retry-After` when the response carries one — so each attempt on those last two causes waits that long; a stream re-dials on its own jittered backoff instead | | 0 | `NETWORK_ERROR` | Yes | Network failure (retried with exponential backoff) | | 0 | `ABORTED` | No | Request canceled via `AbortSignal` | | 0 | `SSE_CONNECT_ERROR` | No | Stream could not be started (e.g. a non-absolute `baseURL`) | diff --git a/internal/config/backends.go b/internal/config/backends.go index 493e9488c..5c1cad8b4 100644 --- a/internal/config/backends.go +++ b/internal/config/backends.go @@ -165,11 +165,16 @@ func checkBackend[T ~string](key, env string, got T, valid []T) error { return fmt.Errorf("%s (%s) %q is not a backend this build has; valid: %s", key, env, got, strings.Join(names, ", ")) } -// embeddedDuplicateWindow mirrors mq.EmbeddedDuplicateWindow, the embedded -// ingest stream's duplicate window (#613 F2). A lease longer than it would let -// the republish of a publish whose outcome was unknown land twice. +// embeddedDuplicateWindow is the embedded ingest stream's duplicate window, +// counted from the stored publish. const embeddedDuplicateWindow = 2 * time.Minute +// maxEmbeddedLease is the longest dedupe.lease that window covers: a client +// that obeys the in-flight 503's Retry-After (the whole lease) after a +// publish whose outcome it never learned republishes up to twice the lease +// after the claim, and DynamoDB rounds a claim's expiry up to the second. +const maxEmbeddedLease = (embeddedDuplicateWindow - time.Second) / 2 + // validateBackends checks every layer's backend and its sub-block, then the // rules that span two layers. func (c *Config) validateBackends() error { @@ -178,8 +183,8 @@ func (c *Config) validateBackends() error { return err } } - if c.MQ.Backend == MQEmbedded && c.Dedupe.Lease > embeddedDuplicateWindow { - return fmt.Errorf("dedupe.lease (WH_DEDUPE_LEASE) %s exceeds the embedded mq's %s duplicate window: a claim must lapse before the queue forgets the publish it guards", c.Dedupe.Lease, embeddedDuplicateWindow) + if c.MQ.Backend == MQEmbedded && c.Dedupe.Lease > maxEmbeddedLease { + return fmt.Errorf("dedupe.lease (WH_DEDUPE_LEASE) %s is over %s with the embedded mq: twice the lease plus 1s must fit its %s duplicate window, since a client obeying the in-flight 503's Retry-After republishes up to twice the lease after the claim", c.Dedupe.Lease, maxEmbeddedLease, embeddedDuplicateWindow) } return nil } diff --git a/internal/config/backends_test.go b/internal/config/backends_test.go index 56f0d083e..f4c01b0c7 100644 --- a/internal/config/backends_test.go +++ b/internal/config/backends_test.go @@ -259,7 +259,8 @@ func TestValidate_Dedupe(t *testing.T) { c.Dedupe.DynamoDB.Endpoint, c.Dedupe.DynamoDB.CreateTable = "http://localhost:8000", true }, ""}, {"the block is not read under pebble", func(c *Config) { c.Dedupe.DynamoDB = DedupeDynamoDBConfig{CreateTable: true} }, ""}, - {"lease at the duplicate window", func(c *Config) { c.Dedupe.Lease = 2 * time.Minute }, ""}, + {"lease just under a minute", func(c *Config) { c.Dedupe.Lease = 59 * time.Second }, ""}, + {"lease at the cap", func(c *Config) { c.Dedupe.Lease = 59*time.Second + 500*time.Millisecond }, ""}, {"create_table without an endpoint", func(c *Config) { dynamo(c) c.Dedupe.DynamoDB.CreateTable = true @@ -275,7 +276,9 @@ func TestValidate_Dedupe(t *testing.T) { {"negative lease", func(c *Config) { c.Dedupe.Lease = -time.Second }, "dedupe.lease (WH_DEDUPE_LEASE) must be > 0"}, {"zero concurrency", func(c *Config) { c.Dedupe.ReserveConcurrency = 0 }, "dedupe.reserve_concurrency (WH_DEDUPE_RESERVE_CONCURRENCY) must be > 0"}, {"negative concurrency", func(c *Config) { c.Dedupe.ReserveConcurrency = -1 }, "dedupe.reserve_concurrency"}, - {"lease past the duplicate window", func(c *Config) { c.Dedupe.Lease = 3 * time.Minute }, "exceeds the embedded mq's 2m0s duplicate window"}, + {"lease of a minute", func(c *Config) { c.Dedupe.Lease = time.Minute }, "dedupe.lease (WH_DEDUPE_LEASE) 1m0s is over 59.5s with the embedded mq: twice the lease plus 1s must fit its 2m0s duplicate window"}, + {"lease just past the cap", func(c *Config) { c.Dedupe.Lease = 59*time.Second + 500*time.Millisecond + 1 }, "is over 59.5s with the embedded mq"}, + {"lease at the duplicate window", func(c *Config) { c.Dedupe.Lease = 2 * time.Minute }, "is over 59.5s with the embedded mq"}, } for _, tc := range cases { t.Run(tc.name, func(t *testing.T) { From 96b99683ba98b88f48ecd3fa333471551bedd75a Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:01:23 -0400 Subject: [PATCH 22/38] docs(dedupe): the DynamoDB key example keeps '-' now keyenc keeps '-' since the merge from the base, so the example id evt-123 is stored as acme/clicks/evt-123, not acme/clicks/evt%2D123. Fixed in the Deployment page's table of attributes, which now also says what the escaping keeps, and in the CHANGELOG entry. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- CHANGELOG.md | 2 +- docs/src/content/docs/deployment.md | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 83a348443..2be43f451 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt%2D123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, retrying for up to eight jittered rounds both the items DynamoDB leaves unprocessed and a batch it throttled whole, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. +- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, retrying for up to eight jittered rounds both the items DynamoDB leaves unprocessed and a batch it throttled whole, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. - **Leases for work that must run in one process at a time, and the sweeper runs under one** (`internal/coord/` (new: `coord.go`, `local.go`, `elect.go`, `coordtest/`, + tests), `internal/app/{app,wire}.go` (+ tests), `internal/config/backends.go`, `internal/ingest/sweeper.go`, `.github/labeler.yml`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development,ingest-pipeline}.md`, `docs/src/content/docs/configuration.mdx`, `config.yaml`, `AGENTS.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `coord.Coordinator` hands out named leases (`TryAcquire` → a `Term` with a strictly increasing fencing `Token`, a `Done` channel and `Resign`; `ErrHeld` while another holder's term is live), and `coord.RunElected` runs a loop only while its process holds the lease, resigning when the loop returns and campaigning again every 2s. `coord.Local` is the in-process implementation, and `coordtest.Conformance` is the suite every implementation runs — the NATS KV backend that lets several replicas share one queue comes next. The sweeper now runs through `RunElected` under the `sweeper` lease; with the in-process coordinator the one process always holds it, so nothing changes for a single-process deployment beyond one `coord: elected` log line at startup. `coord.backend` now selects the coordinator (`local`, the only value), so a `config.Config` built without `config.Load` must name it as well as the other three layers' backends. diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 913ec2946..cff048d65 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -453,7 +453,7 @@ What the backend requires of the table: | Attribute | Type | Role | |---|---|---| -| `pk` | String | Partition key, and the only key: tenant, table and id as readable text, for example `acme/clicks/evt%2D123` (the table and id escaped the way NATS subject tokens are). No sort key. | +| `pk` | String | Partition key, and the only key: tenant, table and id as readable text, for example `acme/clicks/evt-123` (the table and id escaped the way NATS subject tokens are: letters, digits, `_` and `-` kept, every other byte written as `%XX`). No sort key. | | `st` | Number | `1` = pending claim, `2` = committed. | | `ex` | Number | Epoch seconds: the lease end while pending, the retention end once committed; absent = never expires. | | `tk` | Binary | The claim token that `Release` matches. | From a6944186d00d29b1fbe9fc7e8a39c07a5a1d2238 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:07:53 -0400 Subject: [PATCH 23/38] fix(app): the dedupe reload hook makes no DynamoDB call While the table check failed, every settings reload ran it from the AfterAdopt hook, which the registry runs under the lock that serializes reloads: 2.5s per reload against a hung endpoint at the default timeout, the hooks registered after it waiting, and a tenant the reload switched on answering ErrDisabled meanwhile, so its records published un-deduped. The hook now applies every store against the last check's result, which fails a switched-on store closed (ErrUnavailable) while the check has not passed, and wakes the background retry through a one-slot channel, so a reload still retries at once. Only boot and that loop run the check, and once passed it stays passed. A flat directory's boot is unchanged. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/configuration.mdx | 2 +- docs/src/content/docs/deployment.md | 2 +- docs/src/content/docs/settings-directory.mdx | 2 +- internal/app/dedupe_dynamodb_test.go | 82 ++++++++++++++++---- internal/app/wire.go | 70 ++++++++++------- 7 files changed, 118 insertions(+), 44 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 2abb1c9b6..d980aab25 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59.5s` with the embedded queue, so that twice the lease plus a second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish republishes up to twice the lease after the claim, and DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`; no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and on every reload. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. +- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59.5s` with the embedded queue, so that twice the lease plus a second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish republishes up to twice the lease after the claim, and DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`; no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and at once after every reload. A reload never waits on the table: it holds the lock that serializes reloads, so it applies each tenant's switch against the last check's result and only wakes the retry. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. - **A DynamoDB dedupe backend** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods. `Commit` is `BatchWriteItem`, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. `dedupe.backend: dynamodb` selects it (entry above). New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 934ee671e..0d0d7741c 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -93,7 +93,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability (after which each `Config.Warnings` line is logged at `WARN`), ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, the lease coordinator, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. The boot config's `roles` decide which of them a process wires: every process gets the settings registry, observability, the MQ, the coordinator, the reload triggers and a listener; `api` adds schema discovery, the dedupe stores, streaming, auth and the full router; `ingest` adds the ingest worker; `sweeper` adds the sweeper; the ClickHouse pools and the cache come with `api` or `ingest`. A process without `api` serves `api.NewOpsRouter` (probes, `/version`, the metrics path, and the settings reload behind the operator key alone, `wireOpsAuth`) on `server.port`. `config.Validate` refuses a role set the backends cannot serve (a split over the embedded MQ, or `api` without `ingest` and the reverse over a local cache), and `New` refuses a `Config` with no roles, which only one built without `config.Load` can have. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe`, `wireCoord` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where each would be redelivered every ack wait for as long as the tenant is away and would hold that tenant's ack floor, so the sweeper could purge none of its queue past it. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultPolicy` reads it through the registry, where a flat directory's tenant `0` is always served. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own, and runs only while its process holds the `sweeper` lease (`elected`, which wraps `coord.RunElected` over the coordinator `wireCoord` opens; `local`, the only `coord.backend`, keeps leases in the process, so the one process always holds it). The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) and then its table — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The `dynamodb` case builds the same `Stores` over `Dynamo.Tenant`, gated (`Factory.Gated`) on the table's check: boot runs `Dynamo.Check` (after `CreateTable`, when `dedupe.dynamodb.create_table` is on) whether or not any tenant has dedupe on, and a table that fails it follows the same rule as a Pebble instance that cannot open, with the check retried until it passes: by a background component that backs off from one second to thirty (a nested directory has no watcher), and by every reload. It has no Pebble gauges. `wireHTTP` hands the ingest handler `dedupe.lease` (`IngestHandler.DedupeLease`) whichever backend is chosen. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe`, `wireCoord` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where each would be redelivered every ack wait for as long as the tenant is away and would hold that tenant's ack floor, so the sweeper could purge none of its queue past it. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultPolicy` reads it through the registry, where a flat directory's tenant `0` is always served. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own, and runs only while its process holds the `sweeper` lease (`elected`, which wraps `coord.RunElected` over the coordinator `wireCoord` opens; `local`, the only `coord.backend`, keeps leases in the process, so the one process always holds it). The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) and then its table — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The `dynamodb` case builds the same `Stores` over `Dynamo.Tenant`, gated (`Factory.Gated`) on the table's check: boot runs `Dynamo.Check` (after `CreateTable`, when `dedupe.dynamodb.create_table` is on) whether or not any tenant has dedupe on, and a table that fails it follows the same rule as a Pebble instance that cannot open, with the check retried until it passes by a background component that backs off from one second to thirty (a nested directory has no watcher). The `AfterAdopt` hook never runs the check, since it holds the lock that serializes reloads: it applies every store against the last check's result, so a tenant a reload switches on fails closed meanwhile, and wakes the retry, so a reload still retries at once. It has no Pebble gauges. `wireHTTP` hands the ingest handler `dedupe.lease` (`IngestHandler.DedupeLease`) whichever backend is chosen. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 12c0eeb0f..2fd31f258 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -61,7 +61,7 @@ Whether a tenant dedupes, and on which field, are settings-directory keys ([Dedu #### DynamoDB dedupe -Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (String) alone, and TTL off on `ex` is logged as a warning. A table that fails the check, or cannot be reached, refuses boot with a flat settings directory. With a nested one, the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and the check is retried in the background (backing off from one second to thirty) and on every reload, so a table that comes good is picked up without a restart. The check runs in every process running the `api` [role](#process-roles), the one that opens the dedupe stores, whether or not any tenant has dedupe on. +Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (String) alone, and TTL off on `ex` is logged as a warning. A table that fails the check, or cannot be reached, refuses boot with a flat settings directory. With a nested one, the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and the check is retried in the background, backing off from one second to thirty, and at once after every reload, so a table that comes good is picked up without a restart. A reload never waits on the table: it applies each tenant's switch against the last check's result, so a tenant it switches on fails closed until the retry passes. The check runs in every process running the `api` [role](#process-roles), the one that opens the dedupe stores, whether or not any tenant has dedupe on. | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index cf617b0de..f2100f4cb 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -512,7 +512,7 @@ dedupe: region: us-east-1 # or leave empty for AWS_REGION ``` -or `WH_DEDUPE_BACKEND=dynamodb`, `WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod`. A table that is missing, has the wrong key schema, or cannot be reached with the pod's credentials refuses boot over a flat settings directory; over a nested one the pod boots, every tenant with dedupe on fails its ingest closed, and the check is retried in the background (backing off from one second to thirty) and on every reload. No region at all (neither `region` nor one from the SDK chain: `AWS_REGION`, `AWS_DEFAULT_REGION` or a profile) refuses boot in both shapes. The check runs in every pod running the `api` [role](/configuration#process-roles), whether or not any tenant has `dedupe.enabled` on; a pod without it opens no dedupe store. The per-tenant switch stays in each tenant's `config.json`. +or `WH_DEDUPE_BACKEND=dynamodb`, `WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod`. A table that is missing, has the wrong key schema, or cannot be reached with the pod's credentials refuses boot over a flat settings directory; over a nested one the pod boots, every tenant with dedupe on fails its ingest closed, and the check is retried in the background (backing off from one second to thirty, and at once after every reload, which never waits on the table). No region at all (neither `region` nor one from the SDK chain: `AWS_REGION`, `AWS_DEFAULT_REGION` or a profile) refuses boot in both shapes. The check runs in every pod running the `api` [role](/configuration#process-roles), whether or not any tenant has `dedupe.enabled` on; a pod without it opens no dedupe store. The per-tenant switch stays in each tenant's `config.json`. For development against dynamodb-local, set `dedupe.dynamodb.endpoint` (for example `http://localhost:8000`) and `create_table: true`, and give the SDK any static credentials (`AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`) and a region. `create_table` without an `endpoint` refuses boot. diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index dcba15d4e..b6f5397ee 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -185,7 +185,7 @@ What stays in boot config is only what cannot change under a running process — Every per-tenant dedupe knob lives here. Where the seen ids are kept (`dedupe.backend`) and how long a claim is held (`dedupe.lease`) are [boot config](/configuration#dedupe), the same for every tenant. The switch and its fields are resolved per record from one snapshot (table override → global value): -- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store (in the embedded Pebble instance at `/pebble`, or its share of the DynamoDB table under `dedupe.backend: dynamodb`), so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) with `dedupe.backend: pebble`, every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. Under `dedupe.backend: dynamodb` the table check plays the instance's part: the table is checked whether or not any tenant's switch is on, and a table that fails it fails every tenant with dedupe on closed until the check, retried in the background and on every reload, passes ([Configuration](/configuration#dynamodb-dedupe)). +- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store (in the embedded Pebble instance at `/pebble`, or its share of the DynamoDB table under `dedupe.backend: dynamodb`), so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) with `dedupe.backend: pebble`, every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. Under `dedupe.backend: dynamodb` the table check plays the instance's part: the table is checked whether or not any tenant's switch is on, and a table that fails it fails every tenant with dedupe on closed until the check, retried in the background and at once after every reload, passes ([Configuration](/configuration#dynamodb-dedupe)). - `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes once escaped (every byte but an ASCII letter, digit, `_` or `-` takes three) is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its lease (`dedupe.lease`, 30 seconds by default), and a later retry of it is accepted again — counted by `wavehouse_dedupe_commit_failed_total`, which should stay at zero. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. - `dedupe.tables.
.{id_field, require_id}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. diff --git a/internal/app/dedupe_dynamodb_test.go b/internal/app/dedupe_dynamodb_test.go index 6bac9dc23..f3fb1e598 100644 --- a/internal/app/dedupe_dynamodb_test.go +++ b/internal/app/dedupe_dynamodb_test.go @@ -23,10 +23,11 @@ import ( // fakeDynamo answers the DynamoDB JSON protocol for one table, enough for // boot's check, the dev create path, and a claim and its commit. Whether the -// table exists is the test's to switch. +// table exists, and whether the endpoint hangs, are the test's to switch. type fakeDynamo struct { mu sync.Mutex exists bool + hangs bool calls []string } @@ -36,15 +37,24 @@ func (f *fakeDynamo) setExists(v bool) { f.exists = v } -func (f *fakeDynamo) called(op string) bool { +func (f *fakeDynamo) setHangs(v bool) { f.mu.Lock() defer f.mu.Unlock() + f.hangs = v +} + +func (f *fakeDynamo) called(op string) bool { return f.count(op) > 0 } + +func (f *fakeDynamo) count(op string) int { + f.mu.Lock() + defer f.mu.Unlock() + n := 0 for _, c := range f.calls { if c == op { - return true + n++ } } - return false + return n } func (f *fakeDynamo) ServeHTTP(w http.ResponseWriter, r *http.Request) { @@ -55,8 +65,12 @@ func (f *fakeDynamo) ServeHTTP(w http.ResponseWriter, r *http.Request) { if op == "CreateTable" { f.exists = true } - exists := f.exists + exists, hangs := f.exists, f.hangs f.mu.Unlock() + if hangs { + <-r.Context().Done() + return + } w.Header().Set("Content-Type", "application/x-amz-json-1.0") if !exists { w.WriteHeader(http.StatusBadRequest) @@ -144,10 +158,10 @@ func TestNew_DynamoDBDedupeTableMissing(t *testing.T) { }) } }) - t.Run("nested fails closed until a reload passes the check", func(t *testing.T) { + t.Run("nested fails closed", func(t *testing.T) { root := writeNestedSettings(t, map[string]map[string]any{"acme": dedupeOn, "globex": nil}) cfg := testConfig(t, root) - fake := dynamoConfig(t, cfg, false) + dynamoConfig(t, cfg, false) a := newApp(t, cfg, Options{}) acme := a.dedup.For("acme") @@ -156,15 +170,57 @@ func TestNew_DynamoDBDedupeTableMissing(t *testing.T) { require.ErrorIs(t, err, dedupe.ErrUnavailable, "switched on, table missing: ingest fails closed") _, err = dedupetest.Mark(t.Context(), a.dedup.For("globex"), eventKey) require.ErrorIs(t, err, dedupe.ErrDisabled) - - fake.setExists(true) - a.tenants.Reload("test") - assert.True(t, acme.Open(), "the reload checked again and opened the store") - _, err = dedupetest.Mark(context.Background(), acme, eventKey) - require.NoError(t, err) }) } +// The reload hook runs under the lock that serializes reloads, so it never +// calls DynamoDB: against a table that hangs, a reload returns at once, and a +// tenant it switches on fails closed rather than publishing un-deduped. +func TestReload_DynamoDBDedupeMakesNoTableCall(t *testing.T) { + root := writeNestedSettings(t, map[string]map[string]any{"acme": nil}) + cfg := testConfig(t, root) + fake := dynamoConfig(t, cfg, false) + a := newApp(t, cfg, Options{}) + fake.setHangs(true) + before := fake.count("DescribeTable") + + rewriteSettings(t, filepath.Join(root, "acme"), dedupeOn) + start := time.Now() + a.tenants.Reload("test") + assert.Less(t, time.Since(start), time.Second, "a check would wait out its 2.5s deadline") + assert.Equal(t, before, fake.count("DescribeTable"), "the reload made no table call") + + acme := a.dedup.For("acme") + assert.False(t, acme.Open()) + _, err := dedupetest.Mark(t.Context(), acme, eventKey) + require.ErrorIs(t, err, dedupe.ErrUnavailable, "switched on while the table fails: closed, not ErrDisabled") +} + +// A reload wakes the background retry rather than running the check itself. +// The retry's first timed attempt is a second after it starts, and a timer +// never fires early, so an open sooner than that is the reload's doing. +func TestRun_DynamoDBDedupeReloadWakesTheRetry(t *testing.T) { + root := writeNestedSettings(t, map[string]map[string]any{"acme": dedupeOn}) + cfg := testConfig(t, root) + fake := dynamoConfig(t, cfg, false) + var lc net.ListenConfig + ln, err := lc.Listen(t.Context(), "tcp", "127.0.0.1:0") + require.NoError(t, err) + a := newApp(t, cfg, Options{Listener: ln}) + acme := a.dedup.For("acme") + require.False(t, acme.Open()) + + start := time.Now() + _, stop := runApp(t, a, ln) + fake.setExists(true) + a.tenants.Reload("test") + require.Eventually(t, acme.Open, 5*time.Second, 5*time.Millisecond) + assert.Less(t, time.Since(start), time.Second, "opened before the first timed retry") + _, err = dedupetest.Mark(context.Background(), acme, eventKey) + require.NoError(t, err) + require.NoError(t, stop()) +} + // A nested directory has no watcher, so a table that comes good is picked up // by the background retry, not only by a reload someone has to send. func TestRun_DynamoDBDedupeRetriesTheTableCheck(t *testing.T) { diff --git a/internal/app/wire.go b/internal/app/wire.go index 089de185d..08dd37636 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -545,7 +545,10 @@ var errDynamoUnchecked = errors.New("dedupe: dynamodb table not checked yet") // store closed, so its ingest fails closed. Unlike a local disk, a remote // table's failure is usually brief (a throttle, credentials not yet issued // mid-rollout), and a nested directory has no watcher to reload it, so the -// check is also retried in the background, with backoff, until it passes. +// check is then retried in the background, with backoff, until it passes. +// The check is network I/O, so the AfterAdopt hook never runs it: the hook +// holds the lock that serializes reloads. It applies every store against the +// last check's result and wakes the retry, so a reload still retries at once. func (a *App) wireDynamoDedupe(ctx context.Context) error { c := a.cfg.Dedupe.DynamoDB d, err := dedupe.NewDynamo(ctx, dedupe.DynamoConfig{ @@ -557,72 +560,87 @@ func (a *App) wireDynamoDedupe(ctx context.Context) error { return err } var mu sync.Mutex - state := errDynamoUnchecked // nil once the table has passed - check := func(ctx context.Context) error { + state := errDynamoUnchecked // nil once the table has passed, for good + ready := func() error { mu.Lock() defer mu.Unlock() - if state == nil { - return nil - } - if c.CreateTable { - if state = d.CreateTable(ctx); state != nil { - return state - } - } - state = d.Check(ctx) return state } - ready := func() error { + // check is only ever run by boot, then by the retry loop, one at a time. + check := func(ctx context.Context) error { + var err error + if c.CreateTable { + err = d.CreateTable(ctx) + } + if err == nil { + err = d.Check(ctx) + } mu.Lock() defer mu.Unlock() + if state != nil { + state = err + } return state } stores := dedupe.NewStores(dedupe.Factory(d.Tenant).Gated(ready)) a.dedup = stores a.add(component{name: "dedupe", close: withoutContext(stores.Close)}) - var reconciling sync.Mutex // the hook and the retry loop both reconcile - reconcile := func(ctx context.Context) error { + var reconciling sync.Mutex // the hook and the retry loop both apply + apply := func() { reconciling.Lock() defer reconciling.Unlock() if err := stores.Retain(a.served); err != nil { slog.Error("dedupe store close failed", "error", err) } - checkErr := check(ctx) - if checkErr != nil { - slog.Error("dedupe: dynamodb table check failed; ingest with dedupe on fails closed until a reload passes it", - "table", c.Table, "error", checkErr) - } for id, store := range a.tenants.All() { m := stores.For(id) enabled := store.DedupeEnabled() wasOpen := m.Open() - // The one failure an open has is the check's, logged above. + // The one failure an open has is the check's, logged where it ran. _ = m.Apply(enabled) if m.Open() != wasOpen { slog.Info("dedupe store reconciled with settings", "tenant", id, "enabled", enabled) } } - return checkErr } - a.tenants.AfterAdopt(func([]tenant.ID) { _ = reconcile(a.stopCtx) }) - if err := reconcile(ctx); err != nil { + retry := make(chan struct{}, 1) + a.tenants.AfterAdopt(func([]tenant.ID) { + apply() + if ready() != nil { + select { + case retry <- struct{}{}: + default: // a retry is already due + } + } + }) + if err := check(ctx); err != nil { if !a.tenants.Nested() { return fmt.Errorf("dedupe open: %w", err) } + slog.Error("dedupe: dynamodb table check failed; ingest with dedupe on fails closed until a reload passes it", + "table", c.Table, "error", err) a.add(component{name: "dedupe table check", run: func(ctx context.Context) error { for wait := time.Second; ready() != nil; wait = min(2*wait, 30*time.Second) { select { case <-ctx.Done(): return nil case <-time.After(wait): + case <-retry: } - if reconcile(ctx) == nil { - slog.Info("dedupe: dynamodb table check passed", "table", c.Table) + if err := check(ctx); err != nil { + if ctx.Err() == nil { + slog.Error("dedupe: dynamodb table check failed; ingest with dedupe on fails closed until a reload passes it", + "table", c.Table, "error", err) + } + continue } + slog.Info("dedupe: dynamodb table check passed", "table", c.Table) + apply() } return nil }}) } + apply() return nil } From 9b8ea3c5b16693218acfe7984bae9a3b517809e9 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:08:25 -0400 Subject: [PATCH 24/38] fix(app): say what a failed dedupe table check leads to The log said ingest with dedupe on fails closed until a reload passes the check, which is wrong in both shapes: a flat directory refuses boot, and a nested one retries the check in the background. The boot line and the retry's line now say that it is retried and ingest fails closed meanwhile. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- internal/app/wire.go | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/internal/app/wire.go b/internal/app/wire.go index 08dd37636..a65d1bf8d 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -617,7 +617,7 @@ func (a *App) wireDynamoDedupe(ctx context.Context) error { if !a.tenants.Nested() { return fmt.Errorf("dedupe open: %w", err) } - slog.Error("dedupe: dynamodb table check failed; ingest with dedupe on fails closed until a reload passes it", + slog.Error("dedupe: dynamodb table check failed; ingest with dedupe on fails closed while it is retried", "table", c.Table, "error", err) a.add(component{name: "dedupe table check", run: func(ctx context.Context) error { for wait := time.Second; ready() != nil; wait = min(2*wait, 30*time.Second) { @@ -629,7 +629,7 @@ func (a *App) wireDynamoDedupe(ctx context.Context) error { } if err := check(ctx); err != nil { if ctx.Err() == nil { - slog.Error("dedupe: dynamodb table check failed; ingest with dedupe on fails closed until a reload passes it", + slog.Error("dedupe: dynamodb table check failed again; ingest with dedupe on still fails closed", "table", c.Table, "error", err) } continue From 0425a555f9e03bc281c8f507acdce388846d03ae Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:17:32 -0400 Subject: [PATCH 25/38] fix(dedupe): never size the DynamoDB idle pool below the SDK default MaxIdleConnsPerHost was set to ReserveConcurrency outright, so a small ReserveConcurrency (a boot key soon) shrank the pool below the SDK's default of 10 per host. It is now the larger of the two, as MaxIdleConns already was. TestNewDynamo_SizesTheIdlePool adds 4 and asserts both limits stay at or above the SDK's defaults; it failed at 4 and 8 before. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- docs/src/content/docs/architecture.md | 2 +- internal/dedupe/dynamodb.go | 9 +++++---- internal/dedupe/dynamodb_test.go | 4 +++- 3 files changed, 9 insertions(+), 6 deletions(-) diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index cbc0d55cf..7935f25c8 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -137,7 +137,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key every backend stores, as text: `/
/` (for example `acme/clicks/evt-123`, [#222](https://github.com/Wave-RF/WaveHouse/issues/222)). The table and id are escaped and joined (`keyenc.AppendJoin`) by `internal/keyenc` — the escaping NATS subject tokens use — which never writes `/`, and a tenant id cannot hold one, so a table name may hold any byte, NUL included, and neither tenants nor tables share ids; a key is ASCII, so it reads as-is in a console and is a valid DynamoDB String. An id whose escaped form is over 1,024 bytes is stored as `#` plus its SHA-256 in hex (`#` is never written by the escaping), counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. -- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency` (64), with as many idle connections kept per host so a wide `Reserve` reuses them rather than dial. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read, or `Claimed` when it is the put's own pending item (same token): an SDK retry of an attempt DynamoDB applied but whose answer was lost. If any put errors, every put that may have landed is released by its token. A put already sent answers before that undo when a sibling fails, but one cut off by the caller's cancellation or its call deadline can still be applied after its release; it then holds its key `InFlight` until the lease ends, as a crashed request's claim does. `Commit` is `BatchWriteItem`, 25 at a time, retrying with jittered backoff, for up to eight rounds, both the items DynamoDB leaves unprocessed and a batch that failed transiently (a throttle means it processed none of it); the records are already published, and a table that throttles every round delays the ingest response by at most about 3 s at the defaults (eight 250 ms calls and the waits between them) before the commit is given up. `Release` is a `DeleteItem` conditional on the token and the pending state. Each call has a `Timeout` (250 ms) covering the SDK's retries (`MaxAttempts`, 3), which back off with full jitter under a ceiling capped at `Timeout/(2·(MaxAttempts−1))`, so a call's retries wait at most half its timeout and a throttled call fails on its last attempt's answer rather than on the deadline. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. +- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency` (64), with at least as many idle connections kept per host so a wide `Reserve` reuses them rather than dial. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read, or `Claimed` when it is the put's own pending item (same token): an SDK retry of an attempt DynamoDB applied but whose answer was lost. If any put errors, every put that may have landed is released by its token. A put already sent answers before that undo when a sibling fails, but one cut off by the caller's cancellation or its call deadline can still be applied after its release; it then holds its key `InFlight` until the lease ends, as a crashed request's claim does. `Commit` is `BatchWriteItem`, 25 at a time, retrying with jittered backoff, for up to eight rounds, both the items DynamoDB leaves unprocessed and a batch that failed transiently (a throttle means it processed none of it); the records are already published, and a table that throttles every round delays the ingest response by at most about 3 s at the defaults (eight 250 ms calls and the waits between them) before the commit is given up. `Release` is a `DeleteItem` conditional on the token and the pending state. Each call has a `Timeout` (250 ms) covering the SDK's retries (`MaxAttempts`, 3), which back off with full jitter under a ceiling capped at `Timeout/(2·(MaxAttempts−1))`, so a call's retries wait at most half its timeout and a throttled call fails on its last attempt's answer rather than on the deadline. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also reads a lease `<= 0` as `DefaultLease` and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index 2181a1f55..cc3321f7c 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -81,7 +81,8 @@ type DynamoConfig struct { // the client after throttles. RetryMode string // ReserveConcurrency bounds the parallel calls one Reserve, Commit or - // Release makes, and sizes the client's idle connection pool to match. + // Release makes, and sizes the client's idle connection pool to match + // (never below the SDK's default of 10 per host). // 0 = 64. ReserveConcurrency int } @@ -158,11 +159,11 @@ func NewDynamo(ctx context.Context, cfg DynamoConfig, extra ...func(*config.Load } // newHTTPClient keeps an idle connection for every call one Reserve can have -// in flight: with the SDK's default of 10 per host, a wide Reserve would dial -// most of its puts afresh. +// in flight, and never fewer than the SDK's defaults: with its 10 per host, a +// wide Reserve would dial most of its puts afresh. func newHTTPClient(cfg DynamoConfig) *awshttp.BuildableClient { return awshttp.NewBuildableClient().WithTransportOptions(func(tr *http.Transport) { - tr.MaxIdleConnsPerHost = cfg.ReserveConcurrency + tr.MaxIdleConnsPerHost = max(tr.MaxIdleConnsPerHost, cfg.ReserveConcurrency) tr.MaxIdleConns = max(tr.MaxIdleConns, cfg.ReserveConcurrency) }) } diff --git a/internal/dedupe/dynamodb_test.go b/internal/dedupe/dynamodb_test.go index 1a30791be..03a7196f8 100644 --- a/internal/dedupe/dynamodb_test.go +++ b/internal/dedupe/dynamodb_test.go @@ -585,7 +585,7 @@ func TestDynamo_ThrottledCallEndsOnItsLastAttempt(t *testing.T) { // extra (TestDynamo_ThrottledCallEndsOnItsLastAttempt's) replaces it. func TestNewDynamo_SizesTheIdlePool(t *testing.T) { t.Parallel() - for _, n := range []int{0, 8, 200} { + for _, n := range []int{0, 4, 8, 200} { d, err := NewDynamo(t.Context(), DynamoConfig{Table: "dedupe", Region: "us-east-1", ReserveConcurrency: n}) require.NoError(t, err) client, ok := d.api.(*dynamodb.Client).Options().HTTPClient.(*awshttp.BuildableClient) @@ -593,6 +593,8 @@ func TestNewDynamo_SizesTheIdlePool(t *testing.T) { tr := client.GetTransport() assert.GreaterOrEqual(t, tr.MaxIdleConnsPerHost, d.cfg.ReserveConcurrency, "ReserveConcurrency %d", n) assert.GreaterOrEqual(t, tr.MaxIdleConns, d.cfg.ReserveConcurrency, "ReserveConcurrency %d", n) + assert.GreaterOrEqual(t, tr.MaxIdleConnsPerHost, awshttp.DefaultHTTPTransportMaxIdleConnsPerHost, "never below the SDK's default: ReserveConcurrency %d", n) + assert.GreaterOrEqual(t, tr.MaxIdleConns, awshttp.DefaultHTTPTransportMaxIdleConns, "ReserveConcurrency %d", n) } } From 5c64ace0fa3f47e121d4103d4c62d2d28b0365b5 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 05:17:55 -0400 Subject: [PATCH 26/38] docs(dedupe): list the short-circuit counter with the DynamoDB metrics The Metrics bullet on the Deployment page named three of the backend's four metrics; wavehouse_dedupe_dynamodb_short_circuits_total appeared only in passing one bullet earlier. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- docs/src/content/docs/deployment.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index cff048d65..e2be26163 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -510,7 +510,7 @@ data "aws_iam_policy_document" "wavehouse_dedupe" { - **Point-in-time recovery** is not needed. The table records which ids have been seen, so losing it produces duplicate rows, not lost events. - **Cost:** every new event is two writes (the claim, then the commit), and a duplicate is one. On-demand, that is about $1.25 per million new events in us-east-1. Provisioned capacity with auto scaling is cheaper once traffic is steady. Storage is the other line: every distinct id stays in the table (see TTL above), at DynamoDB's per-GB-month rate. - **One table serves every tenant,** so one tenant's burst can throttle the rest. A throttled or unreachable table fails the ingest request closed rather than publishing un-deduped. After five throttled or unreachable claims in a row within one second, the backend stops calling the table for a second and fails every tenant's dedupe requests immediately (`wavehouse_dedupe_dynamodb_short_circuits_total`). A duplicate or in-flight answer is not a failure and resets the count. -- **Metrics:** `wavehouse_dedupe_dynamodb_requests_total{op,outcome}`, `wavehouse_dedupe_dynamodb_request_duration_seconds{op}`, `wavehouse_dedupe_dynamodb_unprocessed_items_total`. The table's own CloudWatch metrics `ThrottledRequests`, `SystemErrors` and `ConsumedWriteCapacityUnits` are worth alerting on too. +- **Metrics:** `wavehouse_dedupe_dynamodb_requests_total{op,outcome}`, `wavehouse_dedupe_dynamodb_request_duration_seconds{op}`, `wavehouse_dedupe_dynamodb_unprocessed_items_total`, `wavehouse_dedupe_dynamodb_short_circuits_total`. The table's own CloudWatch metrics `ThrottledRequests`, `SystemErrors` and `ConsumedWriteCapacityUnits` are worth alerting on too. ## Upgrading across the v2 ingest envelope From 3513823459cec691fb42d7105d39d94bf545b447 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:27:10 -0400 Subject: [PATCH 27/38] fix(config): cap the embedded-mq dedupe lease by lease+ceil(lease)+1s The old check allowed a lease up to 59.5s (2*lease+1s <= 2m), but the DynamoDB claim it bounds rounds its expiry up to the next whole second, and the in-flight 503 it drives can go out a second late. The true worst case is lease + ceil(lease) + 1s, which only stays under the embedded queue's 2-minute duplicate window through 59s exactly: 59.5s (and anything else over 59s) already crosses it once ceil(lease) steps to the next second. Co-Authored-By: Claude Opus 5.5 (1M context) --- internal/config/backends.go | 31 ++++++++++++++++++++++++------- internal/config/backends_test.go | 10 +++++----- 2 files changed, 29 insertions(+), 12 deletions(-) diff --git a/internal/config/backends.go b/internal/config/backends.go index 42d7b705d..b2dc0dd00 100644 --- a/internal/config/backends.go +++ b/internal/config/backends.go @@ -170,11 +170,22 @@ func checkBackend[T ~string](key, env string, got T, valid []T) error { // counted from the stored publish. const embeddedDuplicateWindow = 2 * time.Minute -// maxEmbeddedLease is the longest dedupe.lease that window covers: a client -// that obeys the in-flight 503's Retry-After (the whole lease) after a -// publish whose outcome it never learned republishes up to twice the lease -// after the claim, and DynamoDB rounds a claim's expiry up to the second. -const maxEmbeddedLease = (embeddedDuplicateWindow - time.Second) / 2 +// maxEmbeddedLease is the longest dedupe.lease the duplicate window covers — +// the largest whole second satisfying the rule below. It is informational +// only: validateBackends checks the rule itself, not this constant, since +// the rule's ceiling steps at each whole second rather than moving linearly +// with the lease. +const maxEmbeddedLease = 59 * time.Second + +// ceilSecond rounds d up to the next whole second, as a DynamoDB claim's +// expiry does (epoch seconds, rounded up) — so a claim taken out just before +// the tick it is stamped with can stay live up to a second past the lease. +func ceilSecond(d time.Duration) time.Duration { + if r := d % time.Second; r != 0 { + d += time.Second - r + } + return d +} // validateBackends checks every layer's backend and its sub-block, then the // rules that span two layers. @@ -184,8 +195,14 @@ func (c *Config) validateBackends() error { return err } } - if c.MQ.Backend == MQEmbedded && c.Dedupe.Lease > maxEmbeddedLease { - return fmt.Errorf("dedupe.lease (WH_DEDUPE_LEASE) %s is over %s with the embedded mq: twice the lease plus 1s must fit its %s duplicate window, since a client obeying the in-flight 503's Retry-After republishes up to twice the lease after the claim", c.Dedupe.Lease, maxEmbeddedLease, embeddedDuplicateWindow) + // A client obeying the in-flight 503's Retry-After (the whole lease) + // republishes at t0+lease at the earliest. But a claim can outlive its + // own lease by up to a second (DynamoDB rounds expiry up to the second), + // so the last such 503 can go out at t0+lease+1s, and the republish it + // asks for lands at t0+lease+1s+ceil(lease). That must still fall inside + // the embedded duplicate window: lease + ceil(lease) + 1s <= 2m. + if worst := c.Dedupe.Lease + ceilSecond(c.Dedupe.Lease) + time.Second; c.MQ.Backend == MQEmbedded && worst > embeddedDuplicateWindow { + return fmt.Errorf("dedupe.lease (WH_DEDUPE_LEASE) %s is over %s with the embedded mq: lease + ceil(lease) + 1s (%s) must fit its %s duplicate window, since a client obeying the in-flight 503's Retry-After can republish that late", c.Dedupe.Lease, maxEmbeddedLease, worst, embeddedDuplicateWindow) } return nil } diff --git a/internal/config/backends_test.go b/internal/config/backends_test.go index f4c01b0c7..fd2005a88 100644 --- a/internal/config/backends_test.go +++ b/internal/config/backends_test.go @@ -259,8 +259,8 @@ func TestValidate_Dedupe(t *testing.T) { c.Dedupe.DynamoDB.Endpoint, c.Dedupe.DynamoDB.CreateTable = "http://localhost:8000", true }, ""}, {"the block is not read under pebble", func(c *Config) { c.Dedupe.DynamoDB = DedupeDynamoDBConfig{CreateTable: true} }, ""}, - {"lease just under a minute", func(c *Config) { c.Dedupe.Lease = 59 * time.Second }, ""}, - {"lease at the cap", func(c *Config) { c.Dedupe.Lease = 59*time.Second + 500*time.Millisecond }, ""}, + {"lease at the cap", func(c *Config) { c.Dedupe.Lease = 59 * time.Second }, ""}, + {"lease just past the cap", func(c *Config) { c.Dedupe.Lease = 59*time.Second + 100*time.Millisecond }, "is over 59s with the embedded mq"}, {"create_table without an endpoint", func(c *Config) { dynamo(c) c.Dedupe.DynamoDB.CreateTable = true @@ -276,9 +276,9 @@ func TestValidate_Dedupe(t *testing.T) { {"negative lease", func(c *Config) { c.Dedupe.Lease = -time.Second }, "dedupe.lease (WH_DEDUPE_LEASE) must be > 0"}, {"zero concurrency", func(c *Config) { c.Dedupe.ReserveConcurrency = 0 }, "dedupe.reserve_concurrency (WH_DEDUPE_RESERVE_CONCURRENCY) must be > 0"}, {"negative concurrency", func(c *Config) { c.Dedupe.ReserveConcurrency = -1 }, "dedupe.reserve_concurrency"}, - {"lease of a minute", func(c *Config) { c.Dedupe.Lease = time.Minute }, "dedupe.lease (WH_DEDUPE_LEASE) 1m0s is over 59.5s with the embedded mq: twice the lease plus 1s must fit its 2m0s duplicate window"}, - {"lease just past the cap", func(c *Config) { c.Dedupe.Lease = 59*time.Second + 500*time.Millisecond + 1 }, "is over 59.5s with the embedded mq"}, - {"lease at the duplicate window", func(c *Config) { c.Dedupe.Lease = 2 * time.Minute }, "is over 59.5s with the embedded mq"}, + {"lease of a minute", func(c *Config) { c.Dedupe.Lease = time.Minute }, "dedupe.lease (WH_DEDUPE_LEASE) 1m0s is over 59s with the embedded mq: lease + ceil(lease) + 1s (2m1s) must fit its 2m0s duplicate window"}, + {"lease at the old 59.5s cap", func(c *Config) { c.Dedupe.Lease = 59*time.Second + 500*time.Millisecond }, "is over 59s with the embedded mq"}, + {"lease at the duplicate window", func(c *Config) { c.Dedupe.Lease = 2 * time.Minute }, "is over 59s with the embedded mq"}, } for _, tc := range cases { t.Run(tc.name, func(t *testing.T) { From 2c42f0add33c1b5e9ebfbb1d929b5f9298db91d3 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:27:31 -0400 Subject: [PATCH 28/38] docs(dedupe): sweep the lease cap, the table check's deadline, and what a reload waits on MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Every remaining 59.5s (config.yaml, CHANGELOG.md, configuration.mdx, architecture.md) now reads 59s with the lease+ceil(lease)+1s rule spelled out, matching the previous commit's config change. Adds the boot/background table check's own 10x-timeout deadline (2.5s by default) to the dedupe.dynamodb.timeout row, since that call is not one of the per-request calls the row otherwise describes. Marks the development.md tree line "Pebble or DynamoDB" now that this PR makes the backend selectable, replacing the prior "not yet selectable at boot" framing that architecture.md's dynamodb.go entry already dropped. Qualifies "a reload never waits on the table" everywhere it appears (configuration.mdx, deployment.md, CHANGELOG.md, architecture.md's wire.go paragraph): true for a tenant whose dedupe.enabled did not change, thanks to Managed.Apply's no-op fast path settling under a read lock alone — but switching a tenant's dedupe off is a genuine transition, which takes the write lock and so waits for that tenant's in-flight Reserve/Commit/Release calls to finish first. Co-Authored-By: Claude Opus 5.5 (1M context) --- CHANGELOG.md | 2 +- config.yaml | 2 +- docs/src/content/docs/architecture.md | 4 ++-- docs/src/content/docs/configuration.mdx | 8 ++++---- docs/src/content/docs/deployment.md | 2 +- docs/src/content/docs/development.md | 2 +- 6 files changed, 10 insertions(+), 10 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 5ec73acb3..747fbd8fb 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59.5s` with the embedded queue, so that twice the lease plus a second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish republishes up to twice the lease after the claim, and DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`, which also sizes the DynamoDB client's idle connections per host; its fan-out has no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and at once after every reload. A reload never waits on the table: it holds the lock that serializes reloads, so it applies each tenant's switch against the last check's result and only wakes the retry. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. +- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59s` with the embedded queue, so that the lease plus its own ceiling to the next second plus one more second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`, which also sizes the DynamoDB client's idle connections per host; its fan-out has no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and at once after every reload. A reload makes no table call and does not wait on a tenant whose dedupe setting is unchanged: it holds the lock that serializes reloads, so it applies each tenant's switch against the last check's result and only wakes the retry; switching a tenant's dedupe off is the exception, waiting for that tenant's in-flight calls to finish before its store closes. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. - **A DynamoDB dedupe backend** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, retrying for up to eight jittered rounds both the items DynamoDB leaves unprocessed and a batch it throttled whole, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. `dedupe.backend: dynamodb` selects it (entry above). New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/config.yaml b/config.yaml index 3d24dc797..b45d5b9dc 100644 --- a/config.yaml +++ b/config.yaml @@ -57,7 +57,7 @@ mq: backend: embedded # NATS JetStream under /nats dedupe: backend: pebble # Pebble under /pebble; or dynamodb (below) - lease: 30s # how long a claimed id stays pending; at most 59.5s with the embedded mq (2*lease + 1s within its 2m duplicate window) + lease: 30s # how long a claimed id stays pending; at most 59s with the embedded mq (lease + ceil(lease) + 1s within its 2m duplicate window) reserve_concurrency: 64 # parallel calls per Reserve/Commit/Release to a remote backend, and DynamoDB's idle connections per host; the fan-out has no effect yet (ingest sends one id per call) # dynamodb: # read only when backend is dynamodb; credentials from the AWS SDK chain # table: wavehouse-dedupe-prod diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 83b54a813..6f18bb5a3 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -93,7 +93,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability (after which each `Config.Warnings` line is logged at `WARN`), ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, the lease coordinator, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. The boot config's `roles` decide which of them a process wires: every process gets the settings registry, observability, the MQ, the coordinator, the reload triggers and a listener; `api` adds schema discovery, the dedupe stores, streaming, auth and the full router; `ingest` adds the ingest worker; `sweeper` adds the sweeper; the ClickHouse pools and the cache come with `api` or `ingest`. A process without `api` serves `api.NewOpsRouter` (probes, `/version`, the metrics path, and the settings reload behind the operator key alone, `wireOpsAuth`) on `server.port`. `config.Validate` refuses a role set the backends cannot serve (a split over the embedded MQ, or `api` without `ingest` and the reverse over a local cache), and `New` refuses a `Config` with no roles, which only one built without `config.Load` can have. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe`, `wireCoord` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where each would be redelivered every ack wait for as long as the tenant is away and would hold that tenant's ack floor, so the sweeper could purge none of its queue past it. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultPolicy` reads it through the registry, where a flat directory's tenant `0` is always served. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own, and runs only while its process holds the `sweeper` lease (`elected`, which wraps `coord.RunElected` over the coordinator `wireCoord` opens; `local`, the only `coord.backend`, keeps leases in the process, so the one process always holds it). The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) and then its table — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The `dynamodb` case builds the same `Stores` over `Dynamo.Tenant`, gated (`Factory.Gated`) on the table's check: boot runs `Dynamo.Check` (after `CreateTable`, when `dedupe.dynamodb.create_table` is on) whether or not any tenant has dedupe on, and a table that fails it follows the same rule as a Pebble instance that cannot open, with the check retried until it passes by a background component that backs off from one second to thirty (a nested directory has no watcher). The `AfterAdopt` hook never runs the check, since it holds the lock that serializes reloads: it applies every store against the last check's result, so a tenant a reload switches on fails closed meanwhile, and wakes the retry, so a reload still retries at once. It has no Pebble gauges. `wireHTTP` hands the ingest handler `dedupe.lease` (`IngestHandler.DedupeLease`) whichever backend is chosen. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe`, `wireCoord` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where each would be redelivered every ack wait for as long as the tenant is away and would hold that tenant's ack floor, so the sweeper could purge none of its queue past it. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultPolicy` reads it through the registry, where a flat directory's tenant `0` is always served. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own, and runs only while its process holds the `sweeper` lease (`elected`, which wraps `coord.RunElected` over the coordinator `wireCoord` opens; `local`, the only `coord.backend`, keeps leases in the process, so the one process always holds it). The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) and then its table — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The `dynamodb` case builds the same `Stores` over `Dynamo.Tenant`, gated (`Factory.Gated`) on the table's check: boot runs `Dynamo.Check` (after `CreateTable`, when `dedupe.dynamodb.create_table` is on) whether or not any tenant has dedupe on, and a table that fails it follows the same rule as a Pebble instance that cannot open, with the check retried until it passes by a background component that backs off from one second to thirty (a nested directory has no watcher). The `AfterAdopt` hook never runs the check, since it holds the lock that serializes reloads, and it does not wait on a tenant whose `dedupe.enabled` is unchanged either — `Managed.Apply`'s no-op fast path settles that case under its own read lock, so the hook only takes a store's write lock, and so waits for that tenant's in-flight `Reserve`/`Commit`/`Release` calls to finish, on a genuine flip. It applies every store against the last check's result, so a tenant a reload switches on fails closed meanwhile, and wakes the retry, so a reload still retries at once. It has no Pebble gauges. `wireHTTP` hands the ingest handler `dedupe.lease` (`IngestHandler.DedupeLease`) whichever backend is chosen. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. ### `stream/` — SSE keepalive & fan-out @@ -120,7 +120,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **config.go** — Loads *boot* configuration from a YAML file with environment variable overrides (using [cleanenv](https://github.com/ilyakaznacheev/cleanenv)); every key has a `WH_`-prefixed env var. Boot config is only what can't change under a running process — the implementation each layer runs on, the process's `roles`, resource sizing, listeners, observability exporters, the settings-directory path, and the secrets (`clickhouse.password`, `auth.jwt_secret`, `auth.operator_key`). Everything tenant-tunable lives in the settings directory (`settings/`). Both sources are strict: `Load` refuses to boot naming every YAML key the struct doesn't declare (`strict.go`) and every `WH_*` environment variable no field binds (`check.go`), so a tunable that moved to the settings directory can't be read, ignored, and believed. Boot is the validator for this half — there is no dry-run command. See [Configuration Reference](/configuration). - **check.go** — `rejectUnboundEnv` is the environment half of the strict loader: `unboundEnv` walks the struct's `env` tags (plus the two process-level names, `WH_CONFIG` and `WH_LOG_LEVEL`) against the environment; `CheckDataDir` probes `data_dir` — run by `main` right after `Load` when `Config.NeedsDataDir` says a selected backend keeps state there, so an unusable `data_dir` refuses boot before anything dials out. It refuses an empty value (reachable through `WH_DATA_DIR=`) outright rather than probing the working directory; a path that exists and is not a directory; a dangling symlink at `data_dir` or any component above it (the walk to the nearest existing ancestor uses `Lstat`, so a failed mount is not skipped over as "does not exist"); and a directory the process cannot write to — or, when it does not exist, an unwritable nearest ancestor — probed by creating and removing one temp file. A permission denial, on the probe or on reaching the path through a parent without search permission, carries the UID-65532 hint, since a bind mount owned by root is the typical cause. -- **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` in `Validate`, between `validateRoles` and `validateTopology`), which refuses a value not on the list and names the ones that are. One rule spans two layers: while `mq.backend` is `embedded`, twice `dedupe.lease` plus one second must fit the embedded MQ's 2m duplicate window (`embeddedDuplicateWindow`), a cap of 59.5s (`maxEmbeddedLease`), because a client obeying the in-flight `503`'s `Retry-After` republishes up to twice the lease after the claim and DynamoDB rounds a claim's expiry up to the second. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns the valid combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), which `app.New` logs at `WARN`. +- **backends.go** — the `.backend` keys: one string type per layer (`MQBackend`, `CacheBackend`, `DedupeBackend`, `CoordBackend`), each with its list of the backends this build has, and a `validate` per layer block (`checkBackend`, run by `validateBackends` in `Validate`, between `validateRoles` and `validateTopology`), which refuses a value not on the list and names the ones that are. One rule spans two layers: while `mq.backend` is `embedded`, `dedupe.lease` plus its own ceiling to the next whole second (`ceilSecond`) plus one more second must fit the embedded MQ's 2m duplicate window (`embeddedDuplicateWindow`), a cap of 59s (`maxEmbeddedLease`), because a client obeying the in-flight `503`'s `Retry-After` can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second. A backend's own settings go in a `.` sub-block that is its `validate`'s case to check. `Distributed` reports whether the MQ is shared with other processes, `NeedsDataDir` whether a selected backend keeps state under `data_dir`, and `Warnings` returns the valid combinations that are correct for one replica only (a shared MQ over a local cache or Pebble dedupe), which `app.New` logs at `WARN`. - **config.go**, roles — `roles` (`[]Role`: `api`, `ingest`, `sweeper`; `AllRoles` by default; `Has(Role)`) picks which components `internal/app` wires, and `instance_id` names the process (`-<8 hex>` when empty, resolved in `Load`; today only logged at boot, and a distributed coordinator will record it as a lease's holder). `validateRoles` refuses an empty list, an empty entry, an unknown or a repeated role; `validateTopology` refuses a role set the backends cannot serve: any split over the embedded MQ, and a process with exactly one of `api` and `ingest` over a local cache. `NeedsDataDir` counts Pebble only for a process running `api`, and `Warnings` is empty without `api`, since only that role opens a cache it reads or a dedupe store. - **strict.go** — `rejectUnknownKeys`, the YAML half: re-reads the file as a generic tree and walks it against the struct's `yaml` tags, listing every key the struct doesn't declare. cleanenv itself is lenient by design, which is exactly wrong for boot config once keys have moved to the settings directory. - **persistence.go** — `WarnIfFreshDataDir` logs the startup `WARN` when `data_dir` is missing or empty (on a redeploy, the sign that the volume didn't persist); `LogStorageInitError` attaches the UID-65532 `permissionHint` to a NATS or Pebble open failure that looks like a permission denial — the same hint string `CheckDataDir` uses. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index d89d18331..a1286b53c 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -56,19 +56,19 @@ Whether a tenant dedupes, and on which field, are settings-directory keys ([Dedu | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | -| `dedupe.lease` | `WH_DEDUPE_LEASE` | `30s` | How long a record's id stays claimed while the record is published. Another request carrying the same id meanwhile gets `503` with this as `Retry-After`, in whole seconds; a claim that is neither committed nor released, because its process died mid-publish, lapses after it. With `mq.backend: embedded`, twice the lease plus one second must fit the embedded queue's 2-minute duplicate window, so the lease is at most `59.5s`: a client that obeys `Retry-After` after a publish whose outcome it never learned republishes up to twice the lease after the claim, and DynamoDB rounds a claim's expiry up to the second. A longer lease refuses boot. A Go duration (`30s`, `45s`); `0` refuses boot. | +| `dedupe.lease` | `WH_DEDUPE_LEASE` | `30s` | How long a record's id stays claimed while the record is published. Another request carrying the same id meanwhile gets `503` with this as `Retry-After`, in whole seconds; a claim that is neither committed nor released, because its process died mid-publish, lapses after it. With `mq.backend: embedded`, the lease plus its own ceiling to the next whole second plus one more second must fit the embedded queue's 2-minute duplicate window, so the lease is at most `59s`: a client that obeys `Retry-After` after a publish whose outcome it never learned can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second. A longer lease refuses boot. A Go duration (`30s`, `45s`); `0` refuses boot. | | `dedupe.reserve_concurrency` | `WH_DEDUPE_RESERVE_CONCURRENCY` | `64` | The most parallel calls one Reserve, Commit or Release makes to a remote dedupe backend, and the idle connections per host the DynamoDB client keeps to match. Ingest sends one id per call today, so the fan-out has no effect yet; `pebble` ignores it. `0` refuses boot. | #### DynamoDB dedupe -Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (String) alone, and TTL off on `ex` is logged as a warning. A table that fails the check, or cannot be reached, refuses boot with a flat settings directory. With a nested one, the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and the check is retried in the background, backing off from one second to thirty, and at once after every reload, so a table that comes good is picked up without a restart. A reload never waits on the table: it applies each tenant's switch against the last check's result, so a tenant it switches on fails closed until the retry passes. The check runs in every process running the `api` [role](#process-roles), the one that opens the dedupe stores, whether or not any tenant has dedupe on. +Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (String) alone, and TTL off on `ex` is logged as a warning. A table that fails the check, or cannot be reached, refuses boot with a flat settings directory. With a nested one, the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and the check is retried in the background, backing off from one second to thirty, and at once after every reload, so a table that comes good is picked up without a restart. A reload makes no table call, and does not wait on a tenant whose dedupe setting is unchanged: it applies each tenant's switch against the last check's result, so a tenant it switches on fails closed until the retry passes. Switching a tenant's dedupe off is the exception — it waits for that tenant's in-flight `Reserve`/`Commit`/`Release` calls to finish before the store closes. The check runs in every process running the `api` [role](#process-roles), the one that opens the dedupe stores, whether or not any tenant has dedupe on. | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | | `dedupe.dynamodb.table` | `WH_DEDUPE_DYNAMODB_TABLE` | *(required)* | The shared table. | | `dedupe.dynamodb.region` | `WH_DEDUPE_DYNAMODB_REGION` | *(empty)* | The table's region. Empty uses the SDK chain's (`AWS_REGION`); no region from either refuses boot. | | `dedupe.dynamodb.endpoint` | `WH_DEDUPE_DYNAMODB_ENDPOINT` | *(empty)* | A custom endpoint, for [dynamodb-local](https://docs.aws.amazon.com/amazondynamodb/latest/developerguide/DynamoDBLocal.html) in development and tests. Leave it empty against AWS. | -| `dedupe.dynamodb.timeout` | `WH_DEDUPE_DYNAMODB_TIMEOUT` | `250ms` | Deadline for each DynamoDB call, the SDK's retries included. The retries back off with full jitter, each wait capped at `timeout / (2 × (max_attempts − 1))`, so together they wait at most half of it and a throttled call fails on its last attempt's answer rather than on the deadline. `0` refuses boot. | +| `dedupe.dynamodb.timeout` | `WH_DEDUPE_DYNAMODB_TIMEOUT` | `250ms` | Deadline for each DynamoDB call, the SDK's retries included. The retries back off with full jitter, each wait capped at `timeout / (2 × (max_attempts − 1))`, so together they wait at most half of it and a throttled call fails on its last attempt's answer rather than on the deadline. The boot and background table check (verifying the key schema and TTL) is not one of these calls: it runs under its own deadline of 10 × `timeout` (`2.5s` by default). `0` refuses boot. | | `dedupe.dynamodb.max_attempts` | `WH_DEDUPE_DYNAMODB_MAX_ATTEMPTS` | `3` | Attempts per call, the first included. More attempts share the same half of `timeout` for their waits, so each retry waits less rather than the call running longer. `0` refuses boot. | | `dedupe.dynamodb.retry_mode` | `WH_DEDUPE_DYNAMODB_RETRY_MODE` | `standard` | `standard`, or `adaptive`, which also slows the client down after throttling. Anything else, empty included, refuses boot. | | `dedupe.dynamodb.create_table` | `WH_DEDUPE_DYNAMODB_CREATE_TABLE` | `false` | Development only: create the table at boot if it is missing, with TTL on `ex`. Refused unless `endpoint` is set, so it never creates a table in AWS; the production table belongs to your infrastructure code. | @@ -265,7 +265,7 @@ cache: dedupe: backend: pebble # in-process Pebble under /pebble; or dynamodb - lease: 30s # at most 59.5s with the embedded mq + lease: 30s # at most 59s with the embedded mq reserve_concurrency: 64 # dynamodb: # read only when backend is dynamodb # table: wavehouse-dedupe-prod diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index de5fb64bf..e9aca3710 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -512,7 +512,7 @@ dedupe: region: us-east-1 # or leave empty for AWS_REGION ``` -or `WH_DEDUPE_BACKEND=dynamodb`, `WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod`. A table that is missing, has the wrong key schema, or cannot be reached with the pod's credentials refuses boot over a flat settings directory; over a nested one the pod boots, every tenant with dedupe on fails its ingest closed, and the check is retried in the background (backing off from one second to thirty, and at once after every reload, which never waits on the table). No region at all (neither `region` nor one from the SDK chain: `AWS_REGION`, `AWS_DEFAULT_REGION` or a profile) refuses boot in both shapes. The check runs in every pod running the `api` [role](/configuration#process-roles), whether or not any tenant has `dedupe.enabled` on; a pod without it opens no dedupe store. The per-tenant switch stays in each tenant's `config.json`. +or `WH_DEDUPE_BACKEND=dynamodb`, `WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod`. A table that is missing, has the wrong key schema, or cannot be reached with the pod's credentials refuses boot over a flat settings directory; over a nested one the pod boots, every tenant with dedupe on fails its ingest closed, and the check is retried in the background (backing off from one second to thirty, and at once after every reload). A reload makes no table call itself, and does not wait on a tenant whose dedupe setting is unchanged; switching a tenant's dedupe off waits for that tenant's in-flight calls to finish before its store closes. No region at all (neither `region` nor one from the SDK chain: `AWS_REGION`, `AWS_DEFAULT_REGION` or a profile) refuses boot in both shapes. The check runs in every pod running the `api` [role](/configuration#process-roles), whether or not any tenant has `dedupe.enabled` on; a pod without it opens no dedupe store. The per-tenant switch stays in each tenant's `config.json`. For development against dynamodb-local, set `dedupe.dynamodb.endpoint` (for example `http://localhost:8000`) and `create_table: true`, and give the SDK any static credentials (`AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`) and a region. `create_table` without an `endpoint` refuses boot. diff --git a/docs/src/content/docs/development.md b/docs/src/content/docs/development.md index ab97cb795..4ca7a43d9 100644 --- a/docs/src/content/docs/development.md +++ b/docs/src/content/docs/development.md @@ -459,7 +459,7 @@ WaveHouse/ │ ├── chsql/ # Shared ClickHouse SQL helpers (quoting + bind-safety) │ ├── config/ # YAML + env var configuration │ ├── coord/ # Leases with fencing tokens (in-process Local, RunElected, coordtest suite) -│ ├── dedupe/ # Optional deduplication (Reserve/Commit/Release; Pebble, DynamoDB) +│ ├── dedupe/ # Optional deduplication (Reserve/Commit/Release; Pebble or DynamoDB) │ ├── discovery/ # ClickHouse schema introspection + validation │ ├── ingest/ # Batch buffering + DLQ + Active Sweeper │ ├── keyenc/ # One escaping for composite keys (NATS subject tokens, cache namespace tokens, dedupe keys) From d0a61e4a93554edc0a8e625f22103f90c1ec173a Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:27:45 -0400 Subject: [PATCH 29/38] test(app): a reload must not wait behind a tenant's in-flight DynamoDB call MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Adds a fake per-op hang (fakeDynamo.setHangOn) alongside the existing whole-endpoint one, then exercises Managed.Apply's no-op fast path end to end: a tenant with dedupe on has a Commit in flight against a BatchWriteItem that never answers, and asserts that a reload naming no change for that tenant still returns in well under a second, and that a concurrent Reserve for the same tenant is not blocked either — which it would be if the reload's Apply took the write lock unconditionally, since a pending writer blocks new readers too. Verified by hand: temporarily skipping the settled check in internal/dedupe/managed.go's Apply made the reload assertion fail at 2.475s (Commit's own retry/backoff giving up on the hung call, not a deadlock) instead of passing in 0.02–0.03s; reverted, `git diff` on that file confirmed clean before this commit. Co-Authored-By: Claude Opus 5.5 (1M context) --- internal/app/dedupe_dynamodb_test.go | 64 ++++++++++++++++++++++++++-- 1 file changed, 61 insertions(+), 3 deletions(-) diff --git a/internal/app/dedupe_dynamodb_test.go b/internal/app/dedupe_dynamodb_test.go index f3fb1e598..71bf294c6 100644 --- a/internal/app/dedupe_dynamodb_test.go +++ b/internal/app/dedupe_dynamodb_test.go @@ -23,11 +23,13 @@ import ( // fakeDynamo answers the DynamoDB JSON protocol for one table, enough for // boot's check, the dev create path, and a claim and its commit. Whether the -// table exists, and whether the endpoint hangs, are the test's to switch. +// table exists, and whether the endpoint hangs (every call, or one op +// alone), are the test's to switch. type fakeDynamo struct { mu sync.Mutex exists bool hangs bool + hangOn string // hang calls of this op alone, once set; "" hangs none this way calls []string } @@ -43,6 +45,12 @@ func (f *fakeDynamo) setHangs(v bool) { f.hangs = v } +func (f *fakeDynamo) setHangOn(op string) { + f.mu.Lock() + defer f.mu.Unlock() + f.hangOn = op +} + func (f *fakeDynamo) called(op string) bool { return f.count(op) > 0 } func (f *fakeDynamo) count(op string) int { @@ -58,6 +66,9 @@ func (f *fakeDynamo) count(op string) int { } func (f *fakeDynamo) ServeHTTP(w http.ResponseWriter, r *http.Request) { + // Drained before any hang below: with the body unread, an SDK write + // deadline or the client giving up never reaches this handler, since the + // connection looks like it's still waiting for us to consume it. _, _ = io.Copy(io.Discard, r.Body) _, op, _ := strings.Cut(r.Header.Get("X-Amz-Target"), ".") f.mu.Lock() @@ -65,9 +76,9 @@ func (f *fakeDynamo) ServeHTTP(w http.ResponseWriter, r *http.Request) { if op == "CreateTable" { f.exists = true } - exists, hangs := f.exists, f.hangs + exists, hang := f.exists, f.hangs || op == f.hangOn f.mu.Unlock() - if hangs { + if hang { <-r.Context().Done() return } @@ -259,3 +270,50 @@ func TestNew_DynamoDBDedupeRefusesNoRegion(t *testing.T) { }) } } + +// A reload must not wait behind a tenant's own in-flight DynamoDB call when +// nothing changes for that tenant: Managed.Apply's no-op fast path settles +// under a read lock, so it never contends with a Commit already holding one +// — and, since Go's RWMutex blocks new readers behind a pending writer, a +// concurrent Reserve for the same tenant must also go through, which it +// would not if the reload's Apply took the write lock unconditionally. +func TestReload_DynamoDBDedupeDoesNotWaitOnInFlightCommit(t *testing.T) { + cfg := testConfig(t, writeSettings(t, dedupeOn)) + fake := dynamoConfig(t, cfg, true) + fake.setHangOn("BatchWriteItem") + a := newApp(t, cfg, Options{}) + + store := a.dedup.For(tenant.Default) + require.True(t, store.Open()) + claims, err := store.Reserve(context.Background(), []dedupe.Key{eventKey}, time.Minute) + require.NoError(t, err) + require.Equal(t, dedupe.Claimed, claims[0].Status) + + commitCtx, cancelCommit := context.WithCancel(context.Background()) + defer cancelCommit() + commitDone := make(chan error, 1) + go func() { commitDone <- store.Commit(commitCtx, claims, 0) }() + require.Eventually(t, func() bool { return fake.called("BatchWriteItem") }, time.Second, time.Millisecond, + "commit reached the table and is now hanging on it") + + start := time.Now() + a.tenants.Reload("test") + assert.Less(t, time.Since(start), 500*time.Millisecond, + "a reload that changes nothing for this tenant waited on its in-flight commit") + + otherKey := dedupe.Key{Table: eventKey.Table, ID: "concurrent-reserve"} + reserveDone := make(chan error, 1) + go func() { + _, err := store.Reserve(context.Background(), []dedupe.Key{otherKey}, time.Minute) + reserveDone <- err + }() + select { + case err := <-reserveDone: + require.NoError(t, err, "a Reserve for the same tenant, started right after the reload") + case <-time.After(500 * time.Millisecond): + t.Fatal("a concurrent Reserve for the same tenant was blocked") + } + + cancelCommit() + <-commitDone // let the hung call finish (canceled) before the app closes +} From b16c193ed395c1f02b6b2b5c82e26f5a29e36d42 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:38:22 -0400 Subject: [PATCH 30/38] test(app): drop the concurrent-Reserve half of the reload/commit test MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The Reserve only started after Reload had already returned, so no writer was ever pending when it ran, and it could not fail for the reason its comment gave — with the fast path removed only the reload-time assertion actually failed. No clean sync point exists into "Reload is inside the dedupe hook" without instrumenting production code for the test alone, so drop the block and the sentence claiming it; the reload-time assertion already pins the regression on its own. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- internal/app/dedupe_dynamodb_test.go | 19 ++----------------- 1 file changed, 2 insertions(+), 17 deletions(-) diff --git a/internal/app/dedupe_dynamodb_test.go b/internal/app/dedupe_dynamodb_test.go index 71bf294c6..310d11408 100644 --- a/internal/app/dedupe_dynamodb_test.go +++ b/internal/app/dedupe_dynamodb_test.go @@ -273,10 +273,8 @@ func TestNew_DynamoDBDedupeRefusesNoRegion(t *testing.T) { // A reload must not wait behind a tenant's own in-flight DynamoDB call when // nothing changes for that tenant: Managed.Apply's no-op fast path settles -// under a read lock, so it never contends with a Commit already holding one -// — and, since Go's RWMutex blocks new readers behind a pending writer, a -// concurrent Reserve for the same tenant must also go through, which it -// would not if the reload's Apply took the write lock unconditionally. +// under a read lock alone, so it never contends with a Commit already +// holding one and returns long before the commit does. func TestReload_DynamoDBDedupeDoesNotWaitOnInFlightCommit(t *testing.T) { cfg := testConfig(t, writeSettings(t, dedupeOn)) fake := dynamoConfig(t, cfg, true) @@ -301,19 +299,6 @@ func TestReload_DynamoDBDedupeDoesNotWaitOnInFlightCommit(t *testing.T) { assert.Less(t, time.Since(start), 500*time.Millisecond, "a reload that changes nothing for this tenant waited on its in-flight commit") - otherKey := dedupe.Key{Table: eventKey.Table, ID: "concurrent-reserve"} - reserveDone := make(chan error, 1) - go func() { - _, err := store.Reserve(context.Background(), []dedupe.Key{otherKey}, time.Minute) - reserveDone <- err - }() - select { - case err := <-reserveDone: - require.NoError(t, err, "a Reserve for the same tenant, started right after the reload") - case <-time.After(500 * time.Millisecond): - t.Fatal("a concurrent Reserve for the same tenant was blocked") - } - cancelCommit() <-commitDone // let the hung call finish (canceled) before the app closes } From 674746a78ab6de0049ea23558361af61c2178acd Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 08:38:42 -0400 Subject: [PATCH 31/38] docs(dedupe): the idle-pool floor, and what a reload's wait covers configuration.mdx's dedupe.reserve_concurrency row now says the DynamoDB idle-connection pool never goes below the SDK's own default (10), matching newHTTPClient's max(...) floor. configuration.mdx, deployment.md and CHANGELOG.md said switching a tenant's dedupe off was the one case a reload waits on. It is not the only one: a reload that removes or rejects a tenant whose store was open closes it the same way (Stores.Retain -> Managed.Close -> Apply(false)), taking the same write lock and waiting on the same in-flight calls. Reworded all three to match architecture.md's existing "on a genuine flip" framing, which already covered both cases. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- CHANGELOG.md | 2 +- docs/src/content/docs/configuration.mdx | 4 ++-- docs/src/content/docs/deployment.md | 2 +- 3 files changed, 4 insertions(+), 4 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 747fbd8fb..4ab522f67 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59s` with the embedded queue, so that the lease plus its own ceiling to the next second plus one more second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`, which also sizes the DynamoDB client's idle connections per host; its fan-out has no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and at once after every reload. A reload makes no table call and does not wait on a tenant whose dedupe setting is unchanged: it holds the lock that serializes reloads, so it applies each tenant's switch against the last check's result and only wakes the retry; switching a tenant's dedupe off is the exception, waiting for that tenant's in-flight calls to finish before its store closes. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. +- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59s` with the embedded queue, so that the lease plus its own ceiling to the next second plus one more second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`, which also sizes the DynamoDB client's idle connections per host; its fan-out has no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and at once after every reload. A reload makes no table call and does not wait on a tenant whose dedupe setting is unchanged: it holds the lock that serializes reloads, so it applies each tenant's switch against the last check's result and only wakes the retry; it waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight calls, before its store closes. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. - **A DynamoDB dedupe backend** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, retrying for up to eight jittered rounds both the items DynamoDB leaves unprocessed and a batch it throttled whole, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. `dedupe.backend: dynamodb` selects it (entry above). New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index a1286b53c..bd1781d00 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -57,11 +57,11 @@ Whether a tenant dedupes, and on which field, are settings-directory keys ([Dedu | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | | `dedupe.lease` | `WH_DEDUPE_LEASE` | `30s` | How long a record's id stays claimed while the record is published. Another request carrying the same id meanwhile gets `503` with this as `Retry-After`, in whole seconds; a claim that is neither committed nor released, because its process died mid-publish, lapses after it. With `mq.backend: embedded`, the lease plus its own ceiling to the next whole second plus one more second must fit the embedded queue's 2-minute duplicate window, so the lease is at most `59s`: a client that obeys `Retry-After` after a publish whose outcome it never learned can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second. A longer lease refuses boot. A Go duration (`30s`, `45s`); `0` refuses boot. | -| `dedupe.reserve_concurrency` | `WH_DEDUPE_RESERVE_CONCURRENCY` | `64` | The most parallel calls one Reserve, Commit or Release makes to a remote dedupe backend, and the idle connections per host the DynamoDB client keeps to match. Ingest sends one id per call today, so the fan-out has no effect yet; `pebble` ignores it. `0` refuses boot. | +| `dedupe.reserve_concurrency` | `WH_DEDUPE_RESERVE_CONCURRENCY` | `64` | The most parallel calls one Reserve, Commit or Release makes to a remote dedupe backend, and the idle connections per host the DynamoDB client keeps to match, never fewer than the SDK's own default (10). Ingest sends one id per call today, so the fan-out has no effect yet; `pebble` ignores it. `0` refuses boot. | #### DynamoDB dedupe -Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (String) alone, and TTL off on `ex` is logged as a warning. A table that fails the check, or cannot be reached, refuses boot with a flat settings directory. With a nested one, the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and the check is retried in the background, backing off from one second to thirty, and at once after every reload, so a table that comes good is picked up without a restart. A reload makes no table call, and does not wait on a tenant whose dedupe setting is unchanged: it applies each tenant's switch against the last check's result, so a tenant it switches on fails closed until the retry passes. Switching a tenant's dedupe off is the exception — it waits for that tenant's in-flight `Reserve`/`Commit`/`Release` calls to finish before the store closes. The check runs in every process running the `api` [role](#process-roles), the one that opens the dedupe stores, whether or not any tenant has dedupe on. +Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (String) alone, and TTL off on `ex` is logged as a warning. A table that fails the check, or cannot be reached, refuses boot with a flat settings directory. With a nested one, the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and the check is retried in the background, backing off from one second to thirty, and at once after every reload, so a table that comes good is picked up without a restart. A reload makes no table call, and does not wait on a tenant whose dedupe setting is unchanged: it applies each tenant's switch against the last check's result, so a tenant it switches on fails closed until the retry passes. A reload waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight `Reserve`/`Commit`/`Release` calls, before the store itself closes. The check runs in every process running the `api` [role](#process-roles), the one that opens the dedupe stores, whether or not any tenant has dedupe on. | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index e9aca3710..ddf377953 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -512,7 +512,7 @@ dedupe: region: us-east-1 # or leave empty for AWS_REGION ``` -or `WH_DEDUPE_BACKEND=dynamodb`, `WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod`. A table that is missing, has the wrong key schema, or cannot be reached with the pod's credentials refuses boot over a flat settings directory; over a nested one the pod boots, every tenant with dedupe on fails its ingest closed, and the check is retried in the background (backing off from one second to thirty, and at once after every reload). A reload makes no table call itself, and does not wait on a tenant whose dedupe setting is unchanged; switching a tenant's dedupe off waits for that tenant's in-flight calls to finish before its store closes. No region at all (neither `region` nor one from the SDK chain: `AWS_REGION`, `AWS_DEFAULT_REGION` or a profile) refuses boot in both shapes. The check runs in every pod running the `api` [role](/configuration#process-roles), whether or not any tenant has `dedupe.enabled` on; a pod without it opens no dedupe store. The per-tenant switch stays in each tenant's `config.json`. +or `WH_DEDUPE_BACKEND=dynamodb`, `WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod`. A table that is missing, has the wrong key schema, or cannot be reached with the pod's credentials refuses boot over a flat settings directory; over a nested one the pod boots, every tenant with dedupe on fails its ingest closed, and the check is retried in the background (backing off from one second to thirty, and at once after every reload). A reload makes no table call itself, and does not wait on a tenant whose dedupe setting is unchanged; it waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight calls, before its store closes. No region at all (neither `region` nor one from the SDK chain: `AWS_REGION`, `AWS_DEFAULT_REGION` or a profile) refuses boot in both shapes. The check runs in every pod running the `api` [role](/configuration#process-roles), whether or not any tenant has `dedupe.enabled` on; a pod without it opens no dedupe store. The per-tenant switch stays in each tenant's `config.json`. For development against dynamodb-local, set `dedupe.dynamodb.endpoint` (for example `http://localhost:8000`) and `create_table: true`, and give the SDK any static credentials (`AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`) and a region. `create_table` without an `endpoint` refuses boot. From 7f38b9b350924f4eede8e3f45951edb72c574537 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 09:00:41 -0400 Subject: [PATCH 32/38] fix(app): move DynamoDB dedupe wiring out of wire.go for the e2e gate make ci's e2e coverage measured 59.7% against a 60% floor: the e2e binary always boots with dedupe.backend: pebble, so wireDynamoDedupe, errDynamoUnchecked and the table-check retry component never ran there, same shape as #628's internal/dedupe/dynamodb.go exclusion. Pure move, no behavior change: wireDynamoDedupe and errDynamoUnchecked move verbatim into the new internal/app/wire_dynamodb.go; wireDedupe's switch stays in wire.go untouched. Adds a matching e2e-only exclusion for the new file in .testcoverage.yml, next to dynamodb.go's, so unit and integration keep covering it and the merged total still counts it -- wire.go itself stays out of the exclude list. Updates architecture.md's wire.go bullet (the dynamodb case now points at the new file's own bullet) and the CHANGELOG's file-provenance list for the dedupe.backend: dynamodb entry. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01G6Cz4H5k1spJAPk5CZeGds --- .testcoverage.yml | 6 ++ CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 4 +- internal/app/wire.go | 115 +----------------------- internal/app/wire_dynamodb.go | 125 ++++++++++++++++++++++++++ 5 files changed, 139 insertions(+), 113 deletions(-) create mode 100644 internal/app/wire_dynamodb.go diff --git a/.testcoverage.yml b/.testcoverage.yml index 77b365e9d..afb04f49c 100644 --- a/.testcoverage.yml +++ b/.testcoverage.yml @@ -84,3 +84,9 @@ exclude: # (fake API) and integration (dynamodb-local) suites cover it, and the # merged total still counts it. - ^internal/dedupe/dynamodb\.go$ + # wireDynamoDedupe and its retry component (internal/app/wire_dynamodb.go): + # same reason as dynamodb.go above — the e2e binary never selects + # dedupe.backend: dynamodb, so this file measured 0% there and pulled + # e2e to 59.7%. The unit and integration suites cover it, and the + # merged total still counts it. + - ^internal/app/wire_dynamodb\.go$ diff --git a/CHANGELOG.md b/CHANGELOG.md index 574c8f867..5d405dc6e 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59s` with the embedded queue, so that the lease plus its own ceiling to the next second plus one more second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`, which also sizes the DynamoDB client's idle connections per host; its fan-out has no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and at once after every reload. A reload makes no table call and does not wait on a tenant whose dedupe setting is unchanged: it holds the lock that serializes reloads, so it applies each tenant's switch against the last check's result and only wakes the retry; it waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight calls, before its store closes. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. +- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire,wire_dynamodb}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `.testcoverage.yml`, `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59s` with the embedded queue, so that the lease plus its own ceiling to the next second plus one more second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`, which also sizes the DynamoDB client's idle connections per host; its fan-out has no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and at once after every reload. A reload makes no table call and does not wait on a tenant whose dedupe setting is unchanged: it holds the lock that serializes reloads, so it applies each tenant's switch against the last check's result and only wakes the retry; it waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight calls, before its store closes. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. - **A DynamoDB dedupe backend** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, retrying for up to eight jittered rounds both the items DynamoDB leaves unprocessed and a batch it throttled whole, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. `dedupe.backend: dynamodb` selects it (entry above). New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 6f18bb5a3..579f5a216 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -93,7 +93,9 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request ### `app/` — Process wiring - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability (after which each `Config.Warnings` line is logged at `WARN`), ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, the lease coordinator, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. The boot config's `roles` decide which of them a process wires: every process gets the settings registry, observability, the MQ, the coordinator, the reload triggers and a listener; `api` adds schema discovery, the dedupe stores, streaming, auth and the full router; `ingest` adds the ingest worker; `sweeper` adds the sweeper; the ClickHouse pools and the cache come with `api` or `ingest`. A process without `api` serves `api.NewOpsRouter` (probes, `/version`, the metrics path, and the settings reload behind the operator key alone, `wireOpsAuth`) on `server.port`. `config.Validate` refuses a role set the backends cannot serve (a split over the embedded MQ, or `api` without `ingest` and the reverse over a local cache), and `New` refuses a `Config` with no roles, which only one built without `config.Load` can have. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. -- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe`, `wireCoord` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where each would be redelivered every ack wait for as long as the tenant is away and would hold that tenant's ack floor, so the sweeper could purge none of its queue past it. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultPolicy` reads it through the registry, where a flat directory's tenant `0` is always served. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own, and runs only while its process holds the `sweeper` lease (`elected`, which wraps `coord.RunElected` over the coordinator `wireCoord` opens; `local`, the only `coord.backend`, keeps leases in the process, so the one process always holds it). The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) and then its table — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. The `dynamodb` case builds the same `Stores` over `Dynamo.Tenant`, gated (`Factory.Gated`) on the table's check: boot runs `Dynamo.Check` (after `CreateTable`, when `dedupe.dynamodb.create_table` is on) whether or not any tenant has dedupe on, and a table that fails it follows the same rule as a Pebble instance that cannot open, with the check retried until it passes by a background component that backs off from one second to thirty (a nested directory has no watcher). The `AfterAdopt` hook never runs the check, since it holds the lock that serializes reloads, and it does not wait on a tenant whose `dedupe.enabled` is unchanged either — `Managed.Apply`'s no-op fast path settles that case under its own read lock, so the hook only takes a store's write lock, and so waits for that tenant's in-flight `Reserve`/`Commit`/`Release` calls to finish, on a genuine flip. It applies every store against the last check's result, so a tenant a reload switches on fails closed meanwhile, and wakes the retry, so a reload still retries at once. It has no Pebble gauges. `wireHTTP` hands the ingest handler `dedupe.lease` (`IngestHandler.DedupeLease`) whichever backend is chosen. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. +- **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe`, `wireCoord` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where each would be redelivered every ack wait for as long as the tenant is away and would hold that tenant's ack floor, so the sweeper could purge none of its queue past it. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultPolicy` reads it through the registry, where a flat directory's tenant `0` is always served. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own, and runs only while its process holds the `sweeper` lease (`elected`, which wraps `coord.RunElected` over the coordinator `wireCoord` opens; `local`, the only `coord.backend`, keeps leases in the process, so the one process always holds it). The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) and then its table — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. `wireDedupe`'s `dynamodb` case is `wireDynamoDedupe`, in wire_dynamodb.go (below). `wireHTTP` hands the ingest handler `dedupe.lease` (`IngestHandler.DedupeLease`) whichever backend is chosen. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. + +- **wire_dynamodb.go** — `wireDedupe`'s `dynamodb` case, split out of wire.go so the e2e suite's coverage exclude for it (the e2e binary always runs Pebble dedupe, never DynamoDB) doesn't have to blanket wire.go itself: builds the same `dedupe.Stores` over `Dynamo.Tenant`, gated (`Factory.Gated`) on the table's check: boot runs `Dynamo.Check` (after `CreateTable`, when `dedupe.dynamodb.create_table` is on) whether or not any tenant has dedupe on, and a table that fails it follows the same rule as a Pebble instance that cannot open, with the check retried until it passes by a background component that backs off from one second to thirty (a nested directory has no watcher). The `AfterAdopt` hook never runs the check, since it holds the lock that serializes reloads, and it does not wait on a tenant whose `dedupe.enabled` is unchanged either — `Managed.Apply`'s no-op fast path settles that case under its own read lock, so the hook only takes a store's write lock, and so waits for that tenant's in-flight `Reserve`/`Commit`/`Release` calls to finish, on a genuine flip. It applies every store against the last check's result, so a tenant a reload switches on fails closed meanwhile, and wakes the retry, so a reload still retries at once. It has no Pebble gauges. ### `stream/` — SSE keepalive & fan-out diff --git a/internal/app/wire.go b/internal/app/wire.go index a65d1bf8d..dd48fa070 100644 --- a/internal/app/wire.go +++ b/internal/app/wire.go @@ -532,117 +532,10 @@ func (a *App) wirePebbleDedupe() error { return nil } -// errDynamoUnchecked is a store's open before the first table check has run. -var errDynamoUnchecked = errors.New("dedupe: dynamodb table not checked yet") - -// wireDynamoDedupe builds the dedupe stores over one DynamoDB table that -// every tenant and every process shares (dedupe.Dynamo), so a tenant's store -// opens for free once the table has passed its check. Boot checks it (after -// creating it, with create_table on dynamodb-local) whether or not any tenant -// has dedupe on, and never creates it otherwise. A table that fails the check -// follows the registry's rule for the shape, as Pebble's instance does: a -// flat directory refuses boot; a nested one boots with every switched-on -// store closed, so its ingest fails closed. Unlike a local disk, a remote -// table's failure is usually brief (a throttle, credentials not yet issued -// mid-rollout), and a nested directory has no watcher to reload it, so the -// check is then retried in the background, with backoff, until it passes. -// The check is network I/O, so the AfterAdopt hook never runs it: the hook -// holds the lock that serializes reloads. It applies every store against the -// last check's result and wakes the retry, so a reload still retries at once. -func (a *App) wireDynamoDedupe(ctx context.Context) error { - c := a.cfg.Dedupe.DynamoDB - d, err := dedupe.NewDynamo(ctx, dedupe.DynamoConfig{ - Table: c.Table, Region: c.Region, Endpoint: c.Endpoint, - Timeout: c.Timeout, MaxAttempts: c.MaxAttempts, RetryMode: c.RetryMode, - ReserveConcurrency: a.cfg.Dedupe.ReserveConcurrency, - }) - if err != nil { - return err - } - var mu sync.Mutex - state := errDynamoUnchecked // nil once the table has passed, for good - ready := func() error { - mu.Lock() - defer mu.Unlock() - return state - } - // check is only ever run by boot, then by the retry loop, one at a time. - check := func(ctx context.Context) error { - var err error - if c.CreateTable { - err = d.CreateTable(ctx) - } - if err == nil { - err = d.Check(ctx) - } - mu.Lock() - defer mu.Unlock() - if state != nil { - state = err - } - return state - } - stores := dedupe.NewStores(dedupe.Factory(d.Tenant).Gated(ready)) - a.dedup = stores - a.add(component{name: "dedupe", close: withoutContext(stores.Close)}) - var reconciling sync.Mutex // the hook and the retry loop both apply - apply := func() { - reconciling.Lock() - defer reconciling.Unlock() - if err := stores.Retain(a.served); err != nil { - slog.Error("dedupe store close failed", "error", err) - } - for id, store := range a.tenants.All() { - m := stores.For(id) - enabled := store.DedupeEnabled() - wasOpen := m.Open() - // The one failure an open has is the check's, logged where it ran. - _ = m.Apply(enabled) - if m.Open() != wasOpen { - slog.Info("dedupe store reconciled with settings", "tenant", id, "enabled", enabled) - } - } - } - retry := make(chan struct{}, 1) - a.tenants.AfterAdopt(func([]tenant.ID) { - apply() - if ready() != nil { - select { - case retry <- struct{}{}: - default: // a retry is already due - } - } - }) - if err := check(ctx); err != nil { - if !a.tenants.Nested() { - return fmt.Errorf("dedupe open: %w", err) - } - slog.Error("dedupe: dynamodb table check failed; ingest with dedupe on fails closed while it is retried", - "table", c.Table, "error", err) - a.add(component{name: "dedupe table check", run: func(ctx context.Context) error { - for wait := time.Second; ready() != nil; wait = min(2*wait, 30*time.Second) { - select { - case <-ctx.Done(): - return nil - case <-time.After(wait): - case <-retry: - } - if err := check(ctx); err != nil { - if ctx.Err() == nil { - slog.Error("dedupe: dynamodb table check failed again; ingest with dedupe on still fails closed", - "table", c.Table, "error", err) - } - continue - } - slog.Info("dedupe: dynamodb table check passed", "table", c.Table) - apply() - } - return nil - }}) - } - apply() - return nil -} +// wireDynamoDedupe (dedupe.backend: dynamodb) lives in wire_dynamodb.go, +// excluded from the e2e coverage gate alongside internal/dedupe/dynamodb.go +// (see .testcoverage.yml): the e2e binary always runs Pebble dedupe, so +// nothing there exercises it. wireDedupe above still switches on it. // wireMQ starts the MQ — the one place the implementation is chosen; // everything after it sees mq.Broker. diff --git a/internal/app/wire_dynamodb.go b/internal/app/wire_dynamodb.go new file mode 100644 index 000000000..420111b58 --- /dev/null +++ b/internal/app/wire_dynamodb.go @@ -0,0 +1,125 @@ +package app + +import ( + "context" + "errors" + "fmt" + "log/slog" + "sync" + "time" + + "github.com/Wave-RF/WaveHouse/internal/dedupe" + "github.com/Wave-RF/WaveHouse/internal/tenant" +) + +// errDynamoUnchecked is a store's open before the first table check has run. +var errDynamoUnchecked = errors.New("dedupe: dynamodb table not checked yet") + +// wireDynamoDedupe builds the dedupe stores over one DynamoDB table that +// every tenant and every process shares (dedupe.Dynamo), so a tenant's store +// opens for free once the table has passed its check. Boot checks it (after +// creating it, with create_table on dynamodb-local) whether or not any tenant +// has dedupe on, and never creates it otherwise. A table that fails the check +// follows the registry's rule for the shape, as Pebble's instance does: a +// flat directory refuses boot; a nested one boots with every switched-on +// store closed, so its ingest fails closed. Unlike a local disk, a remote +// table's failure is usually brief (a throttle, credentials not yet issued +// mid-rollout), and a nested directory has no watcher to reload it, so the +// check is then retried in the background, with backoff, until it passes. +// The check is network I/O, so the AfterAdopt hook never runs it: the hook +// holds the lock that serializes reloads. It applies every store against the +// last check's result and wakes the retry, so a reload still retries at once. +func (a *App) wireDynamoDedupe(ctx context.Context) error { + c := a.cfg.Dedupe.DynamoDB + d, err := dedupe.NewDynamo(ctx, dedupe.DynamoConfig{ + Table: c.Table, Region: c.Region, Endpoint: c.Endpoint, + Timeout: c.Timeout, MaxAttempts: c.MaxAttempts, RetryMode: c.RetryMode, + ReserveConcurrency: a.cfg.Dedupe.ReserveConcurrency, + }) + if err != nil { + return err + } + var mu sync.Mutex + state := errDynamoUnchecked // nil once the table has passed, for good + ready := func() error { + mu.Lock() + defer mu.Unlock() + return state + } + // check is only ever run by boot, then by the retry loop, one at a time. + check := func(ctx context.Context) error { + var err error + if c.CreateTable { + err = d.CreateTable(ctx) + } + if err == nil { + err = d.Check(ctx) + } + mu.Lock() + defer mu.Unlock() + if state != nil { + state = err + } + return state + } + stores := dedupe.NewStores(dedupe.Factory(d.Tenant).Gated(ready)) + a.dedup = stores + a.add(component{name: "dedupe", close: withoutContext(stores.Close)}) + var reconciling sync.Mutex // the hook and the retry loop both apply + apply := func() { + reconciling.Lock() + defer reconciling.Unlock() + if err := stores.Retain(a.served); err != nil { + slog.Error("dedupe store close failed", "error", err) + } + for id, store := range a.tenants.All() { + m := stores.For(id) + enabled := store.DedupeEnabled() + wasOpen := m.Open() + // The one failure an open has is the check's, logged where it ran. + _ = m.Apply(enabled) + if m.Open() != wasOpen { + slog.Info("dedupe store reconciled with settings", "tenant", id, "enabled", enabled) + } + } + } + retry := make(chan struct{}, 1) + a.tenants.AfterAdopt(func([]tenant.ID) { + apply() + if ready() != nil { + select { + case retry <- struct{}{}: + default: // a retry is already due + } + } + }) + if err := check(ctx); err != nil { + if !a.tenants.Nested() { + return fmt.Errorf("dedupe open: %w", err) + } + slog.Error("dedupe: dynamodb table check failed; ingest with dedupe on fails closed while it is retried", + "table", c.Table, "error", err) + a.add(component{name: "dedupe table check", run: func(ctx context.Context) error { + for wait := time.Second; ready() != nil; wait = min(2*wait, 30*time.Second) { + select { + case <-ctx.Done(): + return nil + case <-time.After(wait): + case <-retry: + } + if err := check(ctx); err != nil { + if ctx.Err() == nil { + slog.Error("dedupe: dynamodb table check failed again; ingest with dedupe on still fails closed", + "table", c.Table, "error", err) + } + continue + } + slog.Info("dedupe: dynamodb table check passed", "table", c.Table) + apply() + } + return nil + }}) + } + apply() + return nil +} From dc9a331f324cd54f9a64dc5f4b61ff3cad047dcd Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 14:30:43 -0400 Subject: [PATCH 33/38] fix(dedupe): a caller's cancel no longer strands a DynamoDB claim Puts already sent, and the release after them, run on context.WithoutCancel(ctx), each call still bounded by its Timeout. A client that disconnects mid-Reserve no longer leaves an abandoned put to land after its release and hold the id InFlight for the lease; a Reserve whose caller cancelled after every put answered also releases them. The breaker's exemption for cancelled puts is gone: a sent put can no longer be cancelled by its caller, so its answer is always the table's. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01FyrXjhR7iDg33paioLQHFq --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- internal/dedupe/dynamodb.go | 37 ++++++----- internal/dedupe/dynamodb_test.go | 90 ++++++++++++++++++++------- 4 files changed, 89 insertions(+), 42 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 7e2996394..b1f5e3615 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, retrying for up to eight jittered rounds both the items DynamoDB leaves unprocessed and a batch it throttled whole, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. +- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, retrying for up to eight jittered rounds both the items DynamoDB leaves unprocessed and a batch it throttled whole, and `Release` is a `DeleteItem` conditional on the claim's token. A client that disconnects mid-`Reserve` leaves nothing claimed: puts already sent run to their answer and are then released, so its retry is not answered `InFlight` for the lease. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. - **Leases for work that must run in one process at a time, and the sweeper runs under one** (`internal/coord/` (new: `coord.go`, `local.go`, `elect.go`, `coordtest/`, + tests), `internal/app/{app,wire}.go` (+ tests), `internal/config/backends.go`, `internal/ingest/sweeper.go`, `.github/labeler.yml`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development,ingest-pipeline}.md`, `docs/src/content/docs/configuration.mdx`, `config.yaml`, `AGENTS.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `coord.Coordinator` hands out named leases (`TryAcquire` → a `Term` with a strictly increasing fencing `Token`, a `Done` channel and `Resign`; `ErrHeld` while another holder's term is live), and `coord.RunElected` runs a loop only while its process holds the lease, resigning when the loop returns and campaigning again every 2s. `coord.Local` is the in-process implementation, and `coordtest.Conformance` is the suite every implementation runs — the NATS KV backend that lets several replicas share one queue comes next. The sweeper now runs through `RunElected` under the `sweeper` lease; with the in-process coordinator the one process always holds it, so nothing changes for a single-process deployment beyond one `coord: elected` log line at startup. `coord.backend` now selects the coordinator (`local`, the only value), so a `config.Config` built without `config.Load` must name it as well as the other three layers' backends. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 7935f25c8..20e53dcca 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -137,7 +137,7 @@ The SSE fan-out, factored out of `api/` so the delivery hot path ([#294](https:/ - **dedupe.go** — the `Deduplicator` contract, two-phase: `Reserve(ctx, keys, lease)` answers one `Claim` per `Key{Table, ID}`, in order — `Claimed` (first sighting: the caller now holds a pending claim), `Duplicate` (committed earlier, or repeated earlier in the same call) or `InFlight` (another request holds a live claim) — and is atomic per key across every process sharing the backend; `Commit(ctx, claims, retention)` makes the published ids duplicates (retention `0` = forever); `Release(ctx, claims)` gives back ids whose publish failed. A claim neither committed nor released lapses after its lease, so a request that dies mid-publish never strands an id. There is deliberately no read-only check: a separate read is how [#390](https://github.com/Wave-RF/WaveHouse/issues/390) happened. - **key.go** — the key every backend stores, as text: `/
/` (for example `acme/clicks/evt-123`, [#222](https://github.com/Wave-RF/WaveHouse/issues/222)). The table and id are escaped and joined (`keyenc.AppendJoin`) by `internal/keyenc` — the escaping NATS subject tokens use — which never writes `/`, and a tenant id cannot hold one, so a table name may hold any byte, NUL included, and neither tenants nor tables share ids; a key is ASCII, so it reads as-is in a console and is a valid DynamoDB String. An id whose escaped form is over 1,024 bytes is stored as `#` plus its SHA-256 in hex (`#` is never written by the escaping), counted by `wavehouse_dedupe_hashed_id_total`. - **embedded.go** — `Embedded`, the [Pebble](https://github.com/cockroachdb/pebble) (embedded key-value store) implementation: every tenant's seen ids in one instance at `data_dir/pebble` ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 3). Pending claims live in memory beside it, in 64 locked shards: one process owns the instance, so a crash forgetting them is every lease lapsing at once, and the shard lock makes check-and-claim atomic. `Commit` writes every claim it is given in one batch and one fsync. `NewEmbedded(dataDir)` opens nothing; `Tenant(id)` is the `Factory` a `Stores` takes, building the tenant's `Managed` over its share of the instance, which opens with the first tenant store switched on and closes with the last one switched off. `Stats` reports the instance's figures for the system gauges, nil while it is closed. Pebble is per process: two pods on it do not share seen ids. -- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency` (64), with at least as many idle connections kept per host so a wide `Reserve` reuses them rather than dial. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read, or `Claimed` when it is the put's own pending item (same token): an SDK retry of an attempt DynamoDB applied but whose answer was lost. If any put errors, every put that may have landed is released by its token. A put already sent answers before that undo when a sibling fails, but one cut off by the caller's cancellation or its call deadline can still be applied after its release; it then holds its key `InFlight` until the lease ends, as a crashed request's claim does. `Commit` is `BatchWriteItem`, 25 at a time, retrying with jittered backoff, for up to eight rounds, both the items DynamoDB leaves unprocessed and a batch that failed transiently (a throttle means it processed none of it); the records are already published, and a table that throttles every round delays the ingest response by at most about 3 s at the defaults (eight 250 ms calls and the waits between them) before the commit is given up. `Release` is a `DeleteItem` conditional on the token and the pending state. Each call has a `Timeout` (250 ms) covering the SDK's retries (`MaxAttempts`, 3), which back off with full jitter under a ceiling capped at `Timeout/(2·(MaxAttempts−1))`, so a call's retries wait at most half its timeout and a throttled call fails on its last attempt's answer rather than on the deadline. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. +- **dynamodb.go** — `Dynamo`, the DynamoDB implementation, built and tested but not yet selectable at boot (the boot key comes with [#613](https://github.com/Wave-RF/WaveHouse/issues/613)'s boot config): every tenant's ids in one shared table, `pk` (a string) the key above, no sort key. `Reserve` is a conditional `PutItem` per key, run in parallel up to `ReserveConcurrency` (64), with at least as many idle connections kept per host so a wide `Reserve` reuses them rather than dial. It succeeds when no live item holds the key, where an item whose `ex` (epoch seconds, rounded up) has passed counts as absent whether or not TTL has deleted it yet. On a failed condition, the returned old item says `Duplicate` or `InFlight` without a read, or `Claimed` when it is the put's own pending item (same token): an SDK retry of an attempt DynamoDB applied but whose answer was lost. If any put errors, or the caller cancels (a client disconnecting mid-request), the puts not yet sent are skipped and every put that may have landed is released by its token. A put already sent runs to its answer on a context the caller's cancellation does not reach, so it answers before that undo; only its own call deadline can cut it off, and DynamoDB may then apply it after its release, holding its key `InFlight` until the lease ends, as a crashed request's claim does. `Commit` is `BatchWriteItem`, 25 at a time, retrying with jittered backoff, for up to eight rounds, both the items DynamoDB leaves unprocessed and a batch that failed transiently (a throttle means it processed none of it); the records are already published, and a table that throttles every round delays the ingest response by at most about 3 s at the defaults (eight 250 ms calls and the waits between them) before the commit is given up. `Release` is a `DeleteItem` conditional on the token and the pending state. Each call has a `Timeout` (250 ms) covering the SDK's retries (`MaxAttempts`, 3), which back off with full jitter under a ceiling capped at `Timeout/(2·(MaxAttempts−1))`, so a call's retries wait at most half its timeout and a throttled call fails on its last attempt's answer rather than on the deadline. Throttling, server faults, timeouts and connection failures wrap `ErrUnavailable`; a missing table or denied access does not, since those are configuration bugs. Five unavailable claims in a row within a second short-circuit `Reserve` for a second, for every tenant (one breaker per `Dynamo`). `NewDynamo` builds the client from the AWS SDK's default chain (Pod Identity or IRSA), with an `Endpoint` override for dynamodb-local. `Check` verifies the key schema and warns when TTL is off. `CreateTable` is refused unless `Endpoint` is set. `Tenant(id)` is the `Factory`. The table definition and IAM policy are on the [Deployment](/deployment) page. - **managed.go** — `Managed` wraps one store — opened through the function `NewManaged` takes, so the switch semantics are the same for every backend — behind the hot-reloadable `dedupe.enabled` switch: `Apply(enabled)` opens or closes it, idempotently, and in-flight calls are serialized against the swap, so flipping the key is a reload, not a restart. Every call returns `ErrDisabled` while switched off (the ingest handler publishes un-deduped and counts it — a reload-window race, not a mode) and `ErrUnavailable` while switched on but not open (ingest fails closed). `Reserve` also reads a lease `<= 0` as `DefaultLease` and collapses a key repeated in one call before the backend sees it, once for every backend — so a backend may assume distinct keys, a positive lease, and only `Claimed` claims in `Commit` and `Release`. - **dedupetest/** — the conformance suite every backend runs: `Run(t, newHarness)` drives the contract above through a backend's `Factory` (claim, commit, release, lease lapse, retention, one claim among concurrent reserves from two clients, keyspaces, input order, late commit, stale release, failure mid-call); `Harness` optionally injects a clock and a mid-call failure. `Mark` is the old check-and-mark in one call, for tests that only need an id seen. - **stores.go** — `Stores` is one `Managed` per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7), built on first use through a `Factory` (`func(tenant.ID) *Managed`) — whether tenants share a backend is the factory's business (`Embedded.Tenant` puts them all in one Pebble instance), with nothing that holds the `Stores` changing. `For(id)` returns a tenant's store, built closed so a tenant adopted a moment ago answers `ErrDisabled` rather than having no store; `Retain(keep)` closes and forgets the stores of tenants no longer served, touching nothing on disk; `Close()` closes every store. `internal/app` drives it from the registry's `AfterAdopt` hook. diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index cc3321f7c..e9e8cd477 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -303,10 +303,7 @@ func (d *Dynamo) call(ctx context.Context, op string, do func(context.Context) e start := time.Now() err := classify(op, do(ctx)) d.metrics.record(ctx, op, time.Since(start), err) - // A request cancelled because its caller went away (a client - // disconnecting mid-Reserve) says nothing about the table, and must not - // reset the breaker's count. - if op == opReserve && !errors.Is(err, context.Canceled) { + if op == opReserve { d.breaker.record(err) } return err @@ -321,9 +318,10 @@ type dynamoStore struct { // Reserve puts every key's pending item in parallel, each conditional on no // live item holding the key. A failed condition hands back the live item, // whose state says Duplicate or InFlight without a read, or Claimed when its -// token is the put's own. On any error it releases, by token, every put that -// may have landed; what that undo misses (below) holds its key InFlight -// until the lease ends, as a crashed request's claim does. +// token is the put's own. On any error, the caller's cancellation included, +// it releases, by token, every put that may have landed; what that undo +// misses (below) holds its key InFlight until the lease ends, as a crashed +// request's claim does. func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Duration) ([]Claim, error) { if len(keys) == 0 { return []Claim{}, nil @@ -338,12 +336,13 @@ func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Durati claims := make([]Claim, len(keys)) tried := make([]Claim, len(keys)) sent := make([]bool, len(keys)) - // The first failure skips the puts not yet sent: the Reserve fails - // either way, and a throttled table should not take the rest. A put - // already sent runs on ctx, not gctx, so a sibling's failure never cuts - // it off: it answers before the undo below. The caller's cancellation or - // the put's own deadline can, and DynamoDB may then apply it after its - // release. + // The first failure, or the caller's cancellation, skips the puts not + // yet sent: the Reserve fails either way, and a throttled table should + // not take the rest. A put already sent runs on sendCtx, which neither a + // sibling's failure nor the caller's cancellation reaches, so it answers + // before the undo below. Only its own Timeout (call) can cut it off, and + // DynamoDB may then apply it after its release. + sendCtx := context.WithoutCancel(ctx) g, gctx := errgroup.WithContext(ctx) g.SetLimit(s.d.cfg.ReserveConcurrency) for i, k := range keys { @@ -354,7 +353,7 @@ func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Durati return err } sent[i] = true - status, err := s.reserve(ctx, k, token, nowSec, exp) + status, err := s.reserve(sendCtx, k, token, nowSec, exp) if err != nil { return err } @@ -365,7 +364,13 @@ func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Durati return nil }) } - if err := g.Wait(); err != nil { + err := g.Wait() + if err == nil { + // Every sent put answered, but a caller that went away mid-Reserve + // will never commit or release what it claimed. + err = ctx.Err() + } + if err != nil { // A sent put that errored may have landed anyway; releasing a key // its token does not hold is a no-op. Best effort: a put applied // after this, or a release that fails, lapses with the lease. @@ -375,7 +380,7 @@ func (s *dynamoStore) Reserve(ctx context.Context, keys []Key, lease time.Durati undo = append(undo, tried[i]) } } - _ = s.Release(context.WithoutCancel(ctx), undo) + _ = s.Release(sendCtx, undo) return nil, err } return claims, nil diff --git a/internal/dedupe/dynamodb_test.go b/internal/dedupe/dynamodb_test.go index 03a7196f8..99ac8b9a9 100644 --- a/internal/dedupe/dynamodb_test.go +++ b/internal/dedupe/dynamodb_test.go @@ -633,34 +633,76 @@ func TestDynamo_CommitAttemptsEveryChunk(t *testing.T) { assert.Equal(t, 2*batchWriteMax, written, "a failed chunk does not cancel the others: their records are published") } -// A put cancelled by its caller is not an answer from the table: it does -// not reset the breaker's count of throttled puts. -func TestDynamo_CallerCancelDoesNotResetBreaker(t *testing.T) { +// A caller that cancels mid-Reserve (a client disconnecting) leaves nothing +// claimed. The fake applies a put after a delay whatever the caller does, as +// DynamoDB applies a request already on the wire, so a put abandoned on the +// cancel would land after its release and hold its id for the lease. +func TestDynamo_CallerCancelLeavesNothingClaimed(t *testing.T) { t.Parallel() - var hang atomic.Bool - started := make(chan struct{}, 1) - _, m := openFake(t, &fakeDynamo{put: func(ctx context.Context, _ *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { - if hang.Load() { + t.Run("a put still unsent", func(t *testing.T) { + t.Parallel() + assertCancelLeavesNothing(t, keys("k0", "k1", "k2")) + }) + t.Run("every put sent", func(t *testing.T) { + t.Parallel() + assertCancelLeavesNothing(t, keys("k0", "k1")) + }) +} + +// assertCancelLeavesNothing cancels a Reserve of ks once two puts are sent. +func assertCancelLeavesNothing(t *testing.T, ks []Key) { + t.Helper() + var ( + mu sync.Mutex + table = map[string]string{} + sent []string + applies sync.WaitGroup + ) + started := make(chan struct{}, 2) + _, m := openFakeWith(t, &fakeDynamo{ + put: func(ctx context.Context, in *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { + id := idOf(in.Item[attrKey]) + tk := string(in.Item[attrToken].(*types.AttributeValueMemberB).Value) + mu.Lock() + sent = append(sent, id) + mu.Unlock() + applied := make(chan struct{}) + applies.Go(func() { + time.Sleep(20 * time.Millisecond) + mu.Lock() + table[id] = tk + mu.Unlock() + close(applied) + }) started <- struct{}{} - <-ctx.Done() - return nil, ctx.Err() - } - return nil, &types.ProvisionedThroughputExceededException{} - }}) - for range breakerTrips - 1 { - _, err := m.Reserve(t.Context(), keys("a"), time.Minute) - require.ErrorIs(t, err, ErrUnavailable) - } - hang.Store(true) + select { + case <-applied: + return &dynamodb.PutItemOutput{}, nil + case <-ctx.Done(): + return nil, ctx.Err() + } + }, + del: func(_ context.Context, in *dynamodb.DeleteItemInput) (*dynamodb.DeleteItemOutput, error) { + id := idOf(in.Key[attrKey]) + tk := string(in.ExpressionAttributeValues[":tk"].(*types.AttributeValueMemberB).Value) + mu.Lock() + defer mu.Unlock() + if table[id] != tk { + return nil, &types.ConditionalCheckFailedException{} + } + delete(table, id) + return &dynamodb.DeleteItemOutput{}, nil + }, + }, DynamoConfig{Table: "dedupe", ReserveConcurrency: 2}) ctx, cancel := context.WithCancel(t.Context()) - go func() { <-started; cancel() }() - _, err := m.Reserve(ctx, keys("a"), time.Minute) + go func() { <-started; <-started; cancel() }() + _, err := m.Reserve(ctx, ks, time.Minute) require.ErrorIs(t, err, context.Canceled) - hang.Store(false) - _, err = m.Reserve(t.Context(), keys("a"), time.Minute) - require.ErrorIs(t, err, ErrUnavailable) - _, err = m.Reserve(t.Context(), keys("a"), time.Minute) - require.ErrorIs(t, err, errBreakerOpen, "the cancelled put did not reset the count") + applies.Wait() + mu.Lock() + defer mu.Unlock() + assert.Len(t, sent, 2, "a put not yet sent is skipped") + assert.Empty(t, table, "every applied put was released") } func TestDynamo_ReleaseAttemptsEveryClaim(t *testing.T) { From 07e4f77d30dd7200a06a9999d0fffb266a405fb7 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 14:31:07 -0400 Subject: [PATCH 34/38] fix(dedupe): name ErrUnavailable "dedupe store unavailable" A throttle wraps it too, and logged as "dedupe store is not open". The classify comment now says what ingest answers today: 500 for both, until #629 maps ErrUnavailable to a 503. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01FyrXjhR7iDg33paioLQHFq --- internal/dedupe/dynamodb.go | 5 +++-- internal/dedupe/managed.go | 2 +- tests/integration/dedupe_dynamodb_test.go | 4 ++-- 3 files changed, 6 insertions(+), 5 deletions(-) diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index e9e8cd477..88dd557b4 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -574,8 +574,9 @@ func newToken() string { // classify maps a DynamoDB error onto the contract: a condition failure is // returned as is for the caller to read, anything retrying later can cure -// wraps ErrUnavailable (503), and the rest — a missing table, denied access, -// a malformed request — is a configuration bug (500). +// wraps ErrUnavailable, and the rest — a missing table, denied access, a +// malformed request — is a configuration bug. Ingest answers both 500 until +// #629 maps ErrUnavailable to a retryable 503. func classify(op string, err error) error { if err == nil { return nil diff --git a/internal/dedupe/managed.go b/internal/dedupe/managed.go index cd42aeb31..09712c50b 100644 --- a/internal/dedupe/managed.go +++ b/internal/dedupe/managed.go @@ -23,7 +23,7 @@ var ErrDisabled = errors.New("dedupe is disabled") // but the store failed to open, and wrapped by a backend's error when a // retry later can succeed. Ingest fails closed on it — the settings asked // for dedupe, so publishing un-deduped is not a fallback. -var ErrUnavailable = errors.New("dedupe store is not open") +var ErrUnavailable = errors.New("dedupe store unavailable") // hashedIDCounter counts ids stored as their SHA-256 (Key.Hashed): an id // longer than MaxIDBytes is a producer sending something other than an id. diff --git a/tests/integration/dedupe_dynamodb_test.go b/tests/integration/dedupe_dynamodb_test.go index 69be6b16e..a3cf576b6 100644 --- a/tests/integration/dedupe_dynamodb_test.go +++ b/tests/integration/dedupe_dynamodb_test.go @@ -194,7 +194,7 @@ func TestDedupeDynamo_Throttled(t *testing.T) { m := d.Tenant("acme") require.NoError(t, m.Apply(true)) _, err := m.Reserve(t.Context(), []dedupe.Key{{Table: "events", ID: "e1"}}, time.Minute) - require.ErrorIs(t, err, dedupe.ErrUnavailable, "a throttle is worth retrying: 503") + require.ErrorIs(t, err, dedupe.ErrUnavailable, "a throttle is worth retrying") assert.Equal(t, int64(2), sent.Load(), "the SDK retried it once first") }) } @@ -257,7 +257,7 @@ func TestDedupeDynamo_ConfigErrorsAreNotUnavailable(t *testing.T) { require.Error(t, err) var missing *types.ResourceNotFoundException assert.ErrorAs(t, err, &missing) - assert.False(t, errors.Is(err, dedupe.ErrUnavailable), "a missing table is a config bug: 500, not 503") + assert.False(t, errors.Is(err, dedupe.ErrUnavailable), "a missing table is a config bug, not worth retrying") require.Error(t, d.Check(t.Context())) } From 138b166186718d74da9331fd0d0523dd8c059830 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 14:36:14 -0400 Subject: [PATCH 35/38] feat(app): refuse boot only for a misconfigured DynamoDB table in use Over a flat settings directory, boot was refused on any failed table check. Now it is refused only when the failure is a misconfiguration (not ErrUnavailable: a missing table, the wrong key schema, access denied) and a tenant has dedupe on. A transient failure, or a misconfigured table no tenant uses yet, boots with the switched-on stores closed and the check retried in the background, as a nested directory already did; a tenant a reload switches on fails closed until it passes. The misconfiguration is logged at ERROR. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01FyrXjhR7iDg33paioLQHFq --- CHANGELOG.md | 2 +- docs/src/content/docs/architecture.md | 2 +- docs/src/content/docs/configuration.mdx | 2 +- docs/src/content/docs/deployment.md | 2 +- internal/app/dedupe_dynamodb_test.go | 123 ++++++++++++++++++++---- internal/app/wire_dynamodb.go | 38 ++++++-- 6 files changed, 134 insertions(+), 35 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 5d405dc6e..098224c66 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire,wire_dynamodb}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `.testcoverage.yml`, `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59s` with the embedded queue, so that the lease plus its own ceiling to the next second plus one more second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`, which also sizes the DynamoDB client's idle connections per host; its fan-out has no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a failure refuses boot over a flat settings directory and, over a nested one, fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and at once after every reload. A reload makes no table call and does not wait on a tenant whose dedupe setting is unchanged: it holds the lock that serializes reloads, so it applies each tenant's switch against the last check's result and only wakes the retry; it waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight calls, before its store closes. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. +- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire,wire_dynamodb}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `.testcoverage.yml`, `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59s` with the embedded queue, so that the lease plus its own ceiling to the next second plus one more second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`, which also sizes the DynamoDB client's idle connections per host; its fan-out has no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a misconfigured table (missing, the wrong key schema, access denied) refuses boot over a flat settings directory whose tenant has dedupe on and is logged at `ERROR` otherwise; any other failure (a throttle, a timeout, the network), a nested directory, or no tenant deduping yet boots and fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and at once after every reload. A reload makes no table call and does not wait on a tenant whose dedupe setting is unchanged: it holds the lock that serializes reloads, so it applies each tenant's switch against the last check's result and only wakes the retry; it waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight calls, before its store closes. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. - **A DynamoDB dedupe backend** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, retrying for up to eight jittered rounds both the items DynamoDB leaves unprocessed and a batch it throttled whole, and `Release` is a `DeleteItem` conditional on the claim's token. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. `dedupe.backend: dynamodb` selects it (entry above). New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/architecture.md b/docs/src/content/docs/architecture.md index 579f5a216..223b3cecf 100644 --- a/docs/src/content/docs/architecture.md +++ b/docs/src/content/docs/architecture.md @@ -95,7 +95,7 @@ The API layer uses [Chi](https://github.com/go-chi/chi) for routing with Request - **app.go** — `New(ctx, Options)` builds every component from the boot config (`Options.Config`) and the settings directory it names, in dependency order: settings registry, observability (after which each `Config.Warnings` line is logged at `WARN`), ClickHouse pools, schema discovery, the dedupe stores, embedded NATS (ingest + DLQ streams), cache, the lease coordinator, sweeper, streaming (hub, MQ→hub bridge, keepalive wheel), ingest worker, auth, reload triggers, HTTP. The boot config's `roles` decide which of them a process wires: every process gets the settings registry, observability, the MQ, the coordinator, the reload triggers and a listener; `api` adds schema discovery, the dedupe stores, streaming, auth and the full router; `ingest` adds the ingest worker; `sweeper` adds the sweeper; the ClickHouse pools and the cache come with `api` or `ingest`. A process without `api` serves `api.NewOpsRouter` (probes, `/version`, the metrics path, and the settings reload behind the operator key alone, `wireOpsAuth`) on `server.port`. `config.Validate` refuses a role set the backends cannot serve (a split over the embedded MQ, or `api` without `ingest` and the reverse over a local cache), and `New` refuses a `Config` with no roles, which only one built without `config.Load` can have. Each is one `component` value — what it opens, what it loops, what it releases — so a failure part-way releases what was already opened and returns the error. `Run(ctx)` drives every loop under one `errgroup` until `ctx` is canceled (a clean stop: every loop drains, the API server and the ingest worker within `server.shutdown_timeout`; open SSE streams are ended as the drain begins rather than waited on) or a component fails, which stops the rest and returns that error. `Close(ctx)` releases what `New` opened, newest first, under the caller's release budget (`ReleaseTimeout`, 5s), a real bound: a remote implementation's close gives up at the deadline itself, and a close that ignores the context (the local stores) is abandoned at it, with the components below it left unreleased rather than overlapping it, both named in the error — and then flushes telemetry under its own 3s budget, so the flush that reports on the stop is never handed a deadline a slow close already spent. The SIGHUP registration is released last of all. `Handler`, `Registry`, and `MQ` expose the pieces a harness needs; `Options.Listener` lets one serve the API on its own listener instead of `server.port`. - **wire.go** — one `wire*` function per component, each handed the settings registry whole and deriving the per-call getters the internal packages take (`DLQFor`, `DedupeFor`, `GapWindow`, …) and registering its `AfterAdopt` hook there where it has one. A layer with a choice of implementation — `wireMQ`, `wireCache`, `wireDedupe`, `wireCoord` — picks it there and nowhere else, in a `switch` on the boot config's `.backend` with one case per backend; the default case refuses boot, which only a `config.Config` built without `config.Load` reaches, since `Validate` refuses a value no case handles. Those wiring functions are where the per-tenant registry of [#583](https://github.com/Wave-RF/WaveHouse/issues/583) is injected, not `main`: `wireSettings` opens the `settings.Registry`, the HTTP handlers get store-keyed getters (method expressions such as `(*settings.Store).Policy`), and `perTenant` adapts a store accessor into the `func(tenant.ID) T` getter the async packages take, with the tenant each message's `mq.Topic` names for the stream hub and the ingest worker — a tenant the registry is not serving is logged and read as the zero value, except in `dlqFor`, the ingest worker's DLQ switch, where it reads as on so a message the worker cannot read is parked rather than dropped, and a removed or rejected tenant's queued rows are parked rather than left unacked, where each would be redelivered every ack wait for as long as the tenant is away and would hold that tenant's ack floor, so the sweeper could purge none of its queue past it. The ClickHouse pools (`chconn.Pools`) and the per-tenant schema registries (`discoveries`, in `discoveries.go`) are reconciled from `AfterAdopt` after every reload ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 6): `wireClickHouse` builds each served tenant's `chconn.Member` from its store and logs what the reconcile refused; `wireDiscovery` builds a registry over `pools.For` for each newly served tenant — a flat directory's tenant `0` refreshed synchronously first, as before — runs its loop under the App's stop context, stops the loop of a tenant no longer served, and drives the `BootState` from the first tenant's first discovery, sticky from there; before that, a diagnostic naming a tenant a reload stopped serving goes back to the no-tenant one. The handlers resolve both per request through store-keyed getters (`chConnFor`, `registryFor`, `chTargetFor`, `queryTimeout`), the hub and the ingest worker through tenant-keyed ones (`discoveries.For`, `pools.Target`) called with the tenant the message's topic names; a tenant on no pool is an untyped nil connection, the handlers' `503`. The ingest worker is handed the cache through `sharedTables`, which bumps each namespace the worker invalidates under every tenant on the same ClickHouse address and database (`pools.SharingTables`), and the pools hook orphans the table-keyed cache — the structured-query results — of a tenant back on a pool after an absence (`Cache.InvalidateTenant`), since it was out of that fan-out while away, and of a tenant moved to another address or database, since it now reads other tables (both returned by `Pools.Reconcile`). The one setting that still follows the default tenant is read per request, the admin role of a flat directory's ops gate: `defaultPolicy` reads it through the registry, where a flat directory's tenant `0` is always served. The auth verifiers are per tenant: `wireAuth` builds one for each tenant being served, its `AfterAdopt` hook reconfigures the adopted tenants' (rebuilt only when their wiring changed) and prunes the ones no longer served, and the operator key's admin role is read from the request tenant's policy. `wireStreaming`'s hook prunes the stream hub the same way (`Hub.Prune`, with the one `served` predicate the auth and dedupe hooks use too), ending the open streams of a tenant no longer served. One setting is shared by folding over the tenants being served rather than by following tenant `0`: the keepalive wheel runs at the shortest `stream.keepalive_interval` among them (`shortestKeepalive`), re-derived after every reload the registry applies — an adoption, a rejection, or a removal — so a dropped tenant's interval leaves the wheel at once ([#597](https://github.com/Wave-RF/WaveHouse/issues/597)). The sweeper is handed each tenant's own `stream.gap_window_minutes` (`gapWindows`, read every sweep over `Registry.Known`, so a rejected tenant keeps the window its folder last had, and all of its history if the folder has been rejected since boot), since each tenant's events have a queue of their own, and runs only while its process holds the `sweeper` lease (`elected`, which wraps `coord.RunElected` over the coordinator `wireCoord` opens; `local`, the only `coord.backend`, keeps leases in the process, so the one process always holds it). The dedupe stores are per tenant ([#583](https://github.com/Wave-RF/WaveHouse/issues/583) story 7): `wireDedupe`'s `pebble` case builds a `dedupe.Stores` over the `Tenant` factory of the embedded Pebble implementation (`dedupe.NewEmbedded`), handing it `data_dir` once; the implementation decides where every tenant's store lives — one instance, each key led by its tenant (story 3) and then its table — and one reconcile closure, the boot apply and the `AfterAdopt` hook alike, sets every store to what the registry says: open exactly when its tenant is served with `dedupe.enabled` on, closed with its seen ids kept when the tenant is switched off, rejected, or removed. An instance that cannot open follows the registry's rule for the shape: fatal at boot over a flat directory, fail-closed for every tenant with dedupe on over a nested one. The system gauges report that one instance's figures (`Embedded.Stats`), not a sum over tenants. `wireDedupe`'s `dynamodb` case is `wireDynamoDedupe`, in wire_dynamodb.go (below). `wireHTTP` hands the ingest handler `dedupe.lease` (`IngestHandler.DedupeLease`) whichever backend is chosen. The ingest handler picks the tenant's store off the request's `settings.Store` (`Store.Tenant()`). The reload triggers only start in `Run`, after `New` has registered every hook, so the watcher's first reload already drives all of them: SIGHUP in both shapes, the directory watcher for a flat directory only. `wireMQ`'s `embedded` case hands each served tenant's `mq.max_bytes_gb` to `mq.Broker.SetMaxBytes` at boot, under `New`'s context (so a stop signaled mid-boot is not held up by opening many queues), and again after every reload, under the App's stop context; the first apply opens that tenant's queue. A queue that cannot be opened or resized follows the registry's rule for the shape — fatal at boot over a flat directory, logged over a nested one — and is retried by the next reload, a queue that did not open by the next publish too. How the budget is split across the tenant's streams, the time bounds, the rollback, and the dead-letter shrink guard are `internal/mq`'s. -- **wire_dynamodb.go** — `wireDedupe`'s `dynamodb` case, split out of wire.go so the e2e suite's coverage exclude for it (the e2e binary always runs Pebble dedupe, never DynamoDB) doesn't have to blanket wire.go itself: builds the same `dedupe.Stores` over `Dynamo.Tenant`, gated (`Factory.Gated`) on the table's check: boot runs `Dynamo.Check` (after `CreateTable`, when `dedupe.dynamodb.create_table` is on) whether or not any tenant has dedupe on, and a table that fails it follows the same rule as a Pebble instance that cannot open, with the check retried until it passes by a background component that backs off from one second to thirty (a nested directory has no watcher). The `AfterAdopt` hook never runs the check, since it holds the lock that serializes reloads, and it does not wait on a tenant whose `dedupe.enabled` is unchanged either — `Managed.Apply`'s no-op fast path settles that case under its own read lock, so the hook only takes a store's write lock, and so waits for that tenant's in-flight `Reserve`/`Commit`/`Release` calls to finish, on a genuine flip. It applies every store against the last check's result, so a tenant a reload switches on fails closed meanwhile, and wakes the retry, so a reload still retries at once. It has no Pebble gauges. +- **wire_dynamodb.go** — `wireDedupe`'s `dynamodb` case, split out of wire.go so the e2e suite's coverage exclude for it (the e2e binary always runs Pebble dedupe, never DynamoDB) doesn't have to blanket wire.go itself: builds the same `dedupe.Stores` over `Dynamo.Tenant`, gated (`Factory.Gated`) on the table's check: boot runs `Dynamo.Check` (after `CreateTable`, when `dedupe.dynamodb.create_table` is on) whether or not any tenant has dedupe on. Boot is refused only for a misconfigured table (an error that is not `ErrUnavailable`) over a flat directory whose tenant has dedupe on; every other failure boots with the switched-on stores closed, the check retried until it passes by a background component that backs off from one second to thirty (a nested directory has no watcher, and a flat one's table can come good with no settings change). The `AfterAdopt` hook never runs the check, since it holds the lock that serializes reloads, and it does not wait on a tenant whose `dedupe.enabled` is unchanged either — `Managed.Apply`'s no-op fast path settles that case under its own read lock, so the hook only takes a store's write lock, and so waits for that tenant's in-flight `Reserve`/`Commit`/`Release` calls to finish, on a genuine flip. It applies every store against the last check's result, so a tenant a reload switches on fails closed meanwhile, and wakes the retry, so a reload still retries at once. It has no Pebble gauges. ### `stream/` — SSE keepalive & fan-out diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index bd1781d00..4c23d2d97 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -61,7 +61,7 @@ Whether a tenant dedupes, and on which field, are settings-directory keys ([Dedu #### DynamoDB dedupe -Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (String) alone, and TTL off on `ex` is logged as a warning. A table that fails the check, or cannot be reached, refuses boot with a flat settings directory. With a nested one, the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and the check is retried in the background, backing off from one second to thirty, and at once after every reload, so a table that comes good is picked up without a restart. A reload makes no table call, and does not wait on a tenant whose dedupe setting is unchanged: it applies each tenant's switch against the last check's result, so a tenant it switches on fails closed until the retry passes. A reload waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight `Reserve`/`Commit`/`Release` calls, before the store itself closes. The check runs in every process running the `api` [role](#process-roles), the one that opens the dedupe stores, whether or not any tenant has dedupe on. +Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (String) alone, and TTL off on `ex` is logged as a warning. A misconfigured table — missing, with the wrong key schema, or denied to the process's credentials — refuses boot only with a flat settings directory whose tenant has dedupe on, and is logged at `ERROR` otherwise. In every other case — a transient failure (a throttle, a timeout, the network), a nested directory, or no tenant with dedupe on yet — the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and the check is retried in the background, backing off from one second to thirty, and at once after every reload, so a table that comes good is picked up without a restart. A tenant a reload switches dedupe on for fails closed the same way until then. A reload makes no table call, and does not wait on a tenant whose dedupe setting is unchanged: it applies each tenant's switch against the last check's result, so a tenant it switches on fails closed until the retry passes. A reload waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight `Reserve`/`Commit`/`Release` calls, before the store itself closes. The check runs in every process running the `api` [role](#process-roles), the one that opens the dedupe stores, whether or not any tenant has dedupe on. | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index ddf377953..1aea0f34e 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -512,7 +512,7 @@ dedupe: region: us-east-1 # or leave empty for AWS_REGION ``` -or `WH_DEDUPE_BACKEND=dynamodb`, `WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod`. A table that is missing, has the wrong key schema, or cannot be reached with the pod's credentials refuses boot over a flat settings directory; over a nested one the pod boots, every tenant with dedupe on fails its ingest closed, and the check is retried in the background (backing off from one second to thirty, and at once after every reload). A reload makes no table call itself, and does not wait on a tenant whose dedupe setting is unchanged; it waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight calls, before its store closes. No region at all (neither `region` nor one from the SDK chain: `AWS_REGION`, `AWS_DEFAULT_REGION` or a profile) refuses boot in both shapes. The check runs in every pod running the `api` [role](/configuration#process-roles), whether or not any tenant has `dedupe.enabled` on; a pod without it opens no dedupe store. The per-tenant switch stays in each tenant's `config.json`. +or `WH_DEDUPE_BACKEND=dynamodb`, `WH_DEDUPE_DYNAMODB_TABLE=wavehouse-dedupe-prod`. A table that is missing, has the wrong key schema, or refuses the pod's credentials refuses boot over a flat settings directory whose tenant has dedupe on, and is logged at `ERROR` otherwise. In every other case — a throttle or network failure, a nested directory, or no tenant with dedupe on — the pod boots, every tenant with dedupe on (now or after a reload) fails its ingest closed, and the check is retried in the background (backing off from one second to thirty, and at once after every reload). A reload makes no table call itself, and does not wait on a tenant whose dedupe setting is unchanged; it waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight calls, before its store closes. No region at all (neither `region` nor one from the SDK chain: `AWS_REGION`, `AWS_DEFAULT_REGION` or a profile) refuses boot in both shapes. The check runs in every pod running the `api` [role](/configuration#process-roles), whether or not any tenant has `dedupe.enabled` on; a pod without it opens no dedupe store. The per-tenant switch stays in each tenant's `config.json`. For development against dynamodb-local, set `dedupe.dynamodb.endpoint` (for example `http://localhost:8000`) and `create_table: true`, and give the SDK any static credentials (`AWS_ACCESS_KEY_ID`, `AWS_SECRET_ACCESS_KEY`) and a region. `create_table` without an `endpoint` refuses boot. diff --git a/internal/app/dedupe_dynamodb_test.go b/internal/app/dedupe_dynamodb_test.go index 310d11408..b2ab5ab53 100644 --- a/internal/app/dedupe_dynamodb_test.go +++ b/internal/app/dedupe_dynamodb_test.go @@ -1,8 +1,10 @@ package app import ( + "bytes" "context" "io" + "log/slog" "net" "net/http" "net/http/httptest" @@ -23,14 +25,21 @@ import ( // fakeDynamo answers the DynamoDB JSON protocol for one table, enough for // boot's check, the dev create path, and a claim and its commit. Whether the -// table exists, and whether the endpoint hangs (every call, or one op -// alone), are the test's to switch. +// table exists, whether every call is throttled, and whether the endpoint +// hangs (every call, or one op alone), are the test's to switch. type fakeDynamo struct { - mu sync.Mutex - exists bool - hangs bool - hangOn string // hang calls of this op alone, once set; "" hangs none this way - calls []string + mu sync.Mutex + exists bool + throttles bool + hangs bool + hangOn string // hang calls of this op alone, once set; "" hangs none this way + calls []string +} + +func (f *fakeDynamo) setThrottles(v bool) { + f.mu.Lock() + defer f.mu.Unlock() + f.throttles = v } func (f *fakeDynamo) setExists(v bool) { @@ -76,13 +85,18 @@ func (f *fakeDynamo) ServeHTTP(w http.ResponseWriter, r *http.Request) { if op == "CreateTable" { f.exists = true } - exists, hang := f.exists, f.hangs || op == f.hangOn + exists, throttled, hang := f.exists, f.throttles, f.hangs || op == f.hangOn f.mu.Unlock() if hang { <-r.Context().Done() return } w.Header().Set("Content-Type", "application/x-amz-json-1.0") + if throttled { + w.WriteHeader(http.StatusBadRequest) + _, _ = io.WriteString(w, `{"__type":"com.amazonaws.dynamodb.v20120810#ThrottlingException","message":"Rate exceeded"}`) + return + } if !exists { w.WriteHeader(http.StatusBadRequest) _, _ = io.WriteString(w, `{"__type":"com.amazonaws.dynamodb.v20120810#ResourceNotFoundException","message":"Requested resource not found"}`) @@ -154,20 +168,35 @@ func TestNew_DynamoDBDedupeCreatesTheTableOnlyWhenAsked(t *testing.T) { assert.True(t, a.dedup.For(tenant.Default).Open()) } -// A table that fails the check follows the registry's rule for the shape, -// as a Pebble instance that cannot open does. +// A misconfigured table refuses boot only over a flat directory in which a +// tenant has dedupe on; every other failure boots and fails closed. func TestNew_DynamoDBDedupeTableMissing(t *testing.T) { - t.Run("flat refuses boot", func(t *testing.T) { - for name, patch := range map[string]map[string]any{"dedupe on": dedupeOn, "dedupe off": nil} { - t.Run(name, func(t *testing.T) { - guardGlobals(t) - cfg := testConfig(t, writeSettings(t, patch)) - dynamoConfig(t, cfg, false) - _, err := New(t.Context(), Options{Config: cfg}) - require.ErrorContains(t, err, "dedupe open") - require.ErrorContains(t, err, "ResourceNotFoundException") - }) - } + t.Run("flat with dedupe on refuses boot", func(t *testing.T) { + guardGlobals(t) + cfg := testConfig(t, writeSettings(t, dedupeOn)) + dynamoConfig(t, cfg, false) + _, err := New(t.Context(), Options{Config: cfg}) + require.ErrorContains(t, err, "dedupe open") + require.ErrorContains(t, err, "ResourceNotFoundException") + require.NotErrorIs(t, err, dedupe.ErrUnavailable) + }) + t.Run("flat with dedupe off boots, and fails closed once it is on", func(t *testing.T) { + dir := writeSettings(t, nil) + cfg := testConfig(t, dir) + dynamoConfig(t, cfg, false) + logs := bootLogged(t) + a, err := New(t.Context(), Options{Config: cfg}) + require.NoError(t, err) + t.Cleanup(func() { assert.NoError(t, a.Close(context.Background())) }) + assert.Contains(t, logs.String(), `level=ERROR msg="dedupe: dynamodb table is misconfigured`) + + store := a.dedup.For(tenant.Default) + _, err = dedupetest.Mark(t.Context(), store, eventKey) + require.ErrorIs(t, err, dedupe.ErrDisabled) + rewriteSettings(t, dir, dedupeOn) + a.tenants.Reload("test") + _, err = dedupetest.Mark(t.Context(), store, eventKey) + require.ErrorIs(t, err, dedupe.ErrUnavailable, "switched on by a reload while the table is missing: closed, not un-deduped") }) t.Run("nested fails closed", func(t *testing.T) { root := writeNestedSettings(t, map[string]map[string]any{"acme": dedupeOn, "globex": nil}) @@ -184,6 +213,58 @@ func TestNew_DynamoDBDedupeTableMissing(t *testing.T) { }) } +// A transient failure (a throttle) never refuses boot, even over a flat +// directory with dedupe on: the tenant fails closed until the background +// retry's check passes. +func TestRun_DynamoDBDedupeFlatThrottledRecovers(t *testing.T) { + cfg := testConfig(t, writeSettings(t, dedupeOn)) + fake := dynamoConfig(t, cfg, true) + fake.setThrottles(true) + var lc net.ListenConfig + ln, err := lc.Listen(t.Context(), "tcp", "127.0.0.1:0") + require.NoError(t, err) + a := newApp(t, cfg, Options{Listener: ln}) + store := a.dedup.For(tenant.Default) + require.False(t, store.Open()) + _, err = dedupetest.Mark(t.Context(), store, eventKey) + require.ErrorIs(t, err, dedupe.ErrUnavailable, "switched on, table throttled: ingest fails closed") + + _, stop := runApp(t, a, ln) + fake.setThrottles(false) + require.Eventually(t, store.Open, 10*time.Second, 50*time.Millisecond, "the retry opened the store") + _, err = dedupetest.Mark(context.Background(), store, eventKey) + require.NoError(t, err) + require.NoError(t, stop()) +} + +// bootLogged sends the default logger to a buffer for the rest of the test, +// for a boot that logs what it tolerated. +func bootLogged(t *testing.T) *lockedBuffer { + t.Helper() + guardGlobals(t) + buf := &lockedBuffer{} + slog.SetDefault(slog.New(slog.NewTextHandler(buf, nil))) + return buf +} + +// lockedBuffer is a bytes.Buffer safe for the background retry's logging. +type lockedBuffer struct { + mu sync.Mutex + buf bytes.Buffer +} + +func (b *lockedBuffer) Write(p []byte) (int, error) { + b.mu.Lock() + defer b.mu.Unlock() + return b.buf.Write(p) +} + +func (b *lockedBuffer) String() string { + b.mu.Lock() + defer b.mu.Unlock() + return b.buf.String() +} + // The reload hook runs under the lock that serializes reloads, so it never // calls DynamoDB: against a table that hangs, a reload returns at once, and a // tenant it switches on fails closed rather than publishing un-deduped. diff --git a/internal/app/wire_dynamodb.go b/internal/app/wire_dynamodb.go index 420111b58..68f68ce45 100644 --- a/internal/app/wire_dynamodb.go +++ b/internal/app/wire_dynamodb.go @@ -19,13 +19,15 @@ var errDynamoUnchecked = errors.New("dedupe: dynamodb table not checked yet") // every tenant and every process shares (dedupe.Dynamo), so a tenant's store // opens for free once the table has passed its check. Boot checks it (after // creating it, with create_table on dynamodb-local) whether or not any tenant -// has dedupe on, and never creates it otherwise. A table that fails the check -// follows the registry's rule for the shape, as Pebble's instance does: a -// flat directory refuses boot; a nested one boots with every switched-on -// store closed, so its ingest fails closed. Unlike a local disk, a remote -// table's failure is usually brief (a throttle, credentials not yet issued -// mid-rollout), and a nested directory has no watcher to reload it, so the -// check is then retried in the background, with backoff, until it passes. +// has dedupe on, and never creates it otherwise. Boot is refused only when the +// table is misconfigured (a failure that is not ErrUnavailable: missing, the +// wrong key schema, access denied) over a flat directory in which a tenant +// has dedupe on. Otherwise — a transient failure, a nested directory, or no +// tenant deduping yet — the process boots with every switched-on store +// closed, so its ingest fails closed, and the check is retried in the +// background, with backoff, until it passes: a remote table's failure is +// often brief, a nested directory has no watcher to reload it, and a fixed +// table is picked up without a restart. // The check is network I/O, so the AfterAdopt hook never runs it: the hook // holds the lock that serializes reloads. It applies every store against the // last check's result and wakes the retry, so a reload still retries at once. @@ -94,11 +96,17 @@ func (a *App) wireDynamoDedupe(ctx context.Context) error { } }) if err := check(ctx); err != nil { - if !a.tenants.Nested() { + misconfigured := !errors.Is(err, dedupe.ErrUnavailable) + if misconfigured && !a.tenants.Nested() && a.anyDedupeEnabled() { return fmt.Errorf("dedupe open: %w", err) } - slog.Error("dedupe: dynamodb table check failed; ingest with dedupe on fails closed while it is retried", - "table", c.Table, "error", err) + if misconfigured { + slog.Error("dedupe: dynamodb table is misconfigured; ingest with dedupe on fails closed until it is fixed", + "table", c.Table, "error", err) + } else { + slog.Error("dedupe: dynamodb table check failed; ingest with dedupe on fails closed while it is retried", + "table", c.Table, "error", err) + } a.add(component{name: "dedupe table check", run: func(ctx context.Context) error { for wait := time.Second; ready() != nil; wait = min(2*wait, 30*time.Second) { select { @@ -123,3 +131,13 @@ func (a *App) wireDynamoDedupe(ctx context.Context) error { apply() return nil } + +// anyDedupeEnabled reports whether a served tenant has dedupe switched on. +func (a *App) anyDedupeEnabled() bool { + for _, store := range a.tenants.All() { + if store.DedupeEnabled() { + return true + } + } + return false +} From 23ac57a3bdd3b88aef1dc5f257ea669b1b1c6313 Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 14:43:40 -0400 Subject: [PATCH 36/38] test(dedupe): order the cancel test's puts after the cancel, not a sleep No put applies before the caller cancels, so neither subtest depends on scheduling. The CHANGELOG line now names what a cancel can still leave held: a put cut off by its own timeout, or a failed release. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01FyrXjhR7iDg33paioLQHFq --- CHANGELOG.md | 2 +- internal/dedupe/dynamodb_test.go | 9 ++++++--- 2 files changed, 7 insertions(+), 4 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index b1f5e3615..0c9200533 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, retrying for up to eight jittered rounds both the items DynamoDB leaves unprocessed and a batch it throttled whole, and `Release` is a `DeleteItem` conditional on the claim's token. A client that disconnects mid-`Reserve` leaves nothing claimed: puts already sent run to their answer and are then released, so its retry is not answered `InFlight` for the lease. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. +- **A DynamoDB dedupe backend, not yet selectable** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, retrying for up to eight jittered rounds both the items DynamoDB leaves unprocessed and a batch it throttled whole, and `Release` is a `DeleteItem` conditional on the claim's token. A client that disconnects mid-`Reserve` no longer strands its claim: puts already sent run to their answer on a context its cancellation does not reach, and are then released, so its retry is not answered `InFlight` for the lease. Only a put cut off by its own call timeout (which DynamoDB may apply after the release), or a release that fails, still holds its id until the lease ends. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. No boot key selects it yet; that is F5. New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. - **Leases for work that must run in one process at a time, and the sweeper runs under one** (`internal/coord/` (new: `coord.go`, `local.go`, `elect.go`, `coordtest/`, + tests), `internal/app/{app,wire}.go` (+ tests), `internal/config/backends.go`, `internal/ingest/sweeper.go`, `.github/labeler.yml`, `.testcoverage.yml`, `docs/src/content/docs/{architecture,development,ingest-pipeline}.md`, `docs/src/content/docs/configuration.mdx`, `config.yaml`, `AGENTS.md`): part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `coord.Coordinator` hands out named leases (`TryAcquire` → a `Term` with a strictly increasing fencing `Token`, a `Done` channel and `Resign`; `ErrHeld` while another holder's term is live), and `coord.RunElected` runs a loop only while its process holds the lease, resigning when the loop returns and campaigning again every 2s. `coord.Local` is the in-process implementation, and `coordtest.Conformance` is the suite every implementation runs — the NATS KV backend that lets several replicas share one queue comes next. The sweeper now runs through `RunElected` under the `sweeper` lease; with the in-process coordinator the one process always holds it, so nothing changes for a single-process deployment beyond one `coord: elected` log line at startup. `coord.backend` now selects the coordinator (`local`, the only value), so a `config.Config` built without `config.Load` must name it as well as the other three layers' backends. diff --git a/internal/dedupe/dynamodb_test.go b/internal/dedupe/dynamodb_test.go index 99ac8b9a9..a61132590 100644 --- a/internal/dedupe/dynamodb_test.go +++ b/internal/dedupe/dynamodb_test.go @@ -636,7 +636,8 @@ func TestDynamo_CommitAttemptsEveryChunk(t *testing.T) { // A caller that cancels mid-Reserve (a client disconnecting) leaves nothing // claimed. The fake applies a put after a delay whatever the caller does, as // DynamoDB applies a request already on the wire, so a put abandoned on the -// cancel would land after its release and hold its id for the lease. +// cancel would land after its release and hold its id for the lease. No +// put applies before the cancel, so neither case depends on timing. func TestDynamo_CallerCancelLeavesNothingClaimed(t *testing.T) { t.Parallel() t.Run("a put still unsent", func(t *testing.T) { @@ -659,6 +660,7 @@ func assertCancelLeavesNothing(t *testing.T, ks []Key) { applies sync.WaitGroup ) started := make(chan struct{}, 2) + cancelled := make(chan struct{}) _, m := openFakeWith(t, &fakeDynamo{ put: func(ctx context.Context, in *dynamodb.PutItemInput) (*dynamodb.PutItemOutput, error) { id := idOf(in.Item[attrKey]) @@ -668,7 +670,8 @@ func assertCancelLeavesNothing(t *testing.T, ks []Key) { mu.Unlock() applied := make(chan struct{}) applies.Go(func() { - time.Sleep(20 * time.Millisecond) + <-cancelled + time.Sleep(5 * time.Millisecond) // lands after the caller left mu.Lock() table[id] = tk mu.Unlock() @@ -695,7 +698,7 @@ func assertCancelLeavesNothing(t *testing.T, ks []Key) { }, }, DynamoConfig{Table: "dedupe", ReserveConcurrency: 2}) ctx, cancel := context.WithCancel(t.Context()) - go func() { <-started; <-started; cancel() }() + go func() { <-started; <-started; cancel(); close(cancelled) }() _, err := m.Reserve(ctx, ks, time.Minute) require.ErrorIs(t, err, context.Canceled) applies.Wait() From 8ac2fca949789055d9b389de19f05f523c6e5bff Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 15:04:15 -0400 Subject: [PATCH 37/38] fix(dedupe): classify create_table errors; scope the boot rule in docs CreateTable returned its errors unclassified, so with create_table on a dynamodb-local not listening yet read as a misconfiguration and refused boot over a flat directory. Its errors now go through classify, and a test pins that a transient create failure boots and is retried. settings-directory.mdx still said any failed open refuses boot; that is now scoped to Pebble, with the DynamoDB rule stated beside it. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01FyrXjhR7iDg33paioLQHFq --- CHANGELOG.md | 2 +- docs/src/content/docs/settings-directory.mdx | 2 +- internal/app/dedupe_dynamodb_test.go | 21 ++++++++++++++++++++ internal/dedupe/dynamodb.go | 9 +++++---- 4 files changed, 28 insertions(+), 6 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 3adb62980..7fb2552b7 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -10,7 +10,7 @@ The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), ### Added -- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire,wire_dynamodb}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `.testcoverage.yml`, `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59s` with the embedded queue, so that the lease plus its own ceiling to the next second plus one more second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`, which also sizes the DynamoDB client's idle connections per host; its fan-out has no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a misconfigured table (missing, the wrong key schema, access denied) refuses boot over a flat settings directory whose tenant has dedupe on and is logged at `ERROR` otherwise; any other failure (a throttle, a timeout, the network), a nested directory, or no tenant deduping yet boots and fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and at once after every reload. A reload makes no table call and does not wait on a tenant whose dedupe setting is unchanged: it holds the lock that serializes reloads, so it applies each tenant's switch against the last check's result and only wakes the retry; it waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight calls, before its store closes. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. +- **`dedupe.backend: dynamodb` selects the shared DynamoDB dedupe table** (`internal/config/{backends,config}.go` (+ tests), `internal/app/{app,wire,wire_dynamodb}.go` (+ `dedupe_dynamodb_test.go`), `internal/dedupe/{stores,dynamodb}.go` (+ tests), `tests/integration/dedupe_dynamodb_app_test.go` (new), `.testcoverage.yml`, `config.yaml`, `AGENTS.md`, `docs/src/content/docs/{configuration.mdx,settings-directory.mdx,deployment.md,architecture.md,api.md,sdk/reference.md}`): PR F5 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). Pods that set it share seen ids, so an id ingested through one is a duplicate through every other. New boot keys: `dedupe.lease` (`WH_DEDUPE_LEASE`, `30s`, how long a claimed id stays pending and the in-flight `503`'s `Retry-After`; at most `59s` with the embedded queue, so that the lease plus its own ceiling to the next second plus one more second fits its 2-minute duplicate window: a client obeying that `Retry-After` after an uncertain publish can republish as late as the lease plus that ceiling plus a second after the claim, since DynamoDB rounds a claim's expiry up to the second), `dedupe.reserve_concurrency` (`WH_DEDUPE_RESERVE_CONCURRENCY`, `64`, which also sizes the DynamoDB client's idle connections per host; its fan-out has no effect until ingest sends more than one id per call), and the `dedupe.dynamodb` block (`table` (required), `region`, `endpoint`, `timeout` `250ms`, `max_attempts` `3`, `retry_mode` `standard`/`adaptive`, `create_table`), each with its `WH_DEDUPE_DYNAMODB_*` variable. Their defaults are in `defaults()` like every boot key's, so an explicit `0` lease, concurrency, timeout or attempt count, or an empty `retry_mode`, refuses boot rather than becoming the default (the `dynamodb` block's only while `dynamodb` is selected). Credentials come from the AWS SDK's default chain, never from config. Boot checks the table (key schema `pk` String alone; TTL off on `ex` is a warning) in a process running the `api` role, the one that opens the dedupe stores, whether or not a tenant has dedupe on: a misconfigured table (missing, the wrong key schema, access denied) refuses boot over a flat settings directory whose tenant has dedupe on and is logged at `ERROR` otherwise; any other failure (a throttle, a timeout, the network), a nested directory, or no tenant deduping yet boots and fails every switched-on tenant's ingest closed until the check passes, retried in the background (1s backing off to 30s) and at once after every reload. A reload makes no table call and does not wait on a tenant whose dedupe setting is unchanged: it holds the lock that serializes reloads, so it applies each tenant's switch against the last check's result and only wakes the retry; it waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight calls, before its store closes. No region from the config or the SDK chain refuses boot. `create_table` creates a missing table at boot (an endpoint not up yet is a transient failure, retried like the check) and is refused unless `endpoint` is set, so it only ever reaches dynamodb-local. - **A DynamoDB dedupe backend** (`internal/dedupe/dynamodb.go` (new, + tests), `tests/integration/{setup,dedupe_dynamodb}_test.go`, `go.mod`, `AGENTS.md`, `docs/src/content/docs/{architecture,deployment}.md`): PR F3 of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `dedupe.Dynamo` keeps every tenant's seen ids in one shared table, so pods that share the table also share seen ids, which the per-process Pebble store cannot do. The table's only key is a String `pk` holding the readable dedupe key (`acme/clicks/evt-123`), so an item reads as-is in the console. `Reserve` is a conditional `PutItem` that is atomic across pods; an SDK retry of a put DynamoDB applied but whose answer was lost (a `500`, a reset connection) finds its own item by token and keeps the claim, rather than answer `InFlight` and hold the id for the lease. `Commit` is `BatchWriteItem`, retrying for up to eight jittered rounds both the items DynamoDB leaves unprocessed and a batch it throttled whole, and `Release` is a `DeleteItem` conditional on the claim's token. A client that disconnects mid-`Reserve` does not strand its claim: puts already sent run to their answer on a context its cancellation does not reach, and are then released, so its retry is not answered `InFlight` for the lease. Only a put cut off by its own call timeout (which DynamoDB may apply after the release), or a release that fails, still holds its id until the lease ends. An expired item counts as absent without waiting for TTL. Each call has a 250 ms timeout covering the SDK's three attempts, whose retries back off with full jitter under a ceiling that keeps their waits within half the timeout, so a throttled call fails with the throttle as its cause rather than on the deadline. Throttling, timeouts and an unreachable table wrap `ErrUnavailable`, and five such failures in a row within a second short-circuit claims for a second. Credentials come from the AWS SDK's default chain. The HTTP client keeps one idle connection per host for each of the 64 calls a `Reserve`, `Commit` or `Release` runs at once (the SDK's default keeps 10), so a warm 64-key `Reserve` reuses every connection rather than open about 50. WaveHouse never creates the production table: `CreateTable` works only against dynamodb-local, and the Deployment page carries an example Terraform table and IAM policy. The backend passes the `dedupetest` conformance suite against a pinned `amazon/dynamodb-local` container, along with 32 clients racing one id, injected throttles and an unreachable endpoint. New metrics: `wavehouse_dedupe_dynamodb_requests_total`, `_request_duration_seconds`, `_unprocessed_items_total`, `_short_circuits_total`. `dedupe.backend: dynamodb` selects it (entry above). New dependencies: `aws-sdk-go-v2` (`service/dynamodb`, `config`) and what they require. - **One conformance suite for every `mq.Broker`, and a transient broker failure is a `503`** (`internal/mq/mqtest/` (new: the suite and the embedded broker's run of it), `internal/mq/{mq,embedded}.go` (+ tests), `internal/api/ingest.go` (+ tests), `.testcoverage.yml`, `docs/src/content/docs/{api,architecture}.md`, `AGENTS.md`): the first piece of the external-NATS workstream of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). `mqtest.Run` states the `Broker` contract as behavior — round trips with names that need encoding, per-tenant order, `Nak` and `AckWait` redelivery, the trace context reaching `Subscribe`, dead-lettering that keeps the topic and leaves the original unacked, per-tenant dead-letter counts, replay bounds and isolation, and exactly one `failed` report when delivery ends underneath a consumer — through the interfaces alone, so the external backend runs the same cases, with `mqtest.Caps` for the four places where its semantics legitimately differ. The embedded broker passes it; writing it turned up that a durable deleted on several tenants' queues could report on `failed` more than once, which is fixed and pinned by a test that deletes it on one queue after another. It also turned up a replay that lost its connection mid-pull passing for a caught-up one when the pull ended in a timeout; that is an error now, as the `Replayer` contract says. The interface comments now allow a delivery unit that is a partition holding several tenants, a `CreateConsumer` that finds a durable rather than creating one, a `PurgeAcked` that leaves retention to the operator, and zero dead-letter counts where there is no per-tenant queue. A new sentinel, `mq.ErrUnavailable`, is a broker that cannot be reached or does not answer in time: the ingest handler answers it with `503` + `Retry-After: 5` rather than the `500` "publish failed" it would have been. Nothing returns it yet; the external backend of #613 will. - **Process roles: the API and the background workers can run in separate processes** (`internal/config/config.go` (+ `roles_test.go`, `defaults_test.go`), `internal/config/backends.go`, `internal/app/{app,wire}.go` (+ `roles_test.go`), `internal/api/router.go` (+ tests), `tests/integration/{setup,tenants}_test.go`, `config.yaml`, `docs/src/content/docs/{configuration.mdx,deployment.md,architecture.md}`, `AGENTS.md`), part of [#613](https://github.com/Wave-RF/WaveHouse/issues/613). The boot config gains `roles` (`WH_ROLES`, default `api,ingest,sweeper`, set in `defaults()` like every boot default, so an explicit `roles: []` refuses boot) and `instance_id` (`WH_INSTANCE_ID`, default `-<8 hex>`, fresh at every boot; logged at boot, and recorded as a lease holder once a shared `coord.backend` exists). A process wires only what its roles need: `api` runs the HTTP API with schema discovery, the token verifiers, the dedupe stores and the SSE hub (all per API process); `ingest` runs the ingest worker; `sweeper` runs the sweeper under its lease. A process without `api` serves an ops-only listener on `server.port` (the probes and their aliases, `/version`, the same-port metrics path, and `POST /v1/ops/settings/reload`, which takes the operator key alone); every other route answers 404, under `/v1/ops` once the operator key has passed. Boot refuses any split over the embedded MQ, which no other process can reach, and `api` without `ingest` (or the reverse) over a local cache, which the ingest worker's invalidations would never reach. Until a shared `mq.backend` exists, every process therefore runs every role, which is the default, so nothing changes for an existing deployment. `data_dir` is probed for Pebble only in a process running `api`. A `config.Config` built without `config.Load` must now name its roles (`config.AllRoles()` for all of them): `app.New` refuses an empty set. diff --git a/docs/src/content/docs/settings-directory.mdx b/docs/src/content/docs/settings-directory.mdx index e1f0be888..f69d2bbce 100644 --- a/docs/src/content/docs/settings-directory.mdx +++ b/docs/src/content/docs/settings-directory.mdx @@ -185,7 +185,7 @@ What stays in boot config is only what cannot change under a running process — Every per-tenant dedupe knob lives here. Where the seen ids are kept (`dedupe.backend`) and how long a claim is held (`dedupe.lease`) are [boot config](/configuration#dedupe), the same for every tenant. The switch and its fields are resolved per record from one snapshot (table override → global value): -- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store (in the embedded Pebble instance at `/pebble`, or its share of the DynamoDB table under `dedupe.backend: dynamodb`), so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until the next reload or restart — the files asked for dedupe, so publishing un-deduped is not a fallback. At boot a failed open refuses to start, like every other store. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) with `dedupe.backend: pebble`, every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. Under `dedupe.backend: dynamodb` the table check plays the instance's part: the table is checked whether or not any tenant's switch is on, and a table that fails it fails every tenant with dedupe on closed until the check, retried in the background and at once after every reload, passes ([Configuration](/configuration#dynamodb-dedupe)). +- `dedupe.enabled` (seed default `false`) — turns deduplication on. Hot-reloadable: a reload that flips it opens or closes this tenant's store (in the embedded Pebble instance at `/pebble`, or its share of the DynamoDB table under `dedupe.backend: dynamodb`), so no restart is needed; seen ids persist across an off/on cycle. If the store fails to open on a reload, the failure is logged and ingest fails closed (`500 dedupe failed`) until it opens — the files asked for dedupe, so publishing un-deduped is not a fallback. With `dedupe.backend: pebble` that is the next reload or restart, and at boot a failed open over a flat directory refuses to start, like every other store; with `dynamodb` it is the background retry described below. A record that lands in the instant of the flip itself is published un-deduped: if the settings already say on but the store is not yet open, it's counted by `wavehouse_ingest_dedupe_disabled_total`; in the reverse case (settings already say off, store still open) the handler skips dedupe like any other disabled record and nothing is counted. That counter should only ever tick during a reload, so a steadily climbing rate means the store and the settings have come apart. Over [a nested directory](/deployment#the-nested-settings-directory) with `dedupe.backend: pebble`, every tenant's seen ids live in that one instance, each key led by its tenant and table, and it is open while any tenant's switch is on: each tenant's store follows its own folder's `dedupe.enabled` the same way; a tenant's seen ids are never another's; a rejected or removed folder closes its tenant's store and keeps its seen ids for the folder that restores it; and if that instance fails to open, at boot or on reload, every tenant with dedupe on fails closed — its ingest answers `500 dedupe failed` until a reload opens it — while the tenants with dedupe off carry on. Under `dedupe.backend: dynamodb` the table check plays the instance's part, in either shape: the table is checked whether or not any tenant's switch is on, and a table that fails it fails every tenant with dedupe on closed until the check, retried in the background and at once after every reload, passes. Only a misconfigured table (missing, the wrong key schema, access denied) over a flat directory whose tenant has dedupe on refuses boot instead ([Configuration](/configuration#dynamodb-dedupe)). - `dedupe.id_field` (seed default `event_id`) — JSON field name in the ingest body used as the dedup key. An id is a duplicate only within its own tenant and table: the same value in two tables is two ids. An id longer than 1,024 bytes once escaped (every byte but an ASCII letter, digit, `_` or `-` takes three) is stored as its SHA-256, counted by `wavehouse_dedupe_hashed_id_total`. While its record is being published, an id is held for its lease ([`dedupe.lease`](/configuration#dedupe), 30 seconds by default): another request carrying the same id meanwhile gets `503` (`a request with the same dedupe id is in flight`) with the lease, in whole seconds, as `Retry-After` — see [the ingest errors](/api#post-v1ingesttabletable--ingest-data). An id is committed only after its record is published; if that commit fails, the record is still answered `ok`, the id lapses with its lease, and a later retry of it is accepted again — counted by `wavehouse_ingest_dedupe_commit_failed_total`, which should stay at zero. - `dedupe.require_id` (seed default `false`) — controls what happens to a row missing `id_field`, or carrying it as `null` (which can't be deduped, so idempotency wouldn't apply to it). Such a row is always logged at `WARN` and counted by `wavehouse_ingest_dedupe_missing_id_total`, in both modes. `false`: it is then published un-deduped. `true` rejects it instead (`400` for a single insert; a per-record failure in a batch) — a server-side tripwire for producers that must guarantee the id. - `dedupe.tables.
.{id_field, require_id}` — per-table overrides; each entry overrides only the fields it names and inherits the rest. diff --git a/internal/app/dedupe_dynamodb_test.go b/internal/app/dedupe_dynamodb_test.go index b2ab5ab53..e9dd5290c 100644 --- a/internal/app/dedupe_dynamodb_test.go +++ b/internal/app/dedupe_dynamodb_test.go @@ -237,6 +237,27 @@ func TestRun_DynamoDBDedupeFlatThrottledRecovers(t *testing.T) { require.NoError(t, stop()) } +// With create_table on, an endpoint that fails transiently (dynamodb-local +// still starting) boots too, and the retry creates the table once it answers. +func TestRun_DynamoDBDedupeFlatCreateTableRetries(t *testing.T) { + cfg := testConfig(t, writeSettings(t, dedupeOn)) + fake := dynamoConfig(t, cfg, false) + cfg.Dedupe.DynamoDB.CreateTable = true + fake.setThrottles(true) + var lc net.ListenConfig + ln, err := lc.Listen(t.Context(), "tcp", "127.0.0.1:0") + require.NoError(t, err) + a := newApp(t, cfg, Options{Listener: ln}) + store := a.dedup.For(tenant.Default) + require.False(t, store.Open()) + + _, stop := runApp(t, a, ln) + fake.setThrottles(false) + require.Eventually(t, store.Open, 10*time.Second, 50*time.Millisecond, "the retry created the table and opened the store") + assert.True(t, fake.called("CreateTable")) + require.NoError(t, stop()) +} + // bootLogged sends the default logger to a buffer for the rest of the test, // for a boot that logs what it tolerated. func bootLogged(t *testing.T) *lockedBuffer { diff --git a/internal/dedupe/dynamodb.go b/internal/dedupe/dynamodb.go index 8117734ab..18f21232f 100644 --- a/internal/dedupe/dynamodb.go +++ b/internal/dedupe/dynamodb.go @@ -267,7 +267,8 @@ func (d *Dynamo) Check(ctx context.Context) error { var ErrCreateTableNeedsEndpoint = errors.New("dedupe: create_table is for dynamodb-local only; set the endpoint") // CreateTable creates the table on dynamodb-local, with TTL on ex, and waits -// for it. A table that already exists is left as it is. +// for it. A table that already exists is left as it is. Its errors are +// classified as every call's are, so an endpoint not up yet is ErrUnavailable. func (d *Dynamo) CreateTable(ctx context.Context) error { if d.cfg.Endpoint == "" { return ErrCreateTableNeedsEndpoint @@ -283,17 +284,17 @@ func (d *Dynamo) CreateTable(ctx context.Context) error { return nil } if err != nil { - return fmt.Errorf("dedupe: create table %s: %w", d.cfg.Table, err) + return fmt.Errorf("dedupe: create table %s: %w", d.cfg.Table, classify("create_table", err)) } if err := dynamodb.NewTableExistsWaiter(d.api).Wait(ctx, &dynamodb.DescribeTableInput{TableName: &d.cfg.Table}, time.Minute); err != nil { - return fmt.Errorf("dedupe: wait for table %s: %w", d.cfg.Table, err) + return fmt.Errorf("dedupe: wait for table %s: %w", d.cfg.Table, classify("describe_table", err)) } _, err = d.api.UpdateTimeToLive(ctx, &dynamodb.UpdateTimeToLiveInput{ TableName: &d.cfg.Table, TimeToLiveSpecification: &types.TimeToLiveSpecification{AttributeName: aws.String(attrExpiry), Enabled: aws.Bool(true)}, }) if err != nil { - return fmt.Errorf("dedupe: enable ttl on %s: %w", d.cfg.Table, err) + return fmt.Errorf("dedupe: enable ttl on %s: %w", d.cfg.Table, classify("update_time_to_live", err)) } return nil } From 6096ddd61bb40e7f2a25bb86834addf8872c40bb Mon Sep 17 00:00:00 2001 From: Eric Andrechek Date: Sat, 26 Sep 2026 15:13:23 -0400 Subject: [PATCH 38/38] docs(deployment): a mismatched key schema follows the misconfiguration rule The table section still said boot refuses a table whose key schema does not match, in every case. Drop a sentence configuration.mdx said twice. Co-Authored-By: Claude Opus 5.5 (1M context) Claude-Session: https://claude.ai/code/session_01FyrXjhR7iDg33paioLQHFq --- docs/src/content/docs/configuration.mdx | 2 +- docs/src/content/docs/deployment.md | 2 +- 2 files changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/src/content/docs/configuration.mdx b/docs/src/content/docs/configuration.mdx index 4c23d2d97..0895d4ea8 100644 --- a/docs/src/content/docs/configuration.mdx +++ b/docs/src/content/docs/configuration.mdx @@ -61,7 +61,7 @@ Whether a tenant dedupes, and on which field, are settings-directory keys ([Dedu #### DynamoDB dedupe -Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (String) alone, and TTL off on `ex` is logged as a warning. A misconfigured table — missing, with the wrong key schema, or denied to the process's credentials — refuses boot only with a flat settings directory whose tenant has dedupe on, and is logged at `ERROR` otherwise. In every other case — a transient failure (a throttle, a timeout, the network), a nested directory, or no tenant with dedupe on yet — the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and the check is retried in the background, backing off from one second to thirty, and at once after every reload, so a table that comes good is picked up without a restart. A tenant a reload switches dedupe on for fails closed the same way until then. A reload makes no table call, and does not wait on a tenant whose dedupe setting is unchanged: it applies each tenant's switch against the last check's result, so a tenant it switches on fails closed until the retry passes. A reload waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight `Reserve`/`Commit`/`Release` calls, before the store itself closes. The check runs in every process running the `api` [role](#process-roles), the one that opens the dedupe stores, whether or not any tenant has dedupe on. +Read only when `dedupe.backend` is `dynamodb`. Credentials come from the AWS SDK's default chain (EKS Pod Identity or IRSA in a pod; `AWS_*` variables or a profile locally), never from this file; the table and its IAM policy are described in [Deployment](/deployment#a-shared-dedupe-table-on-dynamodb). At boot WaveHouse checks the table: its key schema must be `pk` (String) alone, and TTL off on `ex` is logged as a warning. A misconfigured table — missing, with the wrong key schema, or denied to the process's credentials — refuses boot only with a flat settings directory whose tenant has dedupe on, and is logged at `ERROR` otherwise. In every other case — a transient failure (a throttle, a timeout, the network), a nested directory, or no tenant with dedupe on yet — the process boots, every tenant with dedupe on answers ingest with an error until the check passes, and the check is retried in the background, backing off from one second to thirty, and at once after every reload, so a table that comes good is picked up without a restart. A reload makes no table call, and does not wait on a tenant whose dedupe setting is unchanged: it applies each tenant's switch against the last check's result, so a tenant it switches on fails closed until the retry passes. A reload waits only for a tenant whose store it closes — dedupe switched off, or the tenant removed or rejected — and then only for that tenant's in-flight `Reserve`/`Commit`/`Release` calls, before the store itself closes. The check runs in every process running the `api` [role](#process-roles), the one that opens the dedupe stores, whether or not any tenant has dedupe on. | YAML Key | Env Var | Default | Description | | --- | --- | ------- | ----------- | diff --git a/docs/src/content/docs/deployment.md b/docs/src/content/docs/deployment.md index 1aea0f34e..a078d2a24 100644 --- a/docs/src/content/docs/deployment.md +++ b/docs/src/content/docs/deployment.md @@ -454,7 +454,7 @@ What the backend requires of the table: | `ex` | Number | Epoch seconds: the lease end while pending, the retention end once committed; absent = never expires. | | `tk` | Binary | The claim token that `Release` matches. | -Only `pk` is declared in the table definition. Turn TTL on for `ex`. Correctness never depends on TTL, because a claim whose `ex` has passed counts as absent whether or not DynamoDB has deleted it yet; TTL only reclaims the storage. **Today TTL removes only lapsed claims:** ingest commits every id with no retention, so a committed item carries no `ex` and is kept forever, and the table grows by one item (about 200 bytes) per distinct id. Per-tenant retention is [#220](https://github.com/Wave-RF/WaveHouse/issues/220). Boot checks the table: it refuses one whose key schema does not match, and logs a warning if TTL is off. +Only `pk` is declared in the table definition. Turn TTL on for `ex`. Correctness never depends on TTL, because a claim whose `ex` has passed counts as absent whether or not DynamoDB has deleted it yet; TTL only reclaims the storage. **Today TTL removes only lapsed claims:** ingest commits every id with no retention, so a committed item carries no `ex` and is kept forever, and the table grows by one item (about 200 bytes) per distinct id. Per-tenant retention is [#220](https://github.com/Wave-RF/WaveHouse/issues/220). Boot checks the table and logs a warning if TTL is off; a key schema that does not match is a misconfigured table, handled as described below. An example in Terraform. Replace the tags with your own conventions: